SRE Reliability Engineer

NTT
Bengaluru
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 5+ yearsSkills: ["Troubleshooting","Analytical thinking","Incident coordination","Root-cause analysis","Stakeholder communication"]

Own the reliability, availability, and operational health of business-critical production services as an SRE. Provide L2/L3 production support, incident triage, troubleshooting, and RCA while improving MTTR, SLO adherence, and production stability. Monitor Kubernetes-deployed applications using Datadog/Prometheus, define SLIs/SLOs, and automate remediation and runbooks. Diagnose issues across Java/J2EE services, APIs, microservices, and SQL-backed data layers.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NTT
NTT
1 week ago

SRE Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 13 hours agoStatus: Live

Job Summary

Own the reliability, availability, and operational health of business-critical production services as an SRE. Provide L2/L3 production support, incident triage, troubleshooting, and RCA while improving MTTR, SLO adherence, and production stability. Monitor Kubernetes-deployed applications using Datadog/Prometheus, define SLIs/SLOs, and automate remediation and runbooks. Diagnose issues across Java/J2EE services, APIs, microservices, and SQL-backed data layers.
Location: Bengaluru
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Own reliability, availability, and operational health for business-critical production applications and services.
  • •Provide L2/L3 production support, including incident triage, troubleshooting, resolution, and stakeholder communication.
  • •Monitor and support applications on Kubernetes (pods, deployments, services, ingress, scaling, and cluster issues).
  • •Implement and maintain observability with Datadog/Prometheus, including dashboards, metrics, logs, alerts, and service health indicators.
  • •Lead incident management, RCA, and post-incident reviews; drive permanent remediation and automation to reduce MTTR and production toil.

Key Requirements

  • •5+ years of IT experience with significant SRE, production support, application support, or DevOps experience.
  • •Strong hands-on experience with Kubernetes and containerized applications.
  • •Experience with Datadog and/or Prometheus for monitoring, alerting, metrics, and observability.
  • •Strong experience supporting Java/J2EE or Java-based microservices applications in production, including JVM troubleshooting.
  • •Strong SQL skills to troubleshoot relational databases and production incidents, including incident coordination and RCA.
Experience:5+ yearsSREProduction supportDevOpsMicroservices
Skills:TroubleshootingAnalytical thinkingIncident coordinationRoot-cause analysisStakeholder communication
Tech Stack:KubernetesDatadogPrometheusJavaJ2EEJVMSQLLinux/UnixShell scriptingREST APIsMicroservicesDistributed systemsCI/CDGitLabGitHub ActionsDockerHelmELKOpenSearchSplunk

Company Brief

NTT
Dimension Data, operating under NTT Ltd, provides managed IT services, cloud and data center solutions, networking, cybersecurity, and digital transformation services to enterprise customers worldwide.
Industry: Professional Services
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Valuation: Public Company (Market Cap in USD)
Headquarters: London, United Kingdom
Founded: 1983
WebsiteLinkedIn