SRE Reliability Engineer

NTT
Bengaluru
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 5+ yearsSkills: ["Incident coordination","Root-cause analysis","Problem management","Analytical troubleshooting","Stakeholder communication","Automation mindset"]

Own the reliability, availability, and operational health of business-critical production systems. Provide L2/L3 production support, triaging incidents, troubleshooting Kubernetes-deployed Java applications and microservices, and driving RCA and post-incident reviews. Implement observability with Datadog/Prometheus, define SLOs/SLIs/error budgets, use SQL for incident investigation, and automate operational toil to improve MTTR, performance, capacity, and stability.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NTT
NTT
1 week ago

SRE Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 13 hours agoStatus: Live

Job Summary

Own the reliability, availability, and operational health of business-critical production systems. Provide L2/L3 production support, triaging incidents, troubleshooting Kubernetes-deployed Java applications and microservices, and driving RCA and post-incident reviews. Implement observability with Datadog/Prometheus, define SLOs/SLIs/error budgets, use SQL for incident investigation, and automate operational toil to improve MTTR, performance, capacity, and stability.
Location: Bengaluru
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Own reliability, availability, and operational health of business-critical production applications and services.
  • •Provide L2/L3 production support for incident triage, troubleshooting, resolution, and stakeholder communication.
  • •Monitor and support Kubernetes deployments (pods, deployments, services, ingress, resource utilization, scaling, and cluster issues).
  • •Implement and maintain observability with Datadog/Prometheus, including dashboards, metrics, logs, alerts, and SLO/SLI/error budget tracking.
  • •Troubleshoot production issues across Java applications, APIs/microservices, databases, and infrastructure; lead RCA and post-incident reviews, plus drive permanent remediation and automation.

Key Requirements

  • •5+ years of overall IT experience with significant experience in SRE, production support, application support, or DevOps.
  • •Hands-on experience with Kubernetes and containerized applications.
  • •Experience with Datadog and/or Prometheus for monitoring, alerting, metrics, and observability.
  • •Production support experience with Java/J2EE or Java-based microservices, including JVM troubleshooting and performance issues.
  • •Strong SQL skills for troubleshooting relational databases and diagnosing data-related issues during incidents.
Experience:5+ years
Skills:Incident coordinationRoot-cause analysisProblem managementAnalytical troubleshootingStakeholder communicationAutomation mindset
Tech Stack:KubernetesDatadogPrometheusJavaJ2EESQLJVMLinuxUnixREST APIsMicroservicesDistributed systemsCI/CDRelease managementAWSAzureDockerHelmGitLabGitHub Actions

Company Brief

NTT
Dimension Data, operating under NTT Ltd, provides managed IT services, cloud and data center solutions, networking, cybersecurity, and digital transformation services to enterprise customers worldwide.
Industry: Professional Services
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Valuation: Public Company (Market Cap in USD)
Headquarters: London, United Kingdom
Founded: 1983
WebsiteLinkedIn