SRE Reliability Enginner

NTT
Bengaluru
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 5+ yearsSkills: ["Troubleshooting","Analytical skills","Stakeholder communication","Incident coordination","Root-cause analysis"]

Own reliability and operational health for business-critical production services, providing L2/L3 incident triage, troubleshooting, and stakeholder communication. Monitor Kubernetes-deployed Java/J2EE applications using Datadog and/or Prometheus, define SLOs/SLIs/error budgets, and drive MTTR and availability improvements. Perform performance analysis across JVM and databases using SQL, automate recurring operational work, and support releases, deployments, and post-deployment validation.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NTT
NTT
1 week ago

SRE Reliability Enginner

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 13 hours agoStatus: Live

Job Summary

Own reliability and operational health for business-critical production services, providing L2/L3 incident triage, troubleshooting, and stakeholder communication. Monitor Kubernetes-deployed Java/J2EE applications using Datadog and/or Prometheus, define SLOs/SLIs/error budgets, and drive MTTR and availability improvements. Perform performance analysis across JVM and databases using SQL, automate recurring operational work, and support releases, deployments, and post-deployment validation.
Location: Bengaluru
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Own reliability, availability, and operational health of business-critical production applications and services.
  • •Provide L2/L3 production support including incident triage, troubleshooting, resolution, and stakeholder communication.
  • •Monitor and support Kubernetes-deployed applications, including pods, deployments, services, ingress, utilization, scaling, and cluster-related issues.
  • •Implement and maintain monitoring using Datadog/Prometheus, including dashboards, metrics, logs, alerts, and service-health indicators.
  • •Troubleshoot production issues across Java applications, APIs, microservices, Kubernetes, databases, and infrastructure; perform RCA and post-incident reviews.
  • •Identify recurring production issues and drive permanent remediation through automation and engineering improvements.
  • •Support incident management, major incident calls, and automation of repetitive operational activities.
  • •Support releases, production deployments, rollback activities, and post-deployment validation.

Key Requirements

  • •5+ years of overall IT experience with significant SRE, production support, application support, or DevOps experience.
  • •Strong hands-on experience with Kubernetes and containerized applications in production.
  • •Experience with Datadog and/or Prometheus for monitoring, alerting, metrics, and observability.
  • •Strong production support experience for Java/J2EE or Java-based microservices, including JVM troubleshooting and performance analysis.
  • •Strong SQL skills for investigating relational databases and production incidents, including incident coordination, RCA, and problem management.
Experience:5+ yearsSREProduction supportDevOpsKubernetesMicroservices
Skills:TroubleshootingAnalytical skillsStakeholder communicationIncident coordinationRoot-cause analysis
Tech Stack:KubernetesDatadogPrometheusJavaJ2EESQLJVMREST APIsMicroservicesLinux/UnixShell scriptingCI/CDAWSAzureDockerHelmGitLabGitHub ActionsELKOpenSearch

Company Brief

NTT
Dimension Data, operating under NTT Ltd, provides managed IT services, cloud and data center solutions, networking, cybersecurity, and digital transformation services to enterprise customers worldwide.
Industry: Professional Services
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Valuation: Public Company (Market Cap in USD)
Headquarters: London, United Kingdom
Founded: 1983
WebsiteLinkedIn