Site Reliability Engineer

NTT
Hyderabad
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureEducation: bachelorsSkills: ["Problem-solving","Analytical skills","Communication","Collaboration","Leadership"]

Ensure the reliability, availability, and performance of company systems by monitoring health and performance metrics, responding to incidents, and driving root-cause analysis. Build automation and self-healing capabilities using scripting and infrastructure-as-code to improve resiliency and reduce downtime. Optimize scalability and performance with monitoring and testing tools, collaborate across development and operations, and partner with security teams to implement best practices and compliance controls.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NTT
NTT
4 days ago

Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 23 hours agoStatus: Live

Job Summary

Ensure the reliability, availability, and performance of company systems by monitoring health and performance metrics, responding to incidents, and driving root-cause analysis. Build automation and self-healing capabilities using scripting and infrastructure-as-code to improve resiliency and reduce downtime. Optimize scalability and performance with monitoring and testing tools, collaborate across development and operations, and partner with security teams to implement best practices and compliance controls.
Location: Hyderabad
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure

Key Responsibilities

  • •Monitor system health, performance metrics, and alerts to diagnose and resolve incidents promptly and restore services.
  • •Implement incident response processes to minimize downtime and improve system availability.
  • •Design, develop, and maintain automation tools, scripts, and processes for deployments and configuration changes.
  • •Apply infrastructure-as-code and configuration standards to ensure consistency and repeatability across environments.
  • •Lead and coordinate incident response and post-incident reviews to identify root causes and implement preventive measures, including collaborating with security teams on vulnerabilities and compliance.

Key Requirements

  • •Seasoned hands-on experience in a Site Reliability Engineering role (or related) designing and maintaining highly available, scalable systems.
  • •Strong Linux/Unix, networking, and system administration expertise.
  • •Proficiency in scripting/programming languages such as Python, Go, Java, or Ruby; Bash or PowerShell is beneficial.
  • •Knowledge of cloud platforms (AWS, Azure, Google Cloud) and associated services.
  • •Experience with infrastructure-as-code (e.g., Terraform, CloudFormation), containerization (Docker, Kubernetes), CI/CD, and incident management with post-incident reviews.
Education:Bachelor's in Computer Science, Information Technology, or a related field
Skills:Problem-solvingAnalytical skillsCommunicationCollaborationLeadership
Certifications:AWS Certified DevOps Engineer - ProfessionalGoogle Cloud Professional DevOps EngineerCertified Kubernetes Administrator (CKA)
Tech Stack:Linux/UnixPythonGoJavaRubyAWSAzureGoogle CloudPrometheusGrafanaNew RelicBashPowerShellInfrastructure-as-CodeTerraformCloudFormationDockerKubernetesCI/CDJenkins

Company Brief

NTT
Dimension Data, operating under NTT Ltd, provides managed IT services, cloud and data center solutions, networking, cybersecurity, and digital transformation services to enterprise customers worldwide.
Industry: Professional Services
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Valuation: Public Company (Market Cap in USD)
Headquarters: London, United Kingdom
Founded: 1983
WebsiteLinkedIn