Senior Associate Site Reliability Engineer

NTT
Hyderabad
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureEducation: bachelorsSkills: ["Problem-solving","Analytical skills","Communication","Collaboration","Automation","Continuous improvement"]

Play a key role in ensuring reliability, availability, and performance of company systems and infrastructure. You’ll monitor health and performance, respond to incidents, support deployments and infrastructure changes, and collaborate with development and operations teams to reduce downtime. You’ll help automate routine tasks, assist with capacity planning, document incidents and drive root-cause analysis, and partner with security teams to apply security best practices.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NTT
NTT
4 days ago

Senior Associate Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 23 hours agoStatus: Live

Job Summary

Play a key role in ensuring reliability, availability, and performance of company systems and infrastructure. You’ll monitor health and performance, respond to incidents, support deployments and infrastructure changes, and collaborate with development and operations teams to reduce downtime. You’ll help automate routine tasks, assist with capacity planning, document incidents and drive root-cause analysis, and partner with security teams to apply security best practices.
Location: Hyderabad
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Monitor system health, performance metrics, and alerts to identify and respond to incidents promptly.
  • •Diagnose and troubleshoot issues with senior SREs and teams, and restore services in a timely manner.
  • •Assist with deployment and release of software applications and infrastructure changes while minimizing downtime.
  • •Automate routine tasks with senior SREs to improve operational efficiency and reliability.
  • •Participate in incident documentation, post-incident reviews, root cause analysis, and preventive action implementation.

Key Requirements

  • •Moderate hands-on experience in a Site Reliability Engineering role or related roles, including designing and maintaining highly available, scalable systems.
  • •Moderate experience in incident response and troubleshooting to identify and resolve system issues.
  • •Moderate experience with automation principles and tools, such as Terraform and Jenkins, along with Git for version control.
  • •Developing knowledge of programming/scripting (Python, Bash, PowerShell) and Linux/Unix command-line tools.
  • •Bachelor’s degree (or equivalent) in Computer Science, Information Technology, or a related field; relevant certifications such as AWS Certified DevOps Engineer - Professional, Google Cloud Professional DevOps Engineer, or CKA preferred.
Experience:SRECloudIncident responseAutomation
Education:Bachelor's in Computer Science, Information Technology, or a related field
Skills:Problem-solvingAnalytical skillsCommunicationCollaborationAutomationContinuous improvement
Certifications:AWS Certified DevOps Engineer - ProfessionalGoogle Cloud Professional DevOps EngineerCertified Kubernetes Administrator (CKA)
Tech Stack:AWSAzureGoogle CloudPythonBashPowerShellGitLinux/UnixPrometheusGrafanaNew RelicTerraformJenkinsKubernetes

Company Brief

NTT
Dimension Data, operating under NTT Ltd, provides managed IT services, cloud and data center solutions, networking, cybersecurity, and digital transformation services to enterprise customers worldwide.
Industry: Professional Services
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Valuation: Public Company (Market Cap in USD)
Headquarters: London, United Kingdom
Founded: 1983
WebsiteLinkedIn