Senior Staff Site Reliability Engineer

NVIDIA
Bengaluru, Pune
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 10+ yearsEducation: bachelorsSkills: ["Incident leadership","Cross-team coordination","Executive communication","Technical leadership","Mentoring"]

Lead end-to-end incident response as an Incident Commander while setting technical direction for reliability engineering across NVIDIA’s AI-powered enterprise platforms. Design and operate distributed, Kubernetes- and cloud-native systems, improve observability and signal quality, and build automation that replaces manual runbooks with self-healing processes. Apply AI/data-driven techniques for incident triage and decision support, partner with Cloud/Platform/Security/AI teams on SLOs, and mentor SRE talent as the India practice scales.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
1 day ago

Senior Staff Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 15 hours agoStatus: Live

Job Summary

Lead end-to-end incident response as an Incident Commander while setting technical direction for reliability engineering across NVIDIA’s AI-powered enterprise platforms. Design and operate distributed, Kubernetes- and cloud-native systems, improve observability and signal quality, and build automation that replaces manual runbooks with self-healing processes. Apply AI/data-driven techniques for incident triage and decision support, partner with Cloud/Platform/Security/AI teams on SLOs, and mentor SRE talent as the India practice scales.
Location: Bengaluru, Pune
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Lead major incidents end to end, including triage, cross-team coordination, decision-making, and executive communication across global time zones.
  • •Set technical direction for reliability engineering initiatives to improve reliability, scalability, and developer efficiency and drive adoption.
  • •Design, build, and operate distributed, Kubernetes-based and cloud-native systems that power NVIDIA’s AI-powered enterprise products and services.
  • •Build automation for incident detection, triage, communication, and remediation to replace manual runbooks with self-healing systems.
  • •Improve observability and drive root cause analysis into systemic fixes, prevention mechanisms, and AI-assisted triage and decision support.

Key Requirements

  • •10+ years of experience in Site Reliability Engineering, Production Engineering, Platform Engineering, or Incident Management with technical leadership at scale.
  • •BS or MS in Computer Science, Engineering, or related field (or equivalent practical experience).
  • •Proven experience acting as an Incident Commander or leading major incident response in complex, high-availability environments.
  • •Strong distributed systems and reliability engineering expertise including SLIs/SLOs, error budgets, capacity planning, and graceful degradation.
  • •Strong hands-on skills in programming and cloud/container/infrastructure tooling (e.g., Python/Go/Java, AWS/Azure/GCP, Docker/Kubernetes, Terraform/AWS CDK/CloudFormation, CI/CD, Linux/Unix, observability tools).
Experience:10+ yearsSite reliability engineeringProduction engineeringPlatform engineeringIncident managementDistributed systemsCloud-nativeObservabilityAI/ML operations
Education:Bachelor's in Computer Science, Engineering, or related technical field
Skills:Incident leadershipCross-team coordinationExecutive communicationTechnical leadershipMentoring
Tech Stack:PythonGoJavaKubernetesDockerAWSAzureGCPTerraformAWS CDKCloudFormationCI/CDLinux/UnixOpenTelemetryPrometheusGrafanaPostgreSQLMySQLSQLLLMs

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor