Senior Staff Site Reliability Engineer
Bengaluru, Pune
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 10+ yearsEducation: bachelorsSkills: ["Incident leadership","Cross-team coordination","Executive communication","Technical leadership","Mentoring"]Lead end-to-end incident response as an Incident Commander while setting technical direction for reliability engineering across NVIDIA’s AI-powered enterprise platforms. Design and operate distributed, Kubernetes- and cloud-native systems, improve observability and signal quality, and build automation that replaces manual runbooks with self-healing processes. Apply AI/data-driven techniques for incident triage and decision support, partner with Cloud/Platform/Security/AI teams on SLOs, and mentor SRE talent as the India practice scales.
Loading
Loading job details...
Preparing the role view and application actions.

