Senior Staff Site Reliability Engineer
Bengaluru
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 10+ yearsEducation: bachelorsSkills: ["Problem-solving","Communication","Collaboration","Mentoring","Ownership","Curiosity","Innovation"]Own the technical strategy and roadmap for large-scale SRE initiatives to improve reliability, scalability, and developer efficiency across NVIDIA enterprise systems. Design and modernize resilient distributed architectures, drive automation and observability (including AI workload signals), and build LLM-aware monitoring with autonomous incident response. Partner with Cloud, Platform, Security, and AI/ML teams to deliver high-availability, secure operations, and mentor engineers on AI-assisted engineering practices.
Loading
Loading job details...
Preparing the role view and application actions.

