Site Reliability Engineer
Bengaluru
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureEducation: bachelorsSkills: ["Problem-solving","Curiosity","Communication","Teamwork","Ownership"]Support SRE initiatives to improve reliability, scalability, and developer efficiency across NVIDIA enterprise systems. Build and maintain distributed, cloud-native services, automate database operations (provisioning, scaling, backup, failover), and enhance observability with dashboards, alerts, and automation. Participate in incident response to reduce MTTR and drive post-incident improvements. Collaborate with Cloud, Platform, Security, and AI/ML teams while operating Kubernetes-based infrastructure and adopting AI-assisted engineering practices.
Loading
Loading job details...
Preparing the role view and application actions.

