Senior Staff Site Reliability Engineer

NVIDIA
Bengaluru
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 10+ yearsEducation: bachelorsSkills: ["Problem-solving","Communication","Collaboration","Mentoring","Ownership","Curiosity","Innovation"]

Own the technical strategy and roadmap for large-scale SRE initiatives to improve reliability, scalability, and developer efficiency across NVIDIA enterprise systems. Design and modernize resilient distributed architectures, drive automation and observability (including AI workload signals), and build LLM-aware monitoring with autonomous incident response. Partner with Cloud, Platform, Security, and AI/ML teams to deliver high-availability, secure operations, and mentor engineers on AI-assisted engineering practices.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
36 minutes ago

Senior Staff Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 36 minutes agoStatus: Live

Job Summary

Own the technical strategy and roadmap for large-scale SRE initiatives to improve reliability, scalability, and developer efficiency across NVIDIA enterprise systems. Design and modernize resilient distributed architectures, drive automation and observability (including AI workload signals), and build LLM-aware monitoring with autonomous incident response. Partner with Cloud, Platform, Security, and AI/ML teams to deliver high-availability, secure operations, and mentor engineers on AI-assisted engineering practices.
Location: Bengaluru
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Sr. Manager level

Key Responsibilities

  • •Lead the technical strategy and roadmap for large-scale SRE initiatives to boost reliability and scalability.
  • •Design and build resilient distributed systems, modernizing legacy applications and database systems into scalable architectures.
  • •Drive automation and observability improvements using metrics and analytics, including AI workload quality and model performance telemetry.
  • •Build LLM-aware monitoring and autonomous incident response pipelines to reduce toil and accelerate MTTR.
  • •Collaborate with Cloud, Platform, Security, and AI/ML teams to deliver AI-native SRE elements and mentor engineers on AI-assisted workflows.

Key Requirements

  • •10+ years of experience in Site Reliability Engineering, Platform Engineering, or Cloud Architect roles.
  • •BS degree in Computer Science or related technical field involving coding, or equivalent experience.
  • •Strong programming skills (Python, Typescript/JavaScript, or Go) focused on automation, infrastructure-as-code, and distributed systems debugging.
  • •Experience with infrastructure-as-code tooling such as AWS CDK, AWS CloudFormation, Terraform, or CrossPlane.
  • •Deep expertise in systems architecture, networking, Kubernetes, and public cloud services (AWS, Azure, or GCP).
Experience:10+ yearsPlatform engineeringCloudDistributed systemsKubernetesAI/ML infrastructureInfrastructure-as-code
Education:Bachelor's in Computer Science (or related technical field)
Skills:Problem-solvingCommunicationCollaborationMentoringOwnershipCuriosityInnovation
Tech Stack:PythonTypescriptJavaScriptGoInfrastructure-as-codeAWS CDKAWS CloudFormationTerraformCrossPlaneOpenTelemetryKubernetesAWSAzureGCPLLMLangGraphAutoGenAgent orchestration frameworksRunbook automationMTTR

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor