Principal Site Reliability Engineer

NVIDIA
Santa Clara
Workplace: HybridFull timeUSD 248,000 - 396,750 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 15+ yearsEducation: bachelorsSkills: ["Technical judgment","Communication","Collaboration","Mentoring","Influencing stakeholders"]

Shape the technical vision for reliability across NVIDIA’s AI Platform Runtime, leading architecture and roadmap efforts across multiple organizations. Design and deliver highly available, resilient, secure distributed platforms, and build AI-driven automation to accelerate incident response, troubleshooting, and remediation. Set enterprise reliability standards (SLOs, error budgets, capacity models) and advance observability using OpenTelemetry and anomaly detection. Provide technical leadership during critical incidents and mentor engineering leaders to raise the technical bar.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
3 days ago

Principal Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 3 hours agoStatus: Live

Job Summary

Shape the technical vision for reliability across NVIDIA’s AI Platform Runtime, leading architecture and roadmap efforts across multiple organizations. Design and deliver highly available, resilient, secure distributed platforms, and build AI-driven automation to accelerate incident response, troubleshooting, and remediation. Set enterprise reliability standards (SLOs, error budgets, capacity models) and advance observability using OpenTelemetry and anomaly detection. Provide technical leadership during critical incidents and mentor engineering leaders to raise the technical bar.
Location: Santa Clara
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure

Key Responsibilities

  • •Define and drive the long-term technical vision, architecture, and roadmap for reliability across NVIDIA’s AI Platform Runtime and related enterprise systems.
  • •Architect highly available, resilient, secure, and scalable distributed platforms for next-generation AI-driven products and services.
  • •Lead the design and development of AI agents, AI skills, and intelligent automation to accelerate platform operations and incident response.
  • •Establish platform-wide reliability standards including service-level objectives, error budgets, capacity models, resilience patterns, and operational readiness requirements.
  • •Advance observability across distributed environments using OpenTelemetry (metrics, logs, traces, profiling, analytics) and automated anomaly detection.

Pay and Benefits

Salary: USD 248,000 - 396,750 annually
Equity and Bonus:Equity

Key Requirements

  • •15+ years of experience in Site Reliability Engineering, Platform Engineering, Distributed Systems, Cloud Architecture, or related infrastructure roles.
  • •BS or MS in Computer Science (or related) or equivalent experience involving significant software development.
  • •Demonstrated experience setting technical strategy and leading large-scale engineering initiatives across multiple teams or organizations.
  • •Deep expertise in distributed systems architecture, networking, Linux, Kubernetes, and public cloud platforms such as AWS, Azure, or GCP.
  • •Strong experience with reliability engineering practices (SLOs, error budgets, capacity planning, fault tolerance, disaster recovery, incident management, blameless postmortems).
Experience:15+ yearsDistributed systemsCloudInfrastructure engineeringSREPlatform engineering
Education:Bachelor's in Computer Science
Skills:Technical judgmentCommunicationCollaborationMentoringInfluencing stakeholders
Tech Stack:KubernetesLinuxAWSAzureGCPPythonGoTypeScriptJavaScriptJavaTerraformCrossplaneAWS CDKAWS CloudFormationOpenTelemetry

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor