Senior Site Reliability Engineer, DGX Cloud

NVIDIA
Santa Clara, Texas, Wyoming
Workplace: OnsiteFull timeUSD 168,000 - 270,250 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 8+ yearsEducation: bachelorsSkills: ["Incident management","Triage","Root-cause analysis","Blameless postmortems","Automation mindset"]

Maintain and improve DGX Cloud’s large-scale Kubernetes infrastructure that powers AI workloads across major cloud providers. Own SLO/SLI design, observability, alerting, and reliability reporting; support services through launch readiness and ongoing health monitoring. Optimize GPU workloads on AWS/GCP/Azure/OCI/private clouds, lead incident triage and root-cause analysis, and participate in on-call rotations. Build automation and tooling to improve reliability, velocity, and operational excellence.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
1 day ago

Senior Site Reliability Engineer, DGX Cloud

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 2 hours agoStatus: Live

Job Summary

Maintain and improve DGX Cloud’s large-scale Kubernetes infrastructure that powers AI workloads across major cloud providers. Own SLO/SLI design, observability, alerting, and reliability reporting; support services through launch readiness and ongoing health monitoring. Optimize GPU workloads on AWS/GCP/Azure/OCI/private clouds, lead incident triage and root-cause analysis, and participate in on-call rotations. Build automation and tooling to improve reliability, velocity, and operational excellence.
Location: Santa Clara, Texas, Wyoming
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Build, implement, and support operational reliability for large-scale Kubernetes clusters focused on performance and real-time monitoring, logging, and alerting.
  • •Define SLOs/SLIs, monitor error allowances, and streamline reliability reporting.
  • •Support services before launch via system creation consulting, tooling/platform/framework development, capacity management, and launch reviews.
  • •Maintain live services by measuring and supervising availability, latency, and overall system health.
  • •Lead triage and root-cause analysis for high-severity incidents, participate in on-call rotation, and conduct balanced incident response and blameless postmortems.

Pay and Benefits

Salary: USD 168,000 - 270,250 annually
Equity and Bonus:Equity

Key Requirements

  • •BS in Computer Science or related technical field, or equivalent experience.
  • •8+ years of experience operating production services.
  • •Expert-level knowledge of Kubernetes administration, containerization, and microservices architecture.
  • •Experience with infrastructure automation tools such as Terraform, Ansible, Chef, or Puppet.
  • •Proficiency in at least one programming language (e.g., Python, Go) and strong Linux/networking/cloud security fundamentals.
Experience:8+ yearsAICloud computingKubernetesSRE
Education:Bachelor's in Computer Science
Skills:Incident managementTriageRoot-cause analysisBlameless postmortemsAutomation mindset
Tech Stack:KubernetesLinuxTCP/IPAWSGCPAzureOCITerraformAnsibleChefPuppetPythonGoOpenTelemetryPrometheusGrafanaELK StackLightstepSplunkKubeVirt

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor