Senior Site Reliability Engineer, DGX Cloud

NVIDIA
Zurich
Workplace: RemoteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 10+ yearsEducation: bachelorsSkills: ["Root-cause analysis","Blameless postmortems","Incident response","Capacity management"]

Build and run large-scale Kubernetes clusters for DGX Cloud, with a focus on performance and production reliability. Define SLOs/SLIs, manage error budgets, and improve observability across monitoring, logging, and tracing. Partner with launch teams to advise on system creation, then operate live services by tracking availability, latency, and system health. Lead triage and root-cause analysis during high-severity incidents and participate in on-call to keep GPU workloads running on major clouds.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
1 day ago

Senior Site Reliability Engineer, DGX Cloud

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 3 hours agoStatus: Live

Job Summary

Build and run large-scale Kubernetes clusters for DGX Cloud, with a focus on performance and production reliability. Define SLOs/SLIs, manage error budgets, and improve observability across monitoring, logging, and tracing. Partner with launch teams to advise on system creation, then operate live services by tracking availability, latency, and system health. Lead triage and root-cause analysis during high-severity incidents and participate in on-call to keep GPU workloads running on major clouds.
Location: Zurich
Workplace: Remote
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Build, implement, and support operational reliability for large-scale Kubernetes clusters with performance at scale, including real-time monitoring, logging, and alerting.
  • •Define SLOs/SLIs, monitor error budgets, and streamline reporting for reliability outcomes.
  • •Support services before launch through system creation consulting, developing software tools/platforms/frameworks, capacity management, and launch reviews.
  • •Maintain live services by measuring and monitoring availability, latency, and overall system health.
  • •Lead triage and root-cause analysis for high-severity incidents and participate in on-call rotation to support production services.

Key Requirements

  • •BS in Computer Science or related technical field, or equivalent experience.
  • •10+ years of experience operating production services.
  • •Expert-level knowledge of Kubernetes administration, containerization, and microservices architecture.
  • •Experience with infrastructure automation tools (e.g., Terraform, Ansible, Chef, Puppet).
  • •Proficiency in at least one high-level programming language (e.g., Python, Go).
Experience:10+ yearsCloud infrastructureKubernetesSite reliability engineeringObservability
Education:Bachelor's
Skills:Root-cause analysisBlameless postmortemsIncident responseCapacity management
Tech Stack:KubernetesTerraformAnsibleChefPuppetPythonGoLinuxAWSGCPAzureOCIOpenTelemetryPrometheusGrafanaELK StackLightstepSplunkKubeVirtTemporal

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor