Senior Site Reliability Engineer, DGX Cloud
Santa Clara, Texas, Wyoming
Workplace: OnsiteFull timeUSD 168,000 - 270,250 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 8+ yearsEducation: bachelorsSkills: ["Incident management","Triage","Root-cause analysis","Blameless postmortems","Automation mindset"]Maintain and improve DGX Cloud’s large-scale Kubernetes infrastructure that powers AI workloads across major cloud providers. Own SLO/SLI design, observability, alerting, and reliability reporting; support services through launch readiness and ongoing health monitoring. Optimize GPU workloads on AWS/GCP/Azure/OCI/private clouds, lead incident triage and root-cause analysis, and participate in on-call rotations. Build automation and tooling to improve reliability, velocity, and operational excellence.
Loading
Loading job details...
Preparing the role view and application actions.

