Senior Site Reliability Engineer, DGX Cloud
Zurich
Workplace: RemoteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 10+ yearsEducation: bachelorsSkills: ["Root-cause analysis","Blameless postmortems","Incident response","Capacity management"]Build and run large-scale Kubernetes clusters for DGX Cloud, with a focus on performance and production reliability. Define SLOs/SLIs, manage error budgets, and improve observability across monitoring, logging, and tracing. Partner with launch teams to advise on system creation, then operate live services by tracking availability, latency, and system health. Lead triage and root-cause analysis during high-severity incidents and participate in on-call to keep GPU workloads running on major clouds.
Loading
Loading job details...
Preparing the role view and application actions.

