Senior Site Reliability Engineer, BCM - DGX Cloud
Santa Clara, United States
Workplace: HybridFull timeUSD 168,000 - 333,500 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 8+ yearsEducation: bachelorsSkills: ["Creativity","Autonomy","Incident response","Problem-solving","Operational excellence"]Build and operate large-scale next-generation GPU clusters powering NVIDIA’s Base Command Manager and customer workloads. You’ll handle incidents bridging cluster operations and development, deploy and run systems in production, and extend Base Command Manager with small feature work. Validate complex cluster configurations using Slurm and Kubernetes to ensure performance, scalability, and resilience across real customer scenarios.
Loading
Loading job details...
Preparing the role view and application actions.

