Senior Site Reliability Engineer, BCM - DGX Cloud

NVIDIA
Santa Clara, United States
Workplace: HybridFull timeUSD 168,000 - 333,500 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 8+ yearsEducation: bachelorsSkills: ["Creativity","Autonomy","Incident response","Problem-solving","Operational excellence"]

Build and operate large-scale next-generation GPU clusters powering NVIDIA’s Base Command Manager and customer workloads. You’ll handle incidents bridging cluster operations and development, deploy and run systems in production, and extend Base Command Manager with small feature work. Validate complex cluster configurations using Slurm and Kubernetes to ensure performance, scalability, and resilience across real customer scenarios.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
1 day ago

Senior Site Reliability Engineer, BCM - DGX Cloud

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 30 minutes agoStatus: Live

Job Summary

Build and operate large-scale next-generation GPU clusters powering NVIDIA’s Base Command Manager and customer workloads. You’ll handle incidents bridging cluster operations and development, deploy and run systems in production, and extend Base Command Manager with small feature work. Validate complex cluster configurations using Slurm and Kubernetes to ensure performance, scalability, and resilience across real customer scenarios.
Location: Santa Clara, United States
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Contribute to deployments and daily operations of large scale next-generation GPU platforms.
  • •Handle incidents in GPU clusters, bridging the gap between cluster operations and development.
  • •Design and implement small features in Base Command Manager to learn how the product works.
  • •Validate complex cluster configurations using Slurm and Kubernetes for performance, scalability, and resilience.
  • •Ensure real-world customer scenarios are supported by the cluster configuration and operations.

Pay and Benefits

Salary: USD 168,000 - 333,500 annually
Equity and Bonus:Equity

Key Requirements

  • •Bachelor’s Degree or equivalent experience in Computer Science or related field.
  • •8+ years of experience in site reliability engineering and/or software development roles.
  • •Fluency in Python.
  • •In-depth knowledge of Linux and networking.
  • •Experience with cluster configurations and orchestration (e.g., Slurm and Kubernetes).
Experience:8+ yearsHigh-performance computingCloud infrastructureContainer orchestrationSite reliability engineering
Education:Bachelor's
Skills:CreativityAutonomyIncident responseProblem-solvingOperational excellence
Tech Stack:PythonLinuxNetworkingSlurmKubernetesC++System administrationHigh-performance computingInfiniBandSpectrum-XBase Command ManagerBright Cluster Manager

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor