Senior Software Engineer, DGX Cloud Production Engineering

NVIDIA
Santa Clara
Workplace: RemoteFull timeUSD 184,000 - 356,500 annuallyFunction: Software EngineeringExperience: 8+ yearsEducation: bachelorsSkills: ["Communication","Cross-team collaboration"]

Build and operate automation for large-scale GPU clusters powering DGX Cloud across NVIDIA Cloud Partners (NCP) and on-prem environments. Develop provisioning, validation, upgrade, monitoring, repair, and cluster lifecycle tools, and improve Day 0/1/2 workflows for bringup and production operations. Reduce manual production touches via APIs, GitOps, and agent-assisted automation, and participate in on-call incident response and debugging while partnering across platform and security teams.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
2 days ago

Senior Software Engineer, DGX Cloud Production Engineering

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 16 hours agoStatus: Live
Reposted: similar role first listed 4 months ago

Job Summary

Build and operate automation for large-scale GPU clusters powering DGX Cloud across NVIDIA Cloud Partners (NCP) and on-prem environments. Develop provisioning, validation, upgrade, monitoring, repair, and cluster lifecycle tools, and improve Day 0/1/2 workflows for bringup and production operations. Reduce manual production touches via APIs, GitOps, and agent-assisted automation, and participate in on-call incident response and debugging while partnering across platform and security teams.
Location: Santa Clara
Workplace: Remote
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Build and operate automation for large-scale GPU clusters across NVIDIA Cloud Partners (NCP) and on-prem environments.
  • •Develop tools and services for provisioning, validation, upgrades, monitoring, repair, and cluster lifecycle operations.
  • •Improve Day 0/Day 1/Day 2 workflows for cluster bringup, handoff, and production operations.
  • •Reduce manual production touches using APIs, GitOps, automation, and agent-assisted workflows.
  • •Participate in on-call, incident response, debugging, and durable follow-up work while partnering with platform, storage, networking, and security teams.

Pay and Benefits

Salary: USD 184,000 - 356,500 annually
Equity and Bonus:Equity

Key Requirements

  • •8+ years of experience building or operating production infrastructure.
  • •Strong programming skills in Python, Go, or similar.
  • •Experience with Linux, Kubernetes, containers, cloud infrastructure, or infrastructure automation.
  • •Ability to troubleshoot distributed systems in production.
  • •BS/MS in Computer Science or equivalent experience.
Experience:8+ yearsGPU infrastructureKubernetesCloud infrastructureProduction engineeringDistributed systems
Education:Bachelor's
Skills:CommunicationCross-team collaboration
Tech Stack:PythonGoLinuxKubernetesContainersCloud infrastructureInfrastructure automationGitOpsAPIsTerraformArgoCD

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor