Distinguished Engineer, Production Engineering, Cluster Management

NVIDIA
Santa Clara, Oregon
Workplace: HybridFull timeUSD 320,000 - 488,750 annuallyFunction: Manufacturing & Production OperationsExperience: 18+ yearsSkills: ["Architectural judgment","Technical leadership","Cross-team influence","Operational reliability thinking","Automation mindset"]

Lead production engineering for DGX Cloud GPU capacity as a hands-on technical authority. Define long-range strategy and architectural direction for cluster lifecycle, runtime delivery, restoration, release readiness, and steady-state operability across on-prem, hyperscalers, and NVIDIA Cloud Partner environments. Build durable workflows, APIs, and automation for Kubernetes-based service management and readiness gates, and drive cross-team improvements in reliability, operability, performance, and release safety.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
3 days ago

Distinguished Engineer, Production Engineering, Cluster Management

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 15 hours agoStatus: Live

Job Summary

Lead production engineering for DGX Cloud GPU capacity as a hands-on technical authority. Define long-range strategy and architectural direction for cluster lifecycle, runtime delivery, restoration, release readiness, and steady-state operability across on-prem, hyperscalers, and NVIDIA Cloud Partner environments. Build durable workflows, APIs, and automation for Kubernetes-based service management and readiness gates, and drive cross-team improvements in reliability, operability, performance, and release safety.
Location: Santa Clara, Oregon
Workplace: Hybrid
Employment Type: Full time
Job Function: Manufacturing & Production Operations

Key Responsibilities

  • •Define long-range technical strategy for operating DGX Cloud clusters across local data centers, hyperscalers, and NeoCloud environments.
  • •Set architectural direction and operating standards for cluster lifecycle, runtime delivery, restoration, release readiness, and steady-state operability.
  • •Guide cross-organizational roadmap and execution for investments improving production readiness, operational safety, performance, and coordination.
  • •Make and influence high-impact decisions on how platform, hardware, provider, and service teams work together in production.
  • •Build and evolve automation, APIs, operating workflows, and readiness gates to move new capacity into stable production and keep capacity balanced.

Pay and Benefits

Salary: USD 320,000 - 488,750 annually
Equity and Bonus:Equity

Key Requirements

  • •BS, MS, or PhD in Computer Science, Electrical Engineering, or a related technical field, or equivalent experience.
  • •18+ years of experience building and operating large-scale distributed systems, infrastructure platforms, or production environments.
  • •Confirmed company-level technical leadership at principal, distinguished, or equivalent scope in production engineering, SRE, infrastructure software, or cloud platforms.
  • •Deep experience with Kubernetes-based production systems, infrastructure automation, or distributed systems operations.
  • •Strong software engineering skills in languages such as Python, Go, or similar low-level programming languages.
Experience:18+ yearsDistributed systemsInfrastructureSRECloud platforms
Education:
Skills:Architectural judgmentTechnical leadershipCross-team influenceOperational reliability thinkingAutomation mindset
Tech Stack:DGX CloudGPU capacityKubernetesPythonGoLinuxNetworkingContainersDistributed systemsAPIs

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor