Distinguished Engineer, Production Engineering, Data Center Automation

NVIDIA
Santa Clara, New York, South Dakota, Wyoming
Full timeUSD 320,000 - 488,750 annuallyFunction: Manufacturing & Production OperationsExperience: 18+ yearsSkills: ["Cross-team coordination","Technical leadership","Architectural direction","High-impact decision-making","Engineering collaboration"]

Lead production engineering for DGX Cloud cluster operations, defining long-range technical strategy and architectural direction for consistent runtime delivery, restoration, and release readiness across on-prem, hyperscalers, and NVIDIA Cloud Partner environments. Build robust operating models and workflows that connect Kubernetes production service management, vendor/equipment readiness, and infrastructure operations to sustain reliability, scalability, and availability at scale. Guide roadmaps and high-impact cross-team technical decisions.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
1 week ago

Distinguished Engineer, Production Engineering, Data Center Automation

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 4 hours agoStatus: Live

Job Summary

Lead production engineering for DGX Cloud cluster operations, defining long-range technical strategy and architectural direction for consistent runtime delivery, restoration, and release readiness across on-prem, hyperscalers, and NVIDIA Cloud Partner environments. Build robust operating models and workflows that connect Kubernetes production service management, vendor/equipment readiness, and infrastructure operations to sustain reliability, scalability, and availability at scale. Guide roadmaps and high-impact cross-team technical decisions.
Location: Santa Clara, New York, South Dakota, Wyoming
Employment Type: Full time
Job Function: Manufacturing & Production Operations
Seniority: Sr. Manager level

Key Responsibilities

  • •Define long-range technical strategy for operating DGX Cloud clusters consistently across on-prem, hyperscalers, and NeoCloud environments.
  • •Set architectural vision and core operational guidelines for cluster lifecycle, runtime delivery, restoration, release readiness, and steady-state operability.
  • •Guide roadmaps and execution for cross-organizational investments improving production readiness, operational safety, performance, and cross-team coordination.
  • •Make and guide high-impact technical decisions on how platform, hardware, provider, and service teams coordinate to operate DGX Cloud resources in production.
  • •Develop workflows, interfaces, and collaboration across Kubernetes production service, provider and hardware readiness, on-prem and bare-metal infrastructure operations, and service-layer reliability.

Pay and Benefits

Salary: USD 320,000 - 488,750 annually
Perks:Equity

Key Requirements

  • •BS, MS, or PhD in Computer Science, Electrical Engineering, or a related technical field, or equivalent experience.
  • •18+ years building and operating large-scale distributed systems, infrastructure platforms, or production environments.
  • •Confirmed company-level technical leadership at principal, distinguished, or equivalent scope in production engineering, SRE, infrastructure software, or cloud platforms.
  • •Confirmed experience establishing operating models, architectural direction, and engineering standards across multiple technical domains and organizations.
  • •Consistent record leading large, cross-team technical efforts from concept through production with measurable outcomes.
Experience:18+ yearsDistributed systemsInfrastructureSRECloud platforms
Education:
Skills:Cross-team coordinationTechnical leadershipArchitectural directionHigh-impact decision-makingEngineering collaboration
Tech Stack:KubernetesDGX CloudNeoCloud

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor