Principal Software Engineer, DGX Cloud Production Engineering

NVIDIA
Santa Clara, United States
Workplace: RemoteFull timeUSD 272,000 - 431,250 annuallyFunction: Software EngineeringExperience: 15+ yearsEducation: mastersSkills: ["Communication","Collaboration","Technical leadership","Cross-functional execution","Engineering rigor"]

Build the next-generation Kubernetes platform for self-service GPU infrastructure. Own the architecture and development of core platform capabilities such as cluster management, control plane services, fleet lifecycle, and day-2 operations. Design reliable distributed systems and APIs for provisioning, upgrading, and remediating clusters at scale across cloud and on-prem. Lead diagnosis of complex infrastructure/runtime/networking issues and mentor senior engineers to raise engineering rigor.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
3 hours ago

Principal Software Engineer, DGX Cloud Production Engineering

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 3 hours agoStatus: Live

Job Summary

Build the next-generation Kubernetes platform for self-service GPU infrastructure. Own the architecture and development of core platform capabilities such as cluster management, control plane services, fleet lifecycle, and day-2 operations. Design reliable distributed systems and APIs for provisioning, upgrading, and remediating clusters at scale across cloud and on-prem. Lead diagnosis of complex infrastructure/runtime/networking issues and mentor senior engineers to raise engineering rigor.
Location: Santa Clara, United States
Workplace: Remote
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Lead architecture and development of core Kubernetes platform capabilities including cluster management, control plane services, fleet lifecycle, and day-2 operations.
  • •Design and build reliable distributed systems and APIs for provisioning, managing, upgrading, and remediating Kubernetes clusters at scale.
  • •Define technical requirements, validation criteria, and production-readiness practices, including direction for declarative workflows and automation.
  • •Collaborate across engineering teams to deliver cohesive platform experiences across management APIs, lifecycle orchestration, runtime integration, and fleet consistency.
  • •Diagnose and resolve complex platform issues spanning infrastructure, runtime, networking, hardware, and operations to improve scalability, resilience, and operability.

Pay and Benefits

Salary: USD 272,000 - 431,250 annually
Equity and Bonus:Equity
Perks:Equity

Key Requirements

  • •BS or MS in Computer Science, Computer Engineering, or related field, or equivalent experience.
  • •15+ years of relevant software engineering experience building and operating large-scale production systems.
  • •Deep expertise in Kubernetes internals, including APIs, controllers or operators, and cluster lifecycle management.
  • •Strong distributed systems background focused on reliability, scalability, and failure recovery.
  • •Strong programming skills in one or more systems/cloud-native languages such as Go, Python, Rust, or C++.
Experience:15+ yearsKubernetesPlatform softwareDistributed systemsCloud-nativeManaged servicesAI infrastructureGPUHPC
Education:Master's in Computer Science; Computer Engineering
Skills:CommunicationCollaborationTechnical leadershipCross-functional executionEngineering rigor
Tech Stack:KubernetesGoPythonRustC++APIsDistributed systemsGitOpsControllersOperatorsDeclarative infrastructure

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor