Senior Software Engineer, DGX Cloud Production Engineering

NVIDIA
Santa Clara, United States
Workplace: RemoteFull timeUSD 152,000 - 287,500 annuallyFunction: Software EngineeringExperience: 5+ yearsEducation: bachelorsSkills: ["Technical leadership","Cross-functional collaboration","Communication","Problem solving","Driving initiatives to completion"]

Build production-grade Kubernetes platform capabilities that enable self-service GPU infrastructure, managed control planes, and reliable day-2 operations. Partner across teams to design and implement distributed systems, APIs, and declarative workflows that provision, upgrade, remediate, and operate clusters at scale across cloud and on-prem environments. Drive technical strategy and troubleshooting for complex platform issues spanning infrastructure, runtime, networking, hardware, and operations.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
3 days ago

Senior Software Engineer, DGX Cloud Production Engineering

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 3 hours agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Build production-grade Kubernetes platform capabilities that enable self-service GPU infrastructure, managed control planes, and reliable day-2 operations. Partner across teams to design and implement distributed systems, APIs, and declarative workflows that provision, upgrade, remediate, and operate clusters at scale across cloud and on-prem environments. Drive technical strategy and troubleshooting for complex platform issues spanning infrastructure, runtime, networking, hardware, and operations.
Location: Santa Clara, United States
Workplace: Remote
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Collaborate on architecture and development of core Kubernetes platform capabilities including cluster management, control plane services, fleet lifecycle, and day-2 operations.
  • •Design and build reliable distributed systems and APIs for provisioning, managing, upgrading, and remediating Kubernetes clusters at scale.
  • •Define technical requirements, validation criteria, production-readiness practices, and direction for declarative workflows and automation across the Kubernetes stack.
  • •Collaborate across engineering teams to deliver cohesive platform experiences spanning management APIs, lifecycle orchestration, runtime integration, and fleet consistency.
  • •Diagnose and resolve complex platform issues across infrastructure, runtime, networking, hardware, and operations to improve scalability, resilience, and operability.

Pay and Benefits

Salary: USD 152,000 - 287,500 annually
Equity and Bonus:Equity

Key Requirements

  • •BS or MS in Computer Science, Computer Engineering, or related field, or equivalent experience.
  • •5+ years of relevant software engineering experience building and operating large-scale production systems.
  • •Deep expertise in Kubernetes internals, APIs, controllers or operators, and cluster lifecycle management.
  • •Strong background in distributed systems design, reliability, scalability, and failure recovery.
  • •Strong programming skills in systems/cloud-native languages such as Go, Python, Rust, or C++.
Experience:5+ yearsKubernetesDistributed systemsCloud-nativeManaged servicesAI infrastructure
Education:Bachelor's in Computer Science, Computer Engineering, or related field
Skills:Technical leadershipCross-functional collaborationCommunicationProblem solvingDriving initiatives to completion
Tech Stack:KubernetesGoPythonRustC++APIsControllersOperatorsDistributed systemsGitOpsPolicy-driven automationDeclarative workflows

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor