Principal Software Engineer - Compute Infrastructure

NVIDIA
Santa Clara, United States
Workplace: HybridFull timeUSD 248,000 - 391,000 annuallyFunction: Software EngineeringExperience: 15+ yearsEducation: bachelorsSkills: ["Leadership","Strategic influence","Automation focus","Technical direction leadership","Stakeholder collaboration"]

Lead the architectural vision and operationalization of NVIDIA’s enterprise compute platform and internal AI inference infrastructure. You’ll define service tiers and automated cluster lifecycles across OpenShift/KubeVirt, build remediation and telemetry for pre-release rack-scale GPU systems, and drive capacity planning and migrations into Kubernetes orchestration. Work with highly autonomous engineering teams to deliver self-service architectures, APIs, and infrastructure-as-code providers teams adopt at massive scale.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
4 days ago

Principal Software Engineer - Compute Infrastructure

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 1 hour agoStatus: Live

Job Summary

Lead the architectural vision and operationalization of NVIDIA’s enterprise compute platform and internal AI inference infrastructure. You’ll define service tiers and automated cluster lifecycles across OpenShift/KubeVirt, build remediation and telemetry for pre-release rack-scale GPU systems, and drive capacity planning and migrations into Kubernetes orchestration. Work with highly autonomous engineering teams to deliver self-service architectures, APIs, and infrastructure-as-code providers teams adopt at massive scale.
Location: Santa Clara, United States
Workplace: Hybrid
Employment Type: Full time
Job Function: Software Engineering
Seniority: Sr. Manager level

Key Responsibilities

  • •Define and transform the global enterprise compute platform architecture, including service tiers, SLAs, and automated cluster lifecycles across thousands of nodes, VMs, and containers.
  • •Build the operational foundation for internal AI inference infrastructure, including automated remediation pipelines, hardware watchdogs, and telemetry for rack-scale GPU systems.
  • •Drive capacity planning and scaling strategies under hardware supply constraints, including public cloud bursting and evaluating alternative compute architectures such as ARM.
  • •Develop self-service platform architectures, APIs, and Terraform/OpenTofu providers to enable adoption by autonomous engineering teams.
  • •Evaluate and execute migrations of large legacy workloads into modern Kubernetes orchestration, including large-scale long-running VDI environments.

Pay and Benefits

Salary: USD 248,000 - 391,000 annually
Equity and Bonus:Equity

Key Requirements

  • •Bachelor’s degree in Engineering, Computer Science, Mathematics, or related field, or equivalent experience.
  • •15+ years of experience in compute platform engineering, site reliability, or systems architecture with heavy automation at massive scale.
  • •Deep expertise in Kubernetes architecture and virtualization architectures running VMs inside K8s (KubeVirt, OpenShift).
  • •In-depth knowledge of hardware technologies (GPUs, high-speed backplane networking) including mitigation of hardware-level failures and anomalies.
  • •Proficiency in Go and/or Python and expert-level infrastructure-as-code development (Terraform, Config Management) with GitOps posture (ArgoCD or similar).
Experience:15+ yearsCompute platform engineeringSite reliabilitySystems architectureAI infrastructureKubernetesCloudMicroservices
Education:Bachelor's
Skills:LeadershipStrategic influenceAutomation focusTechnical direction leadershipStakeholder collaboration
Tech Stack:KubernetesOpenShiftKubeVirtBlackwellGoPythonTerraformConfig ManagementOpenTofuArgoCDGitOpsAWSGCPARMNFSv4NVMe/TCPHyperconverged storageTerraform/OpenTofuMicroservicesVDI

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor