Senior Software Engineer, Cloud-Native Stack – CSP Engagements

NVIDIA
Santa Clara, Austin, Redmond, Seattle
Workplace: OnsiteFull timeFunction: Software EngineeringExperience: 6+ yearsEducation: bachelorsSkills: ["Kubernetes","Slurm","Go","Rust","C/C++","Python","Docker","Terraform","Helm","Ansible","GitHub Actions","Tekton","Prometheus","OpenTelemetry","CUDA","RDMA"]

Senior software engineer on the CSP Engagements team building cloud-native stack for AI/ML datacenters. You will Debug multi-rack, multi-tenant clusters, prototype Kubernetes/Slurm extensions, and drive architecture reviews with CSP and platform teams, delivering reproducible testbeds and customer-focused documentation.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
5 months ago

Senior Software Engineer, Cloud-Native Stack – CSP Engagements

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 4 hours agoStatus: Live

Job Summary

Senior software engineer on the CSP Engagements team building cloud-native stack for AI/ML datacenters. You will Debug multi-rack, multi-tenant clusters, prototype Kubernetes/Slurm extensions, and drive architecture reviews with CSP and platform teams, delivering reproducible testbeds and customer-focused documentation.
Location: Santa Clara, Austin, Redmond, Seattle
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering
Seniority: Sr. Manager level

Key Responsibilities

  • •Perform deep-dive debugging of multi-rack, multi-tenant clusters: scheduler behavior, container runtime issues, device-plugin crashes, RDMA/IB fabric anomalies.
  • •Gather customer requirements and prototype feature extensions for Kubernetes operators, Slurm plugins, and custom micro-services exposing new GPU capabilities.
  • •Lead architecture reviews and whiteboard sessions with CSP and platform teams; convert findings into RFCs and upstream PRs.
  • •Create reproducible testbeds (Helm/Ansible/Terraform) mirroring customer environments; automate validation and benchmark suites.
  • •Deliver technical collateral, how-to guides, demo scripts; present at customer on-sites, KubeCon, and SlurmUG.

Pay and Benefits

Perks:EquityHealth InsuranceBenefits

Key Requirements

  • •6+ years of professional software development experience in distributed systems (Go, Rust, C/C++ or Python for tooling)
  • •Strong Kubernetes internals (scheduler, CRI/CNI/CSI, operators) and Slurm knowledge
  • •Experience integrating GPUs into containerized clusters (Blackwell/GB200/GB300)
  • •Familiarity with CI/CD (GitHub Actions, Tekton) and observability (Prometheus, OpenTelemetry)
  • •Customer-facing engineering or solutions-architect background: requirements gathering, PoC ownership
Experience:6+ yearsCloud computingAI infrastructureDistributed systems
Education:Bachelor's
Skills:KubernetesSlurmGoRustC/C++PythonDockerTerraformHelmAnsibleGitHub ActionsTektonPrometheusOpenTelemetryCUDARDMA
Languages:English
Tech Stack:KubernetesSlurmGoRustC/C++PythonDockerTerraformHelmAnsibleGitHub ActionsTektonPrometheusOpenTelemetryRDMACUDA

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor