Senior Software Engineer, Capacity Management - DGX Cloud

NVIDIA
Santa Clara
Full timeUSD 140,000 - 270,250 annuallyFunction: Software EngineeringExperience: 5+ yearsEducation: bachelorsSkills: ["Communication","Cross-functional collaboration","Technical leadership","Debugging/diagnostics","Reliability focus"]

Design and build distributed systems and data pipelines for DGX Cloud capacity planning, allocation, reservations, and utilization. Develop a unified model of GPU capacity across cloud providers, regions, clusters, and products, and automate workflows that are currently manual. Create APIs, tools, and integrations for capacity-aware decisions, while improving forecasting, monitoring, data-quality controls, reliability, and scalability. Lead technical design reviews and mentor engineers.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
1 day ago

Senior Software Engineer, Capacity Management - DGX Cloud

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 2 hours agoStatus: Live

Job Summary

Design and build distributed systems and data pipelines for DGX Cloud capacity planning, allocation, reservations, and utilization. Develop a unified model of GPU capacity across cloud providers, regions, clusters, and products, and automate workflows that are currently manual. Create APIs, tools, and integrations for capacity-aware decisions, while improving forecasting, monitoring, data-quality controls, reliability, and scalability. Lead technical design reviews and mentor engineers.
Location: Santa Clara
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Design and build distributed services and data pipelines for capacity planning, allocation, reservations, and utilization.
  • •Develop a unified model of available, committed, and forecasted GPU capacity across providers, regions, clusters, and products.
  • •Automate capacity-management workflows currently dependent on manual analysis and coordination.
  • •Build APIs, tools, and integrations that enable other DGX Cloud systems and teams to make capacity-aware decisions.
  • •Improve forecasting, scenario planning, and operational visibility with demand signals and infrastructure supply data, including monitoring and data-quality controls.

Pay and Benefits

Salary: USD 140,000 - 270,250 annually

Key Requirements

  • •BS or equivalent experience in Computer Science, Computer Engineering, or a related technical field.
  • •5+ years of software engineering experience building production systems.
  • •Strong programming experience in Python, Go, Java, or similar languages.
  • •Experience designing distributed systems, backend services, APIs, and data-processing pipelines.
  • •Experience working with cloud infrastructure and Kubernetes, compute platforms, or large-scale resource-management systems.
Experience:5+ yearsAI/ML platformsCloud infrastructureHigh-performance computingDistributed systems
Education:Bachelor's
Skills:CommunicationCross-functional collaborationTechnical leadershipDebugging/diagnosticsReliability focus
Tech Stack:PythonGoJavaAPIsDistributed systemsData-processing pipelinesCloud infrastructureKubernetesGPU infrastructureCapacity planningData modelingObservability

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor