Senior Software Engineer, DGX Cloud Orchestration

NVIDIA
United States
Workplace: RemoteFull timeUSD 184,000 - 356,500 annuallyFunction: Software EngineeringExperience: 8+ yearsEducation: bachelorsSkills: ["Communication","Collaboration","Debugging","Problem-solving"]

Design and develop APIs that orchestrate operational workflows for DGX Cloud. Build workflow automation and state management to streamline infrastructure lifecycle processes, and collaborate across teams to codify business processes into scalable systems. Create schema-driven platforms to reduce manual toil while ensuring operational consistency. Integrate with Kubernetes and observability tools (Prometheus, OpenTelemetry, Grafana) and improve reliability and efficiency using telemetry and automated workflows.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
1 week ago

Senior Software Engineer, DGX Cloud Orchestration

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 36 minutes agoStatus: Live

Job Summary

Design and develop APIs that orchestrate operational workflows for DGX Cloud. Build workflow automation and state management to streamline infrastructure lifecycle processes, and collaborate across teams to codify business processes into scalable systems. Create schema-driven platforms to reduce manual toil while ensuring operational consistency. Integrate with Kubernetes and observability tools (Prometheus, OpenTelemetry, Grafana) and improve reliability and efficiency using telemetry and automated workflows.
Location: United States
Workplace: Remote
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Design and develop APIs to orchestrate and integrate operational workflows.
  • •Build state management and workflow automation systems for infrastructure lifecycle processes.
  • •Collaborate across teams to codify business processes into scalable, self-measuring systems.
  • •Develop extensible, schema-driven platforms that reduce manual toil and ensure operational consistency.
  • •Integrate with Kubernetes and observability systems (Prometheus, OpenTelemetry, Grafana) and optimize reliability and efficiency using telemetry.

Pay and Benefits

Salary: USD 184,000 - 356,500 annually
Equity and Bonus:Equity
Perks:Equity

Key Requirements

  • •8+ years of industry experience with a Bachelor’s degree (or equivalent experience); Master’s degree preferred.
  • •Experience designing, building, and operating services in a high-reliability environment.
  • •Proficiency in Go, Java, or Python.
  • •Strong understanding of cloud infrastructure (AWS, GCP, Azure) and container technologies like Docker and Kubernetes.
  • •Experience with high-scale distributed systems and API/data pipeline architectural patterns.
Experience:8+ yearsHigh reliabilityCloud operationsDistributed systemsAI/ML software stack
Education:Bachelor's
Skills:CommunicationCollaborationDebuggingProblem-solving
Tech Stack:APIsGoJavaPythonAWSGCPAzureDockerKubernetesPrometheusOpenTelemetryGrafanaDistributed systemsData pipelinesCUDACuDNN

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor