Senior Platform Engineer, Network Infrastructure - DGX Cloud

NVIDIA
United States
Workplace: RemoteFull timeUSD 168,000 - 333,500 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 8+ yearsEducation: bachelorsSkills: ["Incident response","Root-cause analysis","Driving corrective actions","Problem-solving","Accountability"]

Own the lifecycle and automation of a Kubernetes platform that powers network automation, telemetry, and operations across data center, colocation, and cloud environments. Build production-grade software and GitOps delivery for cluster provisioning, validation, upgrades, remediation, and multi-cluster operations. Provide production support and drive complex incident response, while establishing observability and production-readiness standards across teams.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
1 month ago

Senior Platform Engineer, Network Infrastructure - DGX Cloud

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 4 hours agoStatus: Live

Job Summary

Own the lifecycle and automation of a Kubernetes platform that powers network automation, telemetry, and operations across data center, colocation, and cloud environments. Build production-grade software and GitOps delivery for cluster provisioning, validation, upgrades, remediation, and multi-cluster operations. Provide production support and drive complex incident response, while establishing observability and production-readiness standards across teams.
Location: United States
Workplace: Remote
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Design, build, and operate the Kubernetes platform powering GNI network automation, telemetry, and operations.
  • •Own lifecycle management for Kubernetes environments, including onboarding, upgrades, capacity, availability, and recovery.
  • •Develop production-quality software and automation for cluster provisioning, validation, upgrades, remediation, and GitOps multi-cluster delivery.
  • •Provide production support for network services hosted on the platform, diagnosing cross-platform issues with engineering owners.
  • •Lead incident response during on-call, establish production-readiness and observability standards, and drive corrective actions to completion.

Pay and Benefits

Salary: USD 168,000 - 333,500 annually
Equity and Bonus:Equity

Key Requirements

  • •8+ years building or operating production Kubernetes platforms, network infrastructure, or distributed systems.
  • •Deep Kubernetes expertise at scale, including cluster lifecycle, upgrades, networking, storage, and recovery.
  • •Proficiency in at least one general-purpose programming language such as Go or Python.
  • •Experience with GitOps, infrastructure as code, CI/CD, and automated production delivery.
  • •Experience with production on-call, incident response, root-cause analysis, and driving corrective actions to completion.
Experience:8+ yearsDistributed systemsKubernetesNetwork infrastructure
Education:Bachelor's in Computer Science, Engineering, or a related field
Skills:Incident responseRoot-cause analysisDriving corrective actionsProblem-solvingAccountability
Tech Stack:KubernetesGoPythonGitOpsInfrastructure as codeCI/CDObservabilityCluster APIMetal3

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor