Senior Software Engineer - Cluster Networking

NVIDIA
Santa Clara, Austin, Durham
Workplace: OnsiteFull timeUSD 184,000 - 356,500 annuallyFunction: Software EngineeringExperience: 6+ yearsEducation: mastersSkills: ["Technical judgement","Written communication","Verbal communication","Debugging","Collaboration"]

Own the network architecture for GPU superclusters by designing and operating Kubernetes networking at massive scale. Lead the CNI data plane, overlay mesh, VPN/mesh topologies, and the L7 gateways/load balancers/tunnels that connect control and data planes. Debug hard distributed networking failures across regions and providers, build scale-test validation suites to prevent regressions, and partner with cloud and neocloud teams to bring up new clusters.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
2 hours ago

Senior Software Engineer - Cluster Networking

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 2 hours agoStatus: Live

Job Summary

Own the network architecture for GPU superclusters by designing and operating Kubernetes networking at massive scale. Lead the CNI data plane, overlay mesh, VPN/mesh topologies, and the L7 gateways/load balancers/tunnels that connect control and data planes. Debug hard distributed networking failures across regions and providers, build scale-test validation suites to prevent regressions, and partner with cloud and neocloud teams to bring up new clusters.
Location: Santa Clara, Austin, Durham
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Be the technical owner for how clusters communicate internally and externally, including the CNI data plane, overlay mesh, and inter-cluster connectivity.
  • •Evolve Kubernetes networking architecture for GPU clusters at multi-thousand-node scale.
  • •Design, operate, and scale overlay networks (CNI, mesh, VPN topologies) and the gateways connecting control and data planes.
  • •Design, operate, and scale L7 gateways/load balancers/tunnels (Envoy, Cloudflare).
  • •Eliminate scale ceilings and diagnose complex distributed networking issues by building scale-test environments and validation suites to catch regressions before production.

Pay and Benefits

Salary: USD 184,000 - 356,500 annually
Equity and Bonus:Equity
Perks:Health InsuranceRetirement

Key Requirements

  • •BS/MS in Computer Science, Electrical Engineering, or a related field, or equivalent experience.
  • •6+ years of professional experience in systems, network, or infrastructure software engineering.
  • •Deep knowledge of Kubernetes networking architecture and CNI standards, with production experience operating Calico preferred.
  • •Proficiency designing and maintaining modern mesh and VPN networking topologies (Tailscale, WireGuard, or equivalent).
  • •Strong Linux networking fundamentals (routing, netfilter, iptables/nftables, packet marking, network namespaces) plus the ability to debug distributed network problems at scale.
Experience:6+ yearsAI/MLHigh-performance computingHPCKubernetes networkingCloud networking
Education:Master's
Skills:Technical judgementWritten communicationVerbal communicationDebuggingCollaboration
Tech Stack:KubernetesCNICalicoCiliumTailscaleWireGuardEnvoyCloudflareLinuxIptablesNftablesSlurmGoPythonCInfiniBandRoCERDMAKubernetes networking SIGs

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor