Senior HPC and AI Network Software Architect

NVIDIA
Zurich
Workplace: OnsiteFull timeFunction: Software EngineeringExperience: 5+ yearsEducation: phdSkills: ["Collaboration","Innovation","Continuous learning","Creative problem-solving","Systems thinking"]

Design and evolve scalable software architecture for distributed AI training and real-time inference, optimizing throughput, latency, resiliency, and memory efficiency across cluster-scale deployments. Develop and evaluate communication and runtime capabilities using NCCL, UCX, and UCC, and collaborate with AI framework and internal platform teams to improve end-to-end performance. Work with hardware/system teams across GPUs, DPUs, and interconnects to accelerate data movement and enable robust AI workloads at scale.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
3 days ago

Senior HPC and AI Network Software Architect

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 1 hour agoStatus: Live
Reposted: similar role first listed 5 months ago

Job Summary

Design and evolve scalable software architecture for distributed AI training and real-time inference, optimizing throughput, latency, resiliency, and memory efficiency across cluster-scale deployments. Develop and evaluate communication and runtime capabilities using NCCL, UCX, and UCC, and collaborate with AI framework and internal platform teams to improve end-to-end performance. Work with hardware/system teams across GPUs, DPUs, and interconnects to accelerate data movement and enable robust AI workloads at scale.
Location: Zurich
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering
Seniority: Sr. Manager level

Key Responsibilities

  • •Build and evolve scalable software system architecture for distributed AI training and inference, optimizing throughput, latency, resiliency, and memory efficiency across clusters.
  • •Develop and evaluate next-generation communication and runtime capabilities in libraries such as NCCL, UCX, and UCC to meet frontier AI workload demands.
  • •Partner with AI framework teams (TensorFlow, PyTorch, JAX) and internal platform teams to build integrations and improve end-to-end performance and reliability.
  • •Collaborate with hardware and system-level teams across GPUs, DPUs, and interconnects to speed data movement and enable training, inference, and model serving at scale.
  • •Drive innovation across runtime systems, communication libraries, and AI protocol layers by turning new ideas into practical capabilities and robust implementations.

Key Requirements

  • •Ph.D., or equivalent industry experience, in computer science, computer engineering, or a closely related field.
  • •5+ years designing and building complex systems in systems programming, parallel/distributed computing, high-performance networking, or large-scale data movement.
  • •Strong programming background in C++ and Python, ideally including CUDA or other GPU programming models, with production-quality performance-critical software.
  • •Extensive hands-on experience with AI frameworks (PyTorch, TensorFlow, JAX) and understanding of communication libraries and runtime systems for large-scale training/inference.
  • •Demonstrated success building high-throughput, low-latency systems and reasoning across software stacks, hardware capabilities, and system bottlenecks.
Experience:5+ yearsAIHPCDistributed computingHigh-performance networkingLarge-scale systems
Education:PhD / Doctorate in computer science, computer engineering, or closely related field
Skills:CollaborationInnovationContinuous learningCreative problem-solvingSystems thinking
Tech Stack:C++PythonCUDANCCLUCXUCCTensorFlowPyTorchJAXGPUsDPUsInterconnectsRDMACollective communicationsCongestion-aware transportAccelerator-aware networking

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor