Senior System Software Architect, HPC and AI Networking

NVIDIA
Shanghai
Workplace: HybridFull timeFunction: Software EngineeringExperience: 5+ yearsSkills: ["Communication","Collaboration","Interpersonal skills","Ability to guide and influence","Flexibility"]

Architect scalable HPC and AI inference software systems for distributed training and real-time inference at NVIDIA. Design and prototype systems that optimize throughput, latency, and memory efficiency, and improve communication libraries like NCCL, UCX, and UCC for deep learning workloads. Collaborate with AI framework teams (TensorFlow, PyTorch, JAX) and co-design hardware features in GPUs, DPUs, and interconnects to accelerate data movement for inference and model serving.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
1 day ago

Senior System Software Architect, HPC and AI Networking

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 3 hours agoStatus: Live

Job Summary

Architect scalable HPC and AI inference software systems for distributed training and real-time inference at NVIDIA. Design and prototype systems that optimize throughput, latency, and memory efficiency, and improve communication libraries like NCCL, UCX, and UCC for deep learning workloads. Collaborate with AI framework teams (TensorFlow, PyTorch, JAX) and co-design hardware features in GPUs, DPUs, and interconnects to accelerate data movement for inference and model serving.
Location: Shanghai
Workplace: Hybrid
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Design and prototype scalable software systems to optimize distributed AI training and inference for throughput, latency, and memory efficiency.
  • •Develop and evaluate enhancements to communication libraries (NCCL, UCX, UCC) for deep learning workload demands.
  • •Collaborate with AI framework teams (TensorFlow, PyTorch, JAX) to improve communication backend integration, performance, and reliability.
  • •Co-design hardware features (GPUs, DPUs, interconnects) to accelerate data movement for inference and model serving.
  • •Evolve runtime systems, communication libraries, and AI-specific protocol layers, and collaborate with customers to deliver innovative solutions.

Key Requirements

  • •Ph.D, Masters, or Bachelors in computer science, computer engineering, electrical engineering, or a closely related field.
  • •5+ years of experience in DNNs, scaling, parallelism of deep learning training workloads, and distributed training/inference.
  • •Deep understanding of inference and training workload optimizations such as prefill/decode, data parallelism, tensor parallelism, and FDSP.
  • •Experience with AI network parallelism using collective libraries and RDMA/RoCE.
  • •Strong programming and software development skills, with background in algorithm design, system programming, and computer architecture.
Experience:5+ yearsDeep learningHPCDistributed trainingAI infrastructure
Education:
Skills:CommunicationCollaborationInterpersonal skillsAbility to guide and influenceFlexibility
Tech Stack:NCCLUCXUCCTensorFlowPyTorchJAXRDMARoCEDeep learningCUDAGPUsDPUsInterconnects

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor