Software Engineering Intern, NCCL - 2026

NVIDIA
Shanghai, Beijing
Workplace: OnsiteFull timeFunction: Software EngineeringEducation: phdSkills: ["Debugging","Collaboration","Communication","Ability to work effectively across teams","Interpersonal skills"]

Build and maintain highly-optimized communication runtimes and system software for GPU cluster workloads. Contribute to parallel programming interface specifications such as MPI/OpenSHMEM, enabling efficient GPU-to-GPU and GPU-to-system interactions. Create proof-of-concepts to explore extensions to programming models, runtime designs, and new hardware features, with opportunities spanning deep learning frameworks and HPC communication libraries.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
1 day ago

Software Engineering Intern, NCCL - 2026

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 3 hours agoStatus: Live

Job Summary

Build and maintain highly-optimized communication runtimes and system software for GPU cluster workloads. Contribute to parallel programming interface specifications such as MPI/OpenSHMEM, enabling efficient GPU-to-GPU and GPU-to-system interactions. Create proof-of-concepts to explore extensions to programming models, runtime designs, and new hardware features, with opportunities spanning deep learning frameworks and HPC communication libraries.
Location: Shanghai, Beijing
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering
Seniority: Intern level

Key Responsibilities

  • •Design, implement, and maintain highly-optimized communication runtimes for deep learning frameworks (e.g., NCCL for TensorFlow/PyTorch) on GPU clusters.
  • •Design, implement, and maintain communication runtimes for HPC programming interfaces (e.g., UCX for MPI/OpenSHMEM) on GPU clusters.
  • •Contribute to and participate in parallel programming interface specifications like MPI/OpenSHMEM.
  • •Design, implement, and maintain system software enabling interactions among GPUs and between GPUs and other system components.
  • •Create proof-of-concepts to evaluate extensions to programming models, runtime designs, and new hardware features.

Key Requirements

  • •Pursuing a Ph.D. in CE/CS/EE with strong background in computer architecture, operating systems, and communication libraries and/or AI/ML.
  • •Excellent C/C++ programming and debugging skills.
  • •Strong experience with Linux.
  • •Experience with parallel programming interfaces and communication runtimes.
  • •Ability to work and communicate effectively in a multi-national, multi-time-zone corporate environment.
Experience:AI/MLHigh-performance computingDeep learningParallel programmingGPU clusters
Education:PhD / Doctorate in CE/CS/EE
Skills:DebuggingCollaborationCommunicationAbility to work effectively across teamsInterpersonal skills
Tech Stack:NCCLTensorFlowPyTorchUCXMPIOpenSHMEMLinuxCC++InfiniBandRoCEJAXXLAVLLMSGLangCUDAGPU clustersDeep learning frameworks

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor