Software Engineering Intern, DLFW Comms - 2027

NVIDIA
Shanghai, Shenzhen, Beijing
Full timeFunction: Software EngineeringSkills: ["Adaptability","Passion to learn","Communication","Flexibility","Rapid prototyping"]

Build and optimize deep-learning communication features for AI frameworks and multi-GPU workloads. You’ll analyze training and inference communication needs, integrate new library capabilities from proof-of-concept to performance evaluation, and author custom fused compute/communication kernels for NVIDIA platforms. Work closely with teams developing communication libraries such as NCCL and NVSHMEM, and develop fault-tolerant, elastic solutions for large-scale, dynamic AI workloads across multiple time zones.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
1 day ago

Software Engineering Intern, DLFW Comms - 2027

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 13 hours agoStatus: Live

Job Summary

Build and optimize deep-learning communication features for AI frameworks and multi-GPU workloads. You’ll analyze training and inference communication needs, integrate new library capabilities from proof-of-concept to performance evaluation, and author custom fused compute/communication kernels for NVIDIA platforms. Work closely with teams developing communication libraries such as NCCL and NVSHMEM, and develop fault-tolerant, elastic solutions for large-scale, dynamic AI workloads across multiple time zones.
Location: Shanghai, Shenzhen, Beijing
Employment Type: Full time
Job Function: Software Engineering
Seniority: Intern level

Key Responsibilities

  • •Integrate new communication library features in AI frameworks from proof-of-concept to performance analysis and production.
  • •Analyze AI workloads and frameworks to identify multi-GPU communication requirements and opportunities.
  • •Collaborate hands-on with teams working on the latest AI models.
  • •Author custom communication or fused compute-communication kernels to demonstrate top performance on NVIDIA platforms.
  • •Research and build fault-tolerant, elastic solutions for large-scale or dynamic AI workloads.

Key Requirements

  • •Pursuing an M.S. or Ph.D. in CE/CS/EE with strong background in communication, kernel authoring, and/or AI training/inference.
  • •Rapid prototyping and development with Python, C++, CUDA or related DSLs (e.g., Triton, cuTe).
  • •Solid understanding of LLM models and parallelism techniques.
  • •Ability to adapt quickly and learn new areas and tools.
  • •Flexibility to work and communicate effectively.
Education:
Skills:AdaptabilityPassion to learnCommunicationFlexibilityRapid prototyping
Tech Stack:PyTorchVLLMSGLangTRT-LLMVeRLNCCLNVSHMEMPythonC++CUDATritonCuTeJAXLLMExpert ParallelismTPDPPPCUDA kernel optimizationProfiling

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor