Senior Deep Learning Framework Communications Engineer

NVIDIA
Santa Clara
Workplace: OnsiteFull timeFunction: Communications, PR & CommunityExperience: 5+ yearsEducation: bachelorsSkills: ["Adaptability","Communication","Teamwork","Problem-solving"]

Hands-on engineer for integrating and optimizing deep learning communication libraries within AI frameworks at NVIDIA. You will analyze multi-GPU workloads, enhance compilers for fusion, and build fault-tolerant, scalable solutions across large clusters, collaborating across time zones. Work with PyTorch, JAX, TRT-LLM, NVSHMEM, NCCL, and CUDA to push performance and scalability in AI training and inference.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
6 months ago

Senior Deep Learning Framework Communications Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 4 hours agoStatus: Live

Job Summary

Hands-on engineer for integrating and optimizing deep learning communication libraries within AI frameworks at NVIDIA. You will analyze multi-GPU workloads, enhance compilers for fusion, and build fault-tolerant, scalable solutions across large clusters, collaborating across time zones. Work with PyTorch, JAX, TRT-LLM, NVSHMEM, NCCL, and CUDA to push performance and scalability in AI training and inference.
Location: Santa Clara
Workplace: Onsite
Employment Type: Full time
Job Function: Communications, PR & Community

Key Responsibilities

  • •Integrate new communication libraries features in AI frameworks: from PoC to performance analysis to production
  • •Perform deep analysis of AI workloads and frameworks to identify multi-GPU communication requirements and opportunities. Collaborate hands-on with teams working on the latest AI models
  • •Improve AI compilers to hide communications or perform automatic fusion
  • •Conduct in-depth AI workload performance characterization on multi-GPU clusters
  • •Design fault-tolerant and elastic solutions for large-scale or dynamic AI workloads

Key Requirements

  • •BS, MS, or PhD in Computer Science, or related field (or equivalent experience) with 5+ software engineering and HPC/AI experience
  • •Development or integration experience with Deep Learning Frameworks such as PyTorch, JAX, and Inference Engines such as TRT-LLM, vLLM, SGLang
  • •Rapid prototyping and development with Python, C++, CUDA or related DSLs (Triton, cuTe)
  • •Solid grasp of AI models, parallelisms, and/or compiler technologies (e.g. torch.compile)
  • •Experience conducting performance benchmarking on AI clusters. Familiarity with at least one performance profiler toolchain (PyTorch profiler, NVIDIA Nsight Systems)
Experience:5+ yearsAIHPCDeep LearningMulti-GPU
Education:Bachelor's in Computer Science
Skills:AdaptabilityCommunicationTeamworkProblem-solving
Languages:English
Tech Stack:PyTorchJAXNVIDIATRT-LLMVLLMSGLangPythonC++CUDATritonCuTeTorch.compileNCCLNVSHMEMGPUDirectNsight Systems

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor