Senior Deep Learning Communication Architect

NVIDIA
Santa Clara, Austin
Workplace: OnsiteFull timeFunction: Data Science & Machine LearningExperience: 6+ yearsEducation: bachelorsSkills: ["Communication","Collaboration","Problem-solving"]

Lead and optimize scalable deep learning communication for distributed training/inference on NVIDIA’s systems, working with high‑speed interconnects and libraries to scale DNNs across hundreds of thousands of nodes. Collaborate with hardware/software teams to reduce bottlenecks, design efficient protocols, and evaluate new technologies to improve performance and scalability of DL workloads.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
6 months ago

Senior Deep Learning Communication Architect

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 4 hours agoStatus: Live

Job Summary

Lead and optimize scalable deep learning communication for distributed training/inference on NVIDIA’s systems, working with high‑speed interconnects and libraries to scale DNNs across hundreds of thousands of nodes. Collaborate with hardware/software teams to reduce bottlenecks, design efficient protocols, and evaluate new technologies to improve performance and scalability of DL workloads.
Location: Santa Clara, Austin
Workplace: Onsite
Employment Type: Full time
Job Function: Data Science & Machine Learning

Key Responsibilities

  • •Scale DL models and training/inference frameworks to systems with hundreds of thousands of nodes
  • •Identify and eliminate bottlenecks in data transfer and synchronization during distributed deep learning training and inference
  • •Develop and implement communication algorithms and protocols tailored for deep learning workloads, minimizing overhead and latency
  • •Collaborate with hardware and software teams to craft systems applying high-speed interconnects (NVLink, InfiniBand, SPC-X) and libraries (MPI, NCCL, UCX, UCC, NVSHMEM)
  • •Research and evaluate new communication technologies to enhance performance and scalability of deep learning systems

Pay and Benefits

Equity and Bonus:Equity
Perks:Equity

Key Requirements

  • •Ph.D., Masters, or BS in Computer Science (CS), Electrical Engineering (EE), Computer Science and Electrical Engineering (CSEE), or closely related field or equivalent experience
  • •6+ years of experience in Building DNNs, Scaling of DNNs, Parallelism of DNN frameworks, or deep learning training and inference workloads
  • •Experience in evaluating, analyzing, and optimizing LLM training and inference performance of state-of-the-art models on cutting-edge hardware
  • •Deep understanding of parallelism techniques, including Data Parallelism, Pipeline Parallelism, Tensor Parallelism, Expert Parallelism, and FSDP
  • •Proficiency in developing code for one or more DNN training and Inference frameworks, such as PyTorch, TensorRT-LLM, vLLM, SGLang; strong programming in C++ and Python; familiarity with CUDA/OpenCL and InfiniBand/RoCE
Experience:6+ yearsDeep learningDistributed systemsHigh-performance computing
Education:Bachelor's in CS/EE
Skills:CommunicationCollaborationProblem-solving
Tech Stack:PyTorchTensorRT-LLMVLLMSGLangC++PythonCUDAOpenCLInfiniBandRoCEMPINCCLUCXUCCNVSHMEM

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor