Distinguished Software Architect - Deep Learning and HPC Communications

NVIDIA
Santa Clara
Workplace: OnsiteFull timeUSD 320,000 - 488,750 annuallyFunction: Communications, PR & CommunityEducation: phdSkills: ["Leadership","Communication","Collaboration","Problem-solving"]

Lead co-design of next-generation data center platforms for DL and HPC communications, researching new HPC/ML-enabled communication technologies, and collaborating across GPU, Networking, and Software teams to push performance at scale on thousands of GPUs. Drive adoption of new communication technologies, stay current with DL research, and influence open standards and open source projects.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
6 months ago

Distinguished Software Architect - Deep Learning and HPC Communications

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 4 hours agoStatus: Live

Job Summary

Lead co-design of next-generation data center platforms for DL and HPC communications, researching new HPC/ML-enabled communication technologies, and collaborating across GPU, Networking, and Software teams to push performance at scale on thousands of GPUs. Drive adoption of new communication technologies, stay current with DL research, and influence open standards and open source projects.
Location: Santa Clara
Workplace: Onsite
Employment Type: Full time
Job Function: Communications, PR & Community

Key Responsibilities

  • •Research and design new communication technologies and features for our libraries; expand GPUDirect portfolio and contribute to next-gen platforms.
  • •Co-design HW/SW solutions with GPU, Networking, and SW architects ensuring seamless integration with software stacks.
  • •Inspire changes using quantitative data from proofs-of-concept and analysis/modeling to improve performance.
  • •Drive adoption of new communication technologies across DL/HPC application verticals and across timezones.
  • •Collaborate with DL researchers, customers, and internal/external teams to stay at the forefront of DL research and networking

Pay and Benefits

Salary: USD 320,000 - 488,750 annually
Equity and Bonus:Equity
Perks:Equity

Key Requirements

  • •PhD in Computer Science, Computer Engineering or related field or strong equivalent experience; 15+ years of relevant experience in academia or the industry
  • •Expert in HPC, parallel programming models (MPI, SHMEM), at least one communication runtime (MPI, NCCL, NVSHMEM, OpenSHMEM, UCX, UCC), computer and system architecture, GPU architecture and CUDA
  • •Deep understanding of high performance networking: network technologies (Infiniband, Ethernet), network design, topology, debugging and performance analysis
  • •Strong in ML/DL fundamentals and how they tie to communications, parallel algorithms, fault tolerance and resiliency, performance analysis and optimizations for parallel applications on large clusters, DL Frameworks (PyTorch, TensorFlow)
  • •Programming fluency with C or C++ for systems software development
Experience:HPCDLCommunicationsGPU
Education:PhD / Doctorate
Skills:LeadershipCommunicationCollaborationProblem-solving
Tech Stack:CC++MPISHMEMNCCLNVSHMEMOpenSHMEMUCXUCCCUDAInfinibandEthernetPyTorchTensorFlow

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor