Senior Software Architect - Deep Learning and HPC Communications

NVIDIA
United States
Workplace: OnsiteFull timeUSD 224,000 - 431,250 annuallyFunction: Software EngineeringExperience: 12+ yearsEducation: mastersSkills: ["C++","C","MPI","NCCL","NVSHMEM","OpenSHMEM","UCX","Linux","CUDA","GPU"]

Lead design and implementation of high-performance GPU communication technologies for DL and HPC workloads. Co-design next-generation data-center platforms with GPU, networking, and software teams. Build proofs-of-concept, run experiments, and perform quantitative modeling to identify bottlenecks and accelerate communication across large GPU clusters (scales of 100s to 1000s of GPUs).

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
2 months ago

Senior Software Architect - Deep Learning and HPC Communications

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 7 hours agoStatus: Live
Reposted: similar role first listed 6 months ago

Job Summary

Lead design and implementation of high-performance GPU communication technologies for DL and HPC workloads. Co-design next-generation data-center platforms with GPU, networking, and software teams. Build proofs-of-concept, run experiments, and perform quantitative modeling to identify bottlenecks and accelerate communication across large GPU clusters (scales of 100s to 1000s of GPUs).
Location: United States
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering

Key Responsibilities

  • •Investigate opportunities to improve communication performance by identifying bottlenecks in today’s systems.
  • •Design and implement new communication technologies to accelerate AI and HPC workloads.
  • •Explore innovative solutions in HW and SW for next-generation platforms as part of co-design efforts involving GPU, Networking, and SW architects.
  • •Build proofs-of-concept, conduct experiments, and perform quantitative modeling to evaluate and drive new innovations.
  • •Use simulation to explore performance of large GPU clusters (scales of hundreds to thousands of GPUs).

Pay and Benefits

Salary: USD 224,000 - 431,250 annually
Equity and Bonus:Equity

Key Requirements

  • •MS or PhD in CS/CE or equivalent experience
  • •12+ years of relevant experience
  • •Excellent C/C++ programming and debugging skills
  • •Experience with parallel programming models (MPI, SHMEM) and at least one communication runtime (MPI, NCCL, NVSHMEM, OpenSHMEM, UCX, UCC)
Experience:12+ yearsAIHPCDeep Learning
Education:Master's
Skills:C++CMPINCCLNVSHMEMOpenSHMEMUCXLinuxCUDAGPU
Languages:English
Tech Stack:C++CMPINCCLNVSHMEMOpenSHMEMUCXLinuxCUDANVIDIA GPUs

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor