Senior Software Architect - Deep Learning and HPC Communications

NVIDIA
Santa Clara, Austin, Durham
Workplace: HybridFull timeFunction: Communications, PR & CommunityExperience: 5+ yearsSkills: ["Collaboration","Interpersonal skills","Problem-solving"]

Senior Software Architect to co-design next-gen data center platforms and scalable communication software for DL and HPC workloads. You’ll identify bottlenecks, design new communication technologies, explore HW/SW co-design options, build proofs-of-concept, and simulate performance on large GPU clusters. Strong CS/CE foundations, parallel programming experience, and Linux expertise are essential.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
5 months ago

Senior Software Architect - Deep Learning and HPC Communications

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 4 hours agoStatus: Live

Job Summary

Senior Software Architect to co-design next-gen data center platforms and scalable communication software for DL and HPC workloads. You’ll identify bottlenecks, design new communication technologies, explore HW/SW co-design options, build proofs-of-concept, and simulate performance on large GPU clusters. Strong CS/CE foundations, parallel programming experience, and Linux expertise are essential.
Location: Santa Clara, Austin, Durham
Workplace: Hybrid
Employment Type: Full time
Job Function: Communications, PR & Community

Key Responsibilities

  • •Investigate opportunities to improve communication performance by identifying bottlenecks in today’s systems.
  • •Design and implement new communication technologies to accelerate AI and HPC workloads.
  • •Explore innovative solutions in HW and SW for our next generation platforms as part of co-design efforts involving GPU, Networking, and SW architects.
  • •Build proofs-of-concept, conduct experiments, and perform quantitive modeling to evaluate and drive new innovations.
  • •Use simulation to explore performance of large GPU clusters (think scales of 100s of 1000s of GPUs)

Key Requirements

  • •MS/Ph.D. degree in CS/CE or equivalent experience
  • •5+ years of relevant experience
  • •Excellent C/C++ programming and debugging skills
  • •Experience with parallel programming models (MPI, SHMEM) and at least one communication runtime (MPI, NCCL, NVSHMEM, OpenSHMEM, UCX, UCC)
  • •Deep understanding of operating systems, computer and system architecture
Experience:5+ yearsAIHPCCUDANVIDIA GPUs
Skills:CollaborationInterpersonal skillsProblem-solving
Tech Stack:CC++MPISHMEMNCCLNVSHMEMOpenSHMEMUCXUCCLinuxCUDAInfiniBandRoCENVLinkGPUsPyTorchTensorFlow

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor