Senior Manager, Validation and HPC - NVIS

NVIDIA
United States
Workplace: RemoteFull timeUSD 216,000 - 396,750 annuallyFunction: Solutions Engineering & Sales EngineeringExperience: 10+ yearsEducation: bachelorsSkills: ["Leadership","Team development","Customer focus","Interpersonal communication","Problem-solving"]

Lead and supervise service HPC engineering for NVIDIA’s customer AI high-performance computing systems. You will plan, implement, and validate large-scale AI/HPC deployments, improve hardware/software bring-up and reliability, and drive cluster performance using modern interconnects and provisioning tools. Partner with customers and internal teams, author validation procedures and testing plans, and develop a multi-layered team spanning HPC infrastructure and service quality continuous improvement.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
5 days ago

Senior Manager, Validation and HPC - NVIS

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 8 hours agoStatus: Live

Job Summary

Lead and supervise service HPC engineering for NVIDIA’s customer AI high-performance computing systems. You will plan, implement, and validate large-scale AI/HPC deployments, improve hardware/software bring-up and reliability, and drive cluster performance using modern interconnects and provisioning tools. Partner with customers and internal teams, author validation procedures and testing plans, and develop a multi-layered team spanning HPC infrastructure and service quality continuous improvement.
Location: United States
Workplace: Remote
Employment Type: Full time
Job Function: Solutions Engineering & Sales Engineering
Seniority: Manager level

Key Responsibilities

  • •Direct and supervise service HPC engineering for designing, developing, installing, and validating customer AI HPC systems.
  • •Lead HPC project planning, implementation, and performance improvements across network and server platform bring-up and maintenance.
  • •Drive HPC hardware and software deployment and develop system validation procedures.
  • •Lead team activities, tests, and validation plans for customer HPC AI system implementations, scripts, and testing procedures.
  • •Develop team members, and build relationships with NVIDIA leaders, customers, partners, and collaborators to deliver service quality and continuous improvement.

Pay and Benefits

Salary: USD 216,000 - 396,750 annually
Equity and Bonus:Equity

Key Requirements

  • •10+ years of experience in IT, high-performance computing, or related fields; 3+ years in a management or leadership role.
  • •Expertise in HPC systems design, configuration, planning, and knowledge of HPC storage.
  • •Proficiency with low-latency/high-bandwidth interconnects, including InfiniBand and Ethernet.
  • •Experience with HPC system software cluster management/provisioning tools, including job schedulers (Slurm) and provisioning tools (salt, xCAT).
  • •Hands-on with distributed parallelism and accelerators, including OpenMP, MPI, NCCL, HPL, and GPUs.
Experience:10+ yearsHPCData centersAIInfrastructureLinux/Unix
Education:Bachelor's in computer science, information systems, or a related field
Skills:LeadershipTeam developmentCustomer focusInterpersonal communicationProblem-solving
Tech Stack:InfiniBandEthernetSlurmSaltXCATOpenMPMPINCCLHPLGPUBashPerlPythonLinuxUnixCentOSSolarisAnsiblePuppetLustre

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor