Senior System Software Engineer - GPU Performance

NVIDIA
Santa Clara
Workplace: OnsiteFull timeFunction: Solutions Engineering & Sales EngineeringExperience: 3+ yearsEducation: mastersSkills: ["Collaboration","Communication","Problem-solving","Learning agility","Teamwork"]

Senior HPC Performance Engineer role at NVIDIA focusing on performance characterization and optimization of multi-GPU, multi-node HPC systems. You will analyze HW/SW interactions, benchmark large-scale clusters, triage performance issues, and build tools to visualize performance data, collaborating across time zones on cutting-edge communication libraries.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
5 months ago

Senior System Software Engineer - GPU Performance

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 4 hours agoStatus: Live

Job Summary

Senior HPC Performance Engineer role at NVIDIA focusing on performance characterization and optimization of multi-GPU, multi-node HPC systems. You will analyze HW/SW interactions, benchmark large-scale clusters, triage performance issues, and build tools to visualize performance data, collaborating across time zones on cutting-edge communication libraries.
Location: Santa Clara
Workplace: Onsite
Employment Type: Full time
Job Function: Solutions Engineering & Sales Engineering
Seniority: Sr. Manager level

Key Responsibilities

  • •Conduct in-depth performance characterization and analysis on large multi-GPU and multi-node clusters.
  • •Study the interaction of our libraries with all HW (GPU, CPU, Networking) and SW components in the stack
  • •Evaluate proof-of-concepts, conduct trade-off analysis when multiple solutions are available
  • •Triage and root-cause performance issues reported by our customers
  • •Collect a lot of performance data; build tools and infrastructure to visualize and analyze the information

Pay and Benefits

Equity and Bonus:Equity
Perks:Equity

Key Requirements

  • •MS or PhD in Computer Science or related field with relevant performance engineering and HPC experience
  • •3+ years of experience with parallel programming and at least one communication runtime (MPI, NCCL, UCX, NVSHMEM)
  • •Experience conducting performance benchmarking and triage on large scale HPC clusters
  • •Good understanding of computer system architecture, HW-SW interactions and operating systems principles
  • •Implement micro-benchmarks in C/C++, read and modify the code base when required; Proficient in Python
Experience:3+ yearsHPCGPUParallel computing
Education:Master's in Computer Science
Skills:CollaborationCommunicationProblem-solvingLearning agilityTeamwork
Languages:English
Tech Stack:MPINCCLUCXNVSHMEMC++PythonCUDAKubernetesSLURMAnsibleDocker

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor