Principal Software Engineer, E2E Performance and Goodput — CSP Engagements

NVIDIA
United States
Workplace: HybridFull timeUSD 272,000 - 431,250 annuallyFunction: Software EngineeringExperience: 15+ yearsEducation: bachelorsSkills: ["Python","Nsight Systems","Nsight Compute","DCGM","CUDA","NCCL","GPU","Megatron-LM","DeepSpeed","FSDP","TensorRT","Pandas","Dashboards"]

Lead end-to-end performance characterization for CSP/hyperscale engagements, driving cross-team collaboration to optimize CUDA/NCCL stacks, profiling, and tooling. Own measurement benchmarks, collect workload feedback, and ensure performance targets are met in customer configurations while enabling scalable, reproducible performance improvements across NVIDIA platforms.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
2 months ago

Principal Software Engineer, E2E Performance and Goodput — CSP Engagements

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 18 hours agoStatus: Live

Job Summary

Lead end-to-end performance characterization for CSP/hyperscale engagements, driving cross-team collaboration to optimize CUDA/NCCL stacks, profiling, and tooling. Own measurement benchmarks, collect workload feedback, and ensure performance targets are met in customer configurations while enabling scalable, reproducible performance improvements across NVIDIA platforms.
Location: United States
Workplace: Hybrid
Employment Type: Full time
Job Function: Software Engineering

Key Responsibilities

  • •Drive performance characterization work streams with engineering teams of key CSP/hyperscale customers — ensuring they understand platform performance expectations, profiling methodology, and tuning options for their specific workloads
  • •Gather and synthesize CSP performance feedback — identify gaps between expected and actual throughput, and champion optimization priorities back into NVIDIA's CUDA, NCCL, driver, and firmware teams
  • •Ensure key open-source performance and stress tools (e.g., STREAM, GPU Burn, GPU BLAST) are updated and validated for the latest NVIDIA rack-scale systems, GPU architectures, and CPU platforms — so customers and internal teams have reliable baseline measurements from day one
  • •Work closely with CSPs to ensure their own performance and validation tooling reflects the latest GPU capabilities, memory hierarchy changes, and platform-specific tuning parameters
  • •Conduct cross-CSP performance comparison and pattern analysis — identify configuration, software, or workload differences that explain performance gaps between deployments

Pay and Benefits

Salary: USD 272,000 - 431,250 annually
Equity and Bonus:Equity
Perks:Equity

Key Requirements

  • •15+ years of experience in systems performance engineering, ideally in GPU/HPC/ML infrastructure; BS or MS in Computer Science, Computer Engineering, or related field (or equivalent experience)
Experience:15+ yearsGpuHpcMlInfrastructureDistributed training
Education:Bachelor's
Skills:PythonNsight SystemsNsight ComputeDCGMCUDANCCLGPUMegatron-LMDeepSpeedFSDPTensorRTPandasDashboards
Languages:English
Tech Stack:PythonNsight SystemsNsight ComputeDCGMCUDANCCLGPUMegatron-LMDeepSpeedFSDPTensorRTPandasDashboards

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor