Senior AI Training Performance Architect

NVIDIA
United States
Workplace: OnsiteFull timeUSD 184,000 - 356,500 annuallyFunction: Software EngineeringExperience: 5+ yearsEducation: phdSkills: ["Performance analysis","Performance optimization","Problem solving","Cross-layer systems thinking"]

Design and optimize AI training workloads end-to-end to extract peak performance from NVIDIA’s hardware and software stack. Profile GPU bottlenecks, implement production software across the deep learning platform stack, and build tools for automated analysis and optimization. Support MLPerf Training benchmark submissions and implement training workloads in NVIDIA’s proprietary processor and system simulators to enable future architecture studies.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
1 month ago

Senior AI Training Performance Architect

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 2 hours agoStatus: Live

Job Summary

Design and optimize AI training workloads end-to-end to extract peak performance from NVIDIA’s hardware and software stack. Profile GPU bottlenecks, implement production software across the deep learning platform stack, and build tools for automated analysis and optimization. Support MLPerf Training benchmark submissions and implement training workloads in NVIDIA’s proprietary processor and system simulators to enable future architecture studies.
Location: United States
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Understand, analyze, profile, and optimize AI training workloads on NVIDIA’s hardware and software platforms.
  • •Identify AI training performance bottlenecks on GPUs and solve problems across key training workloads.
  • •Implement production-quality software across multiple layers of the deep learning platform stack, from drivers to DL frameworks.
  • •Build and support NVIDIA submissions for MLPerf Training benchmarks.
  • •Implement DL training workloads in NVIDIA’s proprietary processor and system simulators and develop tools to automate analysis and optimization.

Pay and Benefits

Salary: USD 184,000 - 356,500 annually
Equity and Bonus:Equity

Key Requirements

  • •PhD in CS, EE, or CSEE (or equivalent) with 5+ years of relevant experience, or an MS with 8+ years of experience.
  • •Strong background in deep learning and neural networks, especially training.
  • •Solid understanding of computer architecture with familiarity with GPU architecture fundamentals.
  • •Proven experience analyzing and tuning application performance.
  • •Proficiency in programming with C++, Python, and CUDA, plus processor/system-level performance modeling experience.
Experience:5+ yearsDeep learningNeural networksAI trainingGPUComputer architecture
Education:PhD / Doctorate in CS, EE or CSEE
Skills:Performance analysisPerformance optimizationProblem solvingCross-layer systems thinking
Tech Stack:C++PythonCUDAGPU architectureDeep learning frameworksDL frameworksMLPerf Training

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor