Engineering Manager, Deep Learning Inference

NVIDIA
Santa Clara, District of Columbia, Texas, New York, Washington, Massachusetts
Workplace: HybridFull timeUSD 184,000 - 356,500 annuallyFunction: Data Science & Machine LearningExperience: 6+ yearsEducation: mastersSkills: ["Mentorship","Technical leadership","Collaboration","Engineering management","Technical excellence"]

Lead and grow an engineering team focused on deep learning inference and GPU-accelerated software. Own the strategy, roadmap, and execution behind NVIDIA’s inference frameworks, delivering end-to-end optimized inference pipelines across NVIDIA accelerators. Drive performance tuning, profiling, and optimization for LLM and multimodal generative AI workloads, while partnering with compiler, libraries, and research teams and guiding best practices across CUDA, Triton, CUTLASS, and multi-GPU communications.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
2 hours ago

Engineering Manager, Deep Learning Inference

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 2 hours agoStatus: Live

Job Summary

Lead and grow an engineering team focused on deep learning inference and GPU-accelerated software. Own the strategy, roadmap, and execution behind NVIDIA’s inference frameworks, delivering end-to-end optimized inference pipelines across NVIDIA accelerators. Drive performance tuning, profiling, and optimization for LLM and multimodal generative AI workloads, while partnering with compiler, libraries, and research teams and guiding best practices across CUDA, Triton, CUTLASS, and multi-GPU communications.
Location: Santa Clara, District of Columbia, Texas, New York, Washington, Massachusetts
Workplace: Hybrid
Employment Type: Full time
Job Function: Data Science & Machine Learning
Seniority: Manager level

Key Responsibilities

  • •Lead, mentor, and scale an engineering team focused on deep learning inference and GPU-accelerated software.
  • •Drive strategy, roadmap, and execution for NVIDIA inference frameworks engineering, focusing on Client AI.
  • •Partner with internal compiler, libraries, and research teams to deliver end-to-end optimized inference pipelines.
  • •Oversee performance tuning, profiling, and optimization for LLM, multimodal, and generative AI applications.
  • •Guide engineers in best practices for CUDA, Triton, CUTLASS, and multi-GPU communications (NIXL, NCCL, NVSHMEM).

Pay and Benefits

Salary: USD 184,000 - 356,500 annually
Equity and Bonus:Equity

Key Requirements

  • •MS, PhD, or equivalent experience in Computer Science, Electrical/Computer Engineering, or a related field.
  • •6+ years of software development experience, including 3+ years in technical leadership or engineering management.
  • •Strong C/C++ software design and development experience; Python proficiency is a plus.
  • •Hands-on GPU programming and performance optimization experience with CUDA, Triton, and CUTLASS.
  • •Proven experience deploying or optimizing deep learning models in production environments.
Experience:6+ yearsDeep learningLLM servingGPU programmingOpen source
Education:Master's
Skills:MentorshipTechnical leadershipCollaborationEngineering managementTechnical excellence
Tech Stack:C/C++PythonCUDATritonCUTLASSNIXLNCCLNVSHMEMVLLMSGLangFlashInferPyTorchTensorRT-LLMAgile

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor