Engineering Manager, Deep Learning Inference

NVIDIA
United States
Workplace: HybridFull timeUSD 224,000 - 431,250 annuallyFunction: Data Science & Machine LearningExperience: 6+ yearsEducation: mastersSkills: ["Leadership","Mentorship","Collaboration","Technical excellence","Continuous innovation"]

Lead and scale the Deep Learning Inference engineering team to advance AI model deployment on NVIDIA GPUs. Guide strategy, roadmap, and OSS framework execution for scalable inference using SGLang, vLLM, FlashInfer, and related technologies. Partner across compilers, libraries, and research teams to deliver optimized end-to-end inference pipelines, driving performance tuning, profiling, and multi-GPU acceleration (CUDA, Triton, CUTLASS, NCCL/NVSHMEM).

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
1 month ago

Engineering Manager, Deep Learning Inference

âś“ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 17 hours agoStatus: Live

Job Summary

Lead and scale the Deep Learning Inference engineering team to advance AI model deployment on NVIDIA GPUs. Guide strategy, roadmap, and OSS framework execution for scalable inference using SGLang, vLLM, FlashInfer, and related technologies. Partner across compilers, libraries, and research teams to deliver optimized end-to-end inference pipelines, driving performance tuning, profiling, and multi-GPU acceleration (CUDA, Triton, CUTLASS, NCCL/NVSHMEM).
Location: United States
Workplace: Hybrid
Employment Type: Full time
Job Function: Data Science & Machine Learning
Seniority: Manager level

Key Responsibilities

  • •Lead, mentor, and scale an engineering team focused on deep learning inference and GPU-accelerated software.
  • •Own the strategy, roadmap, and execution for NVIDIA’s OSS inference frameworks engineering.
  • •Partner with internal compiler, libraries, and research teams to deliver optimized inference pipelines across NVIDIA accelerators.
  • •Drive performance tuning, profiling, and optimization for LLM, multimodal, and generative AI inference workloads.
  • •Guide engineering best practices for CUDA, Triton, CUTLASS, and multi-GPU communications; represent team in roadmap and planning.
  • •Foster a culture of technical excellence, open collaboration, and continuous innovation.

Pay and Benefits

Salary: USD 224,000 - 431,250 annually

Key Requirements

  • •MS, PhD, or equivalent experience in Computer Science, Electrical/Computer Engineering, or related field.
  • •6+ years of software development experience, including 3+ years in technical leadership or engineering management.
  • •Strong C/C++ software design and development skills; Python proficiency is a plus.
  • •Hands-on GPU programming and performance optimization experience (CUDA, Triton, CUTLASS).
  • •Proven track record deploying or optimizing deep learning models in production environments.
Experience:6+ yearsDeep learningAI model deploymentGPU programmingLLM servingOpen source
Education:Master's
Skills:LeadershipMentorshipCollaborationTechnical excellenceContinuous innovation
Tech Stack:CC++PythonCUDATritonCUTLASSNIXLNCCLNVSHMEMSGLangVLLMFlashInferPyTorchTensorRT-LLMAgileMulti-GPUProfilingPerformance optimization

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor