Engineering Manager, Deep Learning Inference
NVIDIA
Santa Clara, Georgia, District of Columbia, Illinois, California, Massachusetts
Workplace: HybridFull timeUSD 224,000 - 431,250 annuallyFunction: Data Science & Machine LearningExperience: 6+ yearsEducation: mastersSkills: ["Leadership","Mentorship","Collaboration","Technical excellence","Continuous innovation"]Lead and scale the Deep Learning Inference engineering team to advance AI model deployment on NVIDIA GPUs. Guide strategy, roadmap, and OSS framework execution for scalable inference using SGLang, vLLM, FlashInfer, and related technologies. Partner across compilers, libraries, and research teams to deliver optimized end-to-end inference pipelines, driving performance tuning, profiling, and multi-GPU acceleration (CUDA, Triton, CUTLASS, NCCL/NVSHMEM).

