Engineering Manager, Deep Learning Inference
Santa Clara, District of Columbia, Texas, New York, Washington, Massachusetts
Workplace: HybridFull timeUSD 184,000 - 356,500 annuallyFunction: Data Science & Machine LearningExperience: 6+ yearsEducation: mastersSkills: ["Mentorship","Technical leadership","Collaboration","Engineering management","Technical excellence"]Lead and grow an engineering team focused on deep learning inference and GPU-accelerated software. Own the strategy, roadmap, and execution behind NVIDIA’s inference frameworks, delivering end-to-end optimized inference pipelines across NVIDIA accelerators. Drive performance tuning, profiling, and optimization for LLM and multimodal generative AI workloads, while partnering with compiler, libraries, and research teams and guiding best practices across CUDA, Triton, CUTLASS, and multi-GPU communications.
Loading
Loading job details...
Preparing the role view and application actions.

