Senior Deep Learning Algorithm Engineer

NVIDIA
Santa Clara
Workplace: HybridFull timeUSD 152,000 - 287,500 annuallyFunction: Data Science & Machine LearningExperience: 3+ yearsEducation: phdSkills: ["High agency","Leading end-to-end","Performance engineering","Collaboration"]

Build and maintain Dynamo integrations for open-source inference frameworks like vLLM, SGLang, and TensorRT-LLM. Lead architecture and performance work across distributed inference, partnering with research, software, systems, and hardware teams to improve latency, throughput, reliability, and efficiency. Identify and eliminate bottlenecks across runtimes and orchestration, and develop inference optimizations for scheduling, disaggregation, KV caching, and autoscaling.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
2 hours ago

Senior Deep Learning Algorithm Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 2 hours agoStatus: Live

Job Summary

Build and maintain Dynamo integrations for open-source inference frameworks like vLLM, SGLang, and TensorRT-LLM. Lead architecture and performance work across distributed inference, partnering with research, software, systems, and hardware teams to improve latency, throughput, reliability, and efficiency. Identify and eliminate bottlenecks across runtimes and orchestration, and develop inference optimizations for scheduling, disaggregation, KV caching, and autoscaling.
Location: Santa Clara
Workplace: Hybrid
Employment Type: Full time
Job Function: Data Science & Machine Learning
Seniority: Mid level

Key Responsibilities

  • •Design, build, and maintain Dynamo integrations for open-source frameworks vLLM, SGLang, and TRTLLM.
  • •Partner with open-source communities to improve latency, throughput, reliability, and efficiency.
  • •Push performance on public/private benchmarks focused on token/watt leadership.
  • •Identify and remove bottlenecks across runtimes, kernels, networking, routing, and orchestration.
  • •Develop inference optimizations for scheduling, disaggregation, KV caching, and autoscaling.

Pay and Benefits

Salary: USD 152,000 - 287,500 annually
Equity and Bonus:Equity

Key Requirements

  • •BS, MS, or PhD in Computer Science, Electrical Engineering, Computer Engineering, or a related field (or equivalent experience).
  • •3+ years building, profiling, and debugging performance-critical distributed or ML systems.
  • •Strong programming skills in Python and/or Rust, C++.
  • •Understanding of modern ML architectures and inference techniques.
  • •Experience leading ambiguous work end to end or contributing to open source frameworks/communities.
Experience:3+ yearsDistributed systemsML inferenceOpen sourceAI acceleratorsDeep learning
Education:PhD / Doctorate in Computer Science, Electrical Engineering, Computer Engineering, or related field
Skills:High agencyLeading end-to-endPerformance engineeringCollaboration
Tech Stack:DynamoVLLMSGLangTRTLLMTensorRT-LLMPythonRustC++

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor