Senior System Software Engineer, Agentic Inference - Dynamo

NVIDIA
United States
Workplace: HybridFull timeUSD 224,000 - 431,250 annuallyFunction: Software EngineeringEducation: mastersSkills: []

Build GPU-accelerated inference infrastructure for Dynamo, developing open-source components that serve trained AI models and support agentic workloads. You’ll advance disaggregated serving across vLLM, SGLang, and TensorRT-LLM, improve inference-state management with KV/prefix cache reuse via NIXL, and evolve a distributed inference frontend with stateful, multi-turn semantics. Partner with team leads to prioritize features, load-balance async requests, and optimize throughput under latency constraints.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
1 month ago

Senior System Software Engineer, Agentic Inference - Dynamo

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 14 hours agoStatus: Live

Job Summary

Build GPU-accelerated inference infrastructure for Dynamo, developing open-source components that serve trained AI models and support agentic workloads. You’ll advance disaggregated serving across vLLM, SGLang, and TensorRT-LLM, improve inference-state management with KV/prefix cache reuse via NIXL, and evolve a distributed inference frontend with stateful, multi-turn semantics. Partner with team leads to prioritize features, load-balance async requests, and optimize throughput under latency constraints.
Location: United States
Workplace: Hybrid
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Develop open-source software to serve inference of trained AI models running on GPUs.
  • •Contribute to disaggregated serving for Dynamo-supported inference engines (vLLM, SGLang, TRT-LLM) and extend to agentic inference workloads (long-horizon reasoning, tool calling, stateful multi-turn).
  • •Innovate inference-state management for long-running agents, including KV/prefix-cache reuse and transfer across heterogeneous memory and storage hierarchies with NIXL.
  • •Build and evolve Dynamo’s distributed inference frontend across vLLM, SGLang, and TensorRT-LLM, including day-0 support for new models and stateful Responses API semantics.
  • •Balance performance goals by building robust scalable components, load-balancing async requests, optimizing throughput under latency constraints, and integrating the latest open-source technology.

Pay and Benefits

Salary: USD 224,000 - 431,250 annually
Equity and Bonus:Equity

Key Requirements

  • •Masters or PhD (or equivalent) experience.
  • •10+ years in Computer Science, Computer Engineering, or a related field.
  • •Ability to work in a fast-paced, agile team environment.
  • •Excellent Rust/Python programming and software design skills, including debugging, performance analysis, and test design.
  • •Understanding of modern LLM API semantics, including structured outputs, tool calling, reasoning controls, token accounting, context management, and multimodal inputs.
Experience:Deep learningGenerative AILLM inferenceOpen source
Education:Master's
Tech Stack:RustPythonDynamoGPUsVLLMSGLangTRT-LLMTensorRT-LLMNIXLKV cachePrefix cacheLLM APITool callingStructured outputs

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor