Senior System Software Engineer - Dynamo-Triton Inference Server

NVIDIA
Santa Clara, California, Washington
Workplace: HybridFull timeUSD 224,000 - 356,500 annuallyFunction: Software EngineeringExperience: 12+ yearsEducation: mastersSkills: ["Communication","Debugging","Performance analysis","Test design","Agile teamwork"]

Develop GPU-accelerated AI inference serving software for the Dynamo-Triton Inference Server, driving high-performance, production-ready inference across LLM and non-LLM workloads. Contribute to feature development, unify NVIDIA Triton Inference Server and Dynamo stacks for feature parity, and optimize prediction throughput and latency. Collaborate in an agile team environment and actively engage with the open source deep learning software community.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
1 day ago

Senior System Software Engineer - Dynamo-Triton Inference Server

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 7 hours agoStatus: Live

Job Summary

Develop GPU-accelerated AI inference serving software for the Dynamo-Triton Inference Server, driving high-performance, production-ready inference across LLM and non-LLM workloads. Contribute to feature development, unify NVIDIA Triton Inference Server and Dynamo stacks for feature parity, and optimize prediction throughput and latency. Collaborate in an agile team environment and actively engage with the open source deep learning software community.
Location: Santa Clara, California, Washington
Workplace: Hybrid
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Develop GPU-accelerated AI inference serving software.
  • •Contribute to feature development and drive customer adoption.
  • •Drive convergence of the Triton Inference Server and NVIDIA Dynamo stacks into a unified, high-performance inference platform.
  • •Optimize prediction throughput and latency for production server or cloud deployments.
  • •Participate actively in the open source deep learning software engineering community.

Pay and Benefits

Salary: USD 224,000 - 356,500 annually
Equity and Bonus:Equity

Key Requirements

  • •MS or PhD in Computer Science or a relevant field (or equivalent experience).
  • •12+ years of professional experience working on deep learning software.
  • •Strong Rust and C++ skills, familiarity with Python, and solid software design and debugging ability.
  • •Experience with high-scale distributed systems and ML systems.
  • •Excellent communication skills and ability to work in a fast-paced, agile team environment.
Experience:12+ yearsDeep learningDistributed systemsML systemsOpen source
Education:Master's in Computer Science
Skills:CommunicationDebuggingPerformance analysisTest designAgile teamwork
Tech Stack:RustC++PythonTriton Inference ServerNVIDIA DynamoTensorRTPyTorchONNXOpenVINOVLLMTRT-LLMGitHub

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor