Senior Software Engineer – TensorRT Edge-LLM

NVIDIA
United States
Workplace: HybridFull timeUSD 152,000 - 287,500 annuallyFunction: Software EngineeringExperience: 4+ yearsEducation: bachelorsSkills: ["Collaboration","Software design","Execution"]

Build and evolve NVIDIA TensorRT Edge-LLM inference capabilities for real-time edge AI in automotive and robotics. Develop C++ framework extensions for autoregressive serving (speculative decoding, LoRA, MoE, KV cache), and implement compiler/runtime optimizations for transformer models on constrained platforms. Collaborate with teams across CUDA and robotics, contribute CUDA kernels/operators, and benchmark/profile performance to deliver production-ready low-latency generative AI on-device.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
6 months ago

Senior Software Engineer – TensorRT Edge-LLM

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 22 hours agoStatus: Live

Job Summary

Build and evolve NVIDIA TensorRT Edge-LLM inference capabilities for real-time edge AI in automotive and robotics. Develop C++ framework extensions for autoregressive serving (speculative decoding, LoRA, MoE, KV cache), and implement compiler/runtime optimizations for transformer models on constrained platforms. Collaborate with teams across CUDA and robotics, contribute CUDA kernels/operators, and benchmark/profile performance to deliver production-ready low-latency generative AI on-device.
Location: United States
Workplace: Hybrid
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Develop and evolve TensorRT Edge-LLM inference framework extensions for autoregressive serving, including speculative decoding, LoRA, MoE, and KV cache management.
  • •Design and implement compiler and runtime optimizations for transformer models on constrained, real-time edge platforms.
  • •Collaborate with teams across CUDA, kernel libraries, compilers, and robotics to deliver high-performance, production-ready solutions.
  • •Contribute to CUDA kernel and operator development for transformer components such as attention, GEMM, and MoE.
  • •Benchmark, profile, and optimize inference performance across embedded and automotive environments.

Pay and Benefits

Salary: USD 152,000 - 287,500 annually
Equity and Bonus:Equity

Key Requirements

  • •4+ years of relevant software development experience.
  • •BS, MS, PhD, or equivalent experience in Computer Science, Electrical/Computer Engineering, or a related field.
  • •Deep understanding of transformer models and inference optimization (e.g., quantization, tensor parallelism, memory-efficient scheduling).
  • •Proficient in modern C++ (C++11/14/17 and beyond).
  • •Familiarity with LLM inference frameworks/libraries such as TensorRT, TensorRT-LLM, vLLM, SGLang, MLC-LLM, or FlashInfer.
Experience:4+ years
Education:Bachelor's
Skills:CollaborationSoftware designExecution
Tech Stack:C++CUDATensorRTTensorRT-LLMVLLMSGLangMLC-LLMFlashInferSpeculative decodingLoRAMoEKV cache managementQuantizationTensor parallelism

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor