DL Performance Software Engineer - LLM Inference

NVIDIA
Toronto
Workplace: HybridFull timeCAD 135,000 - 220,000 annuallyFunction: Software EngineeringExperience: 5+ yearsEducation: bachelorsSkills: ["Debugging","Problem-solving","Communication"]

Architect and implement high-performance LLM inference systems for NVIDIA’s large-scale models. Work on vLLM to add features for the latest NVIDIA GPUs, optimize inference frameworks with techniques like speculative decoding and parallelism, and design runtime/kernels for benchmarking and efficiency. Develop and optimize GPU kernels using CUDA and profiling tools, and collaborate with inference performance, kernels, training, serving, and research teams to advance ML systems research into production-grade open source software.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
NVIDIA
NVIDIA
16 hours ago

DL Performance Software Engineer - LLM Inference

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 16 hours agoStatus: Live

Job Summary

Architect and implement high-performance LLM inference systems for NVIDIA’s large-scale models. Work on vLLM to add features for the latest NVIDIA GPUs, optimize inference frameworks with techniques like speculative decoding and parallelism, and design runtime/kernels for benchmarking and efficiency. Develop and optimize GPU kernels using CUDA and profiling tools, and collaborate with inference performance, kernels, training, serving, and research teams to advance ML systems research into production-grade open source software.
Location: Toronto
Workplace: Hybrid
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Contribute features to vLLM that enable newer models with NVIDIA GPU hardware features and serving runtime algorithms.
  • •Optimize the vLLM inference framework using methods like speculative decoding, 5D Parallelism, and prefill-decode disaggregation.
  • •Architect runtime optimizations for inference infrastructure, benchmarking, and kernels.
  • •Develop, optimize, and benchmark GPU kernels using fusion, autotuning, and memory/layout optimization techniques.
  • •Conduct and publish ML Systems research, then integrate research ideas and prototypes into production-grade open-source software.

Pay and Benefits

Salary: CAD 135,000 - 220,000 annually
Equity and Bonus:Equity

Key Requirements

  • •Bachelor’s, Master’s, or PhD degree in Computer Science, Computer Engineering, or Software Engineering.
  • •5+ years of industry software engineering experience (or equivalent research experience).
  • •Strong programming in Python and one of C/C++, Go, or Rust.
  • •Solid CS fundamentals: algorithms & data structures, operating systems, computer architecture, parallel programming, software engineering, distributed systems, and deep learning theories.
  • •Knowledge and passion for performance engineering in ML frameworks and inference engines (e.g., PyTorch, vLLM, SGLang), plus profiling/debugging skills (e.g., Nsight Systems/Compute).
Experience:5+ years
Education:Bachelor's
Skills:DebuggingProblem-solvingCommunication
Tech Stack:PythonCC++GoRustPyTorchVLLMSGLangCUDANCCLNsight SystemsNsight ComputeTritonMLIRLLVMXLACUTLASSCUDA GraphTensor CoresSpeculative decoding

Company Brief

NVIDIA
Designs and manufactures GPUs, AI accelerators, and system-on-chip products for gaming, data centers, professional visualization, and automotive markets, enabling advanced graphics, AI, and high-performance computing solutions worldwide.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1993
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor