DL Performance Software Engineer - LLM Inference
Toronto
Workplace: HybridFull timeCAD 135,000 - 220,000 annuallyFunction: Software EngineeringExperience: 5+ yearsEducation: bachelorsSkills: ["Debugging","Problem-solving","Communication"]Architect and implement high-performance LLM inference systems for NVIDIA’s large-scale models. Work on vLLM to add features for the latest NVIDIA GPUs, optimize inference frameworks with techniques like speculative decoding and parallelism, and design runtime/kernels for benchmarking and efficiency. Develop and optimize GPU kernels using CUDA and profiling tools, and collaborate with inference performance, kernels, training, serving, and research teams to advance ML systems research into production-grade open source software.
Loading
Loading job details...
Preparing the role view and application actions.

