Senior Software Engineer, Quantized Inference
NVIDIA
Redmond, Santa Clara
Workplace: OnsiteFull timeUSD 152,000 - 287,500 annuallyFunction: Software EngineeringExperience: 4+ yearsEducation: mastersSkills: ["Software engineering fundamentals","Written communication","Verbal communication","Debugging","Collaboration"]Accelerate the discovery and deployment of efficient quantized and sparse inference “recipes” for LLMs. Translate recipe specifications into performant inference-engine implementations across vLLM, TRT-LLM, and SGLang—writing Triton kernels, inserting quantize/dequantize nodes, and handling MoE scaling correctly. Own model export pipelines (ModelOpt, Megatron-LM, HuggingFace), build benchmarking harnesses and numerics debugging tools, and improve developer productivity through CI and build system enhancements.

