Staff Software Engineer, GPU Inference
Toronto, Sunnyvale
Workplace: HybridFull timeFunction: Software EngineeringExperience: 8+ yearsEducation: bachelorsSkills: ["Technical leadership","Communication","Cross-functional collaboration"]Build and optimize a GPU inference serving stack for disaggregated AI inference systems, focusing on the full GPU prefill path from API services through vLLM, PyTorch, and the ROCm stack. Drive production reliability with operational readiness, automated recovery, and incident response. Improve latency, throughput, GPU utilization, and memory efficiency, while debugging performance across application, runtime, distributed systems, and hardware layers.

