Staff Software Engineer, Inference API
Toronto
Workplace: HybridFull timeFunction: Software EngineeringExperience: 5+ yearsEducation: bachelorsSkills: ["Communication","Cross-functional execution"]Build and evolve the ML Inference API layer for a disaggregated serving system that combines GPU prefill with ultra-fast Cerebras decode. Own production inference APIs for chat, generation, streaming, tool use, structured outputs, multimodal inputs, and model configuration—delivering consistent semantics across heterogeneous backends. Collaborate across model enablement, compiler/runtime, cloud infrastructure, and product to ensure compatibility, reliability, correctness, observability, and strong developer experience.
Loading
Loading job details...
Preparing the role view and application actions.

