Staff Software Engineer, Inference API
Toronto
Workplace: HybridFull timeFunction: Software EngineeringExperience: 5+ yearsEducation: bachelorsSkills: ["Cross-functional execution","Communication"]Build and evolve the ML inference API layer for disaggregated AI serving, creating consistent request/response semantics across GPU prefill, Cerebras decode, and other heterogeneous backends. You’ll integrate model-serving runtimes (including vLLM and PyTorch), enable new LLM capabilities (streaming, sampling, tool use, structured outputs, multimodal inputs), and improve performance, correctness, reliability, and observability for production inference traffic.
Loading
Loading job details...
Preparing the role view and application actions.

