Staff Software Engineer, Inference API

Cerebras
Toronto
Workplace: HybridFull timeFunction: Software EngineeringExperience: 5+ yearsEducation: bachelorsSkills: ["Cross-functional execution","Communication"]

Build and evolve the ML inference API layer for disaggregated AI serving, creating consistent request/response semantics across GPU prefill, Cerebras decode, and other heterogeneous backends. You’ll integrate model-serving runtimes (including vLLM and PyTorch), enable new LLM capabilities (streaming, sampling, tool use, structured outputs, multimodal inputs), and improve performance, correctness, reliability, and observability for production inference traffic.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Cerebras
Cerebras
15 hours ago

Staff Software Engineer, Inference API

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 5 hours agoStatus: Live

Job Summary

Build and evolve the ML inference API layer for disaggregated AI serving, creating consistent request/response semantics across GPU prefill, Cerebras decode, and other heterogeneous backends. You’ll integrate model-serving runtimes (including vLLM and PyTorch), enable new LLM capabilities (streaming, sampling, tool use, structured outputs, multimodal inputs), and improve performance, correctness, reliability, and observability for production inference traffic.
Location: Toronto
Workplace: Hybrid
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Build production ML inference APIs for chat completions, text generation, streaming, model configuration, tool calling, structured outputs, multimodal inputs, and emerging inference capabilities.
  • •Deliver a unified serving experience by defining consistent request and response semantics across GPU prefill, Cerebras decode, and other heterogeneous inference backends.
  • •Enable new models and capabilities by integrating foundation models, tokenizers, prompt formats, sampling methods, attention variants, multimodal inputs, and model-specific features into the serving platform.
  • •Own API compatibility and evolution through versioning, deprecation, validation, and backward-compatibility practices for widely adopted inference interfaces and Cerebras extensions.
  • •Integrate with inference runtimes and disaggregated inference components (including vLLM and custom inference services), and improve performance, correctness, reliability, observability, and testing/qualification for production traffic.

Key Requirements

  • •5+ years of software engineering experience with substantial individual-contributor ownership of production software or distributed systems.
  • •Strong programming ability in Python, and experience developing performance-sensitive or highly concurrent services in C++, Go, or similar systems languages.
  • •Hands-on experience with model-serving frameworks such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, Hugging Face Text Generation Inference, or an equivalent platform.
  • •Understanding of modern LLM inference concepts, including tokenization, prompt formatting, sampling, streaming generation, continuous batching, KV-cache management, and model configuration.
  • •Experience building stable, observable, versioned APIs with strong validation, error handling, and compatibility practices; plus production experience with Linux, containers, and Kubernetes (or comparable orchestration) and CI/CD.
Experience:5+ years
Education:Bachelor's
Skills:Cross-functional executionCommunication
Tech Stack:PythonC++GoVLLMSGLangTensorRT-LLMTriton Inference ServerHugging FacePyTorchROCmAMD ROCmLinuxContainersKubernetesCI/CDGRPCREST

Company Brief

Cerebras
Designs and builds wafer-scale AI accelerators and systems for large-scale deep learning workloads, delivering specialized hardware and software to accelerate model training and inference for enterprises and research institutions.
Industry: Hardware Devices
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Headquarters: Sunnyvale, United States
Founded: 2016
WebsiteLinkedIn