Staff Software Engineer, Inference API

Cerebras
Toronto
Workplace: HybridFull timeFunction: Software EngineeringExperience: 5+ yearsEducation: bachelorsSkills: ["Communication","Cross-functional execution"]

Build and evolve the ML Inference API layer for a disaggregated serving system that combines GPU prefill with ultra-fast Cerebras decode. Own production inference APIs for chat, generation, streaming, tool use, structured outputs, multimodal inputs, and model configuration—delivering consistent semantics across heterogeneous backends. Collaborate across model enablement, compiler/runtime, cloud infrastructure, and product to ensure compatibility, reliability, correctness, observability, and strong developer experience.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Cerebras
Cerebras
4 days ago

Staff Software Engineer, Inference API

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 22 hours agoStatus: Live
Reposted: similar role first listed 6 days ago

Job Summary

Build and evolve the ML Inference API layer for a disaggregated serving system that combines GPU prefill with ultra-fast Cerebras decode. Own production inference APIs for chat, generation, streaming, tool use, structured outputs, multimodal inputs, and model configuration—delivering consistent semantics across heterogeneous backends. Collaborate across model enablement, compiler/runtime, cloud infrastructure, and product to ensure compatibility, reliability, correctness, observability, and strong developer experience.
Location: Toronto
Workplace: Hybrid
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Build production ML inference APIs for chat completions, text generation, streaming, model configuration, tool calling, structured outputs, and multimodal inputs.
  • •Create consistent request/response semantics across GPU prefill, Cerebras decode, and other heterogeneous inference backends.
  • •Integrate emerging models and capabilities into the serving platform, including tokenizers, prompt formats, sampling methods, attention variants, and model-specific features.
  • •Maintain API compatibility and evolution, including versioning, deprecation, validation, and backward compatibility practices.
  • •Extend and integrate inference services and runtimes (including vLLM, PyTorch, Hugging Face, ROCm stack, and Cerebras runtime components) and support disaggregated prefill/decode control and data paths.

Key Requirements

  • •5+ years of software engineering experience with substantial individual-contributor ownership of production software or distributed systems.
  • •Strong programming ability in Python and Go, and experience developing performance-sensitive or highly concurrent services in C++, Rust, or similar systems languages.
  • •Experience building stable APIs with validation, error handling, observability, compatibility, and versioning practices.
  • •Experience integrating software across service, framework, runtime, and infrastructure boundaries.
  • •Bachelor’s degree in computer science or a related discipline (or equivalent practical experience).
Experience:5+ yearsDistributed systemsMachine learningML inferenceLLMsAPI developmentModel serving
Education:Bachelor's in Computer Science, Computer Engineering, Electrical Engineering, or a related discipline
Skills:CommunicationCross-functional execution
Tech Stack:PythonGoC++RustVLLMPyTorchHugging FaceROCmAMD ROCmLinuxContainersKubernetesGRPCRESTStreamingCI/CDOpenAI-compatible APIsTritonTensorRT-LLMSGLang

Company Brief

Cerebras
Designs and builds wafer-scale AI accelerators and systems for large-scale deep learning workloads, delivering specialized hardware and software to accelerate model training and inference for enterprises and research institutions.
Industry: Hardware Devices
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Headquarters: Sunnyvale, United States
Founded: 2016
WebsiteLinkedIn