Software Engineer, Inference Platform

Cerebras
Sunnyvale, Toronto
Workplace: OnsiteFull timeFunction: Software EngineeringExperience: 3+ yearsSkills: ["Engineering judgment","Operational rigor","Cross-team collaboration","Technical influence"]

Build and maintain production software for the Inference Platform team, which orchestrates inference on datacenter clusters by connecting cloud components with ML services. You’ll shape platform direction, define Kubernetes custom resource behaviors and system boundaries, and drive reliability through active-active architectures, rapid failover, SLOs, and incident-driven improvements. Partner with ML, Product, Infrastructure, and Cloud teams to translate requirements into scalable distributed system designs.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Cerebras
Cerebras
2 months ago

Software Engineer, Inference Platform

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 16 hours agoStatus: Live

Job Summary

Build and maintain production software for the Inference Platform team, which orchestrates inference on datacenter clusters by connecting cloud components with ML services. You’ll shape platform direction, define Kubernetes custom resource behaviors and system boundaries, and drive reliability through active-active architectures, rapid failover, SLOs, and incident-driven improvements. Partner with ML, Product, Infrastructure, and Cloud teams to translate requirements into scalable distributed system designs.
Location: Sunnyvale, Toronto
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Design, develop, test, and maintain production software across testing, continuous development, observability, security, networking, debugging, and productionization.
  • •Help shape the technical direction of the Inference Platform, including Kubernetes custom resource definitions, failure domains, service boundaries, and long-term system evolution.
  • •Architect active-active systems with rapid failover, graceful degradation, and clear SLOs; improve latency, throughput, capacity efficiency, and resilience.
  • •Write and review production code on critical parts of the platform and set the technical bar through design and code reviews.
  • •Lead hardest production issues and cross-system bottlenecks, driving observability, incident response, capacity planning, and post-incident improvements.

Key Requirements

  • •3+ years of software engineering experience building and operating large-scale distributed systems or cloud infrastructure.
  • •Experience with distributed systems, ideally with Kubernetes.
  • •Experience building highly available, latency-sensitive systems at scale.
  • •Experience with security including certificates, TLS, and mTLS.
  • •Strong proficiency in backend or systems languages such as Go or C++.
  • •Experience optimizing latency, throughput, and efficiency in high-QPS systems, including TTFT and tail-latency reduction (preferred).
Experience:3+ yearsDistributed systemsCloud infrastructureKubernetesHigh-QPS systemsML inference infrastructureGPU-accelerated workloadsModel serving
Skills:Engineering judgmentOperational rigorCross-team collaborationTechnical influence
Tech Stack:KubernetesGoC++TLSMTLSCI/CDKubernetes custom resource definitionsDistributed systems

Company Brief

Cerebras
Designs and builds wafer-scale AI accelerators and systems for large-scale deep learning workloads, delivering specialized hardware and software to accelerate model training and inference for enterprises and research institutions.
Industry: Hardware Devices
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Headquarters: Sunnyvale, United States
Founded: 2016
WebsiteLinkedIn