Staff Site Reliability Engineer – Automation and Platform

Cerebras
Sunnyvale, Toronto
Workplace: HybridFull timeFunction: DevOps, Cloud & InfrastructureExperience: 8+ yearsSkills: ["Lead complex projects","Cross-functional influence","Clear technical communication","Mentorship"]

Build and lead a high-performance SRE function for ultra-reliable AI inference infrastructure powered by the Wafer-Scale Engine. Drive toil elimination at scale with self-service delivery pipelines and shared observability, then architect the “tomorrow” layer: declarative GitOps-driven CD, capacity provisioning, and cluster upgrades. Partner with an early-career SRE sub-team, mentor engineers, and define reliability practices using SLOs/SLIs, error budgets, chaos testing, and capacity forecasting.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Cerebras
Cerebras
11 months ago

Staff Site Reliability Engineer – Automation and Platform

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 16 hours agoStatus: Live

Job Summary

Build and lead a high-performance SRE function for ultra-reliable AI inference infrastructure powered by the Wafer-Scale Engine. Drive toil elimination at scale with self-service delivery pipelines and shared observability, then architect the “tomorrow” layer: declarative GitOps-driven CD, capacity provisioning, and cluster upgrades. Partner with an early-career SRE sub-team, mentor engineers, and define reliability practices using SLOs/SLIs, error budgets, chaos testing, and capacity forecasting.
Location: Sunnyvale, Toronto
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Define and implement a strategy for delivering and running software reliably and at scale across multiple datacenters and cloud-based solutions.
  • •Architect self-service platforms and internal tooling for safely triggering and observing critical workflows with minimal handoffs.
  • •Define and evolve reliability practices for inference workloads, including SLOs/SLIs, error budgets, blameless postmortems, chaos testing, and capacity forecasting.
  • •Mentor mid-level SREs and support critical incident escalations while prioritizing high-leverage automation based on production pain points.
  • •Measure and drive impact using metrics such as toil reduction, deployment velocity, SLO compliance, MTTR, and adoption of self-service workflows.

Key Requirements

  • •8+ years in SRE, infrastructure engineering, or platform engineering, with a strong record improving automation and reliability at large scale.
  • •Deep expertise operating large-scale heterogeneous clusters with a proprietary cloud control plane.
  • •Design and deliver CI/CD or GitOps systems using Argo CD (or similar) with safety and observability built in.
  • •Hands-on experience with observability tools such as Loki, Tempo, Mimir, and Prometheus.
  • •Ability to lead complex end-to-end projects, influence cross-functional stakeholders, and communicate technical direction clearly.
Experience:8+ yearsSREPlatform engineeringInfrastructure engineeringFAANGHyperscaler
Skills:Lead complex projectsCross-functional influenceClear technical communicationMentorship
Tech Stack:GitOpsArgo CDCI/CDLokiTempoMimirPrometheusSLOSLIGitOps-driven CDCapacity forecastingChaos testingBazelWafer-Scale Engine (WSE)

Company Brief

Cerebras
Designs and builds wafer-scale AI accelerators and systems for large-scale deep learning workloads, delivering specialized hardware and software to accelerate model training and inference for enterprises and research institutions.
Industry: Hardware Devices
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Headquarters: Sunnyvale, United States
Founded: 2016
WebsiteLinkedIn