Staff Site Reliability Engineer – Automation and Platform
Cerebras
Sunnyvale, Toronto
Workplace: HybridFull timeFunction: DevOps, Cloud & InfrastructureExperience: 8+ yearsSkills: ["Lead complex projects","Cross-functional influence","Clear technical communication","Mentorship"]Build and lead a high-performance SRE function for ultra-reliable AI inference infrastructure powered by the Wafer-Scale Engine. Drive toil elimination at scale with self-service delivery pipelines and shared observability, then architect the “tomorrow” layer: declarative GitOps-driven CD, capacity provisioning, and cluster upgrades. Partner with an early-career SRE sub-team, mentor engineers, and define reliability practices using SLOs/SLIs, error budgets, chaos testing, and capacity forecasting.

