Staff Site Reliability Engineer – Automation and Platform
Sunnyvale, Toronto
Workplace: HybridFull timeFunction: DevOps, Cloud & InfrastructureExperience: 8+ yearsSkills: ["Lead complex projects","Cross-functional influence","Clear technical communication","Mentorship"]Build and lead a high-performance SRE function for ultra-reliable AI inference infrastructure powered by the Wafer-Scale Engine. Drive toil elimination at scale with self-service delivery pipelines and shared observability, then architect the “tomorrow” layer: declarative GitOps-driven CD, capacity provisioning, and cluster upgrades. Partner with an early-career SRE sub-team, mentor engineers, and define reliability practices using SLOs/SLIs, error budgets, chaos testing, and capacity forecasting.
Loading
Loading job details...
Preparing the role view and application actions.

