Site Reliability Engineer - Ops & Automation
Cerebras
Toronto
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 2+ yearsSkills: ["Hands-on operational execution","Automation mindset","Collaboration","Communication","Impact measurement"]Own real production systems for a rapidly growing AI inference service, helping operate the current environment and bring up new capacity in high-stakes setups. Build and extend continuous delivery and self-service pipelines, improve telemetry/observability/alerting for reliability at scale, and reduce operational toil with reusable automation and internal developer tools. Collaborate with Cluster Ops and development teams on SLOs, post-mortems, and capacity planning (no 24/7 on-call).

