Manager, Cloud Services and Site Reliability
Ottawa
Workplace: HybridFull timeFunction: DevOps, Cloud & InfrastructureExperience: 5+ yearsSkills: ["People leadership","Communication","Structured problem solving","Collaboration","Ownership"]Lead a hybrid SRE team responsible for reliability, availability, scalability, and operational excellence for high-volume SaaS services. Drive reliability engineering practices including SLOs/SLIs, monitoring, alerting, capacity planning, and service health reporting. Own incident management and post-incident reviews, champion automation and tooling, and use operational metrics and risk indicators to prioritize improvements while partnering across engineering, product, platform, and security.
Loading
Loading job details...
Preparing the role view and application actions.

