Manager, Cloud Services and Site Reliability

Barracuda Networks
Ottawa
Workplace: HybridFull timeFunction: DevOps, Cloud & InfrastructureExperience: 5+ yearsSkills: ["People leadership","Communication","Structured problem solving","Collaboration","Ownership"]

Lead a hybrid SRE team responsible for reliability, availability, scalability, and operational excellence for high-volume SaaS services. Drive reliability engineering practices including SLOs/SLIs, monitoring, alerting, capacity planning, and service health reporting. Own incident management and post-incident reviews, champion automation and tooling, and use operational metrics and risk indicators to prioritize improvements while partnering across engineering, product, platform, and security.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Barracuda Networks
Barracuda Networks
4 days ago

Manager, Cloud Services and Site Reliability

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 1 hour agoStatus: Live

Job Summary

Lead a hybrid SRE team responsible for reliability, availability, scalability, and operational excellence for high-volume SaaS services. Drive reliability engineering practices including SLOs/SLIs, monitoring, alerting, capacity planning, and service health reporting. Own incident management and post-incident reviews, champion automation and tooling, and use operational metrics and risk indicators to prioritize improvements while partnering across engineering, product, platform, and security.
Location: Ottawa
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Manager level

Key Responsibilities

  • •Lead, coach, and develop an SRE team, setting expectations and fostering ownership, collaboration, and continuous improvement.
  • •Drive reliability engineering practices across critical services, including SLOs/SLIs, monitoring, alerting, capacity planning, and service health reporting.
  • •Partner with engineering and platform teams to improve the design, operation, and scalability of cloud-based systems with a focus on reliability and maintainability.
  • •Own and improve incident management, including major incident coordination, post-incident reviews, follow-up actions, and systemic reliability improvements.
  • •Champion automation and tooling to reduce manual effort and scale production support; use operational data and metrics to identify and prioritize reliability gaps.

Pay and Benefits

Perks:Health InsuranceRetirementEmployer MatchEquityPaid LeaveFlexible Time

Key Requirements

  • •5+ years of experience in SRE, DevOps, infrastructure, cloud operations, or related technical operations, including experience leading/managing technical teams.
  • •Strong understanding of cloud platforms, distributed systems, production operations, and modern reliability practices.
  • •Experience implementing or improving SLOs, SLIs, monitoring, alerting, incident response, and post-incident review practices.
  • •Ability to hire, mentor, coach, and develop engineers while building a healthy, accountable, inclusive team culture.
  • •Experience with infrastructure automation, CI/CD, disaster recovery, cost optimization, or multi-cloud operations.
Experience:5+ years
Skills:People leadershipCommunicationStructured problem solvingCollaborationOwnership
Tech Stack:SREDevOpsCloud platformsDistributed systemsSLOSLIMonitoringAlertingCI/CDDisaster recoveryMulti-cloud operations

Company Brief

Barracuda Networks
Provides cloud-enabled security and data protection solutions including email protection, network and application security, and backup and recovery services for businesses, service providers, and government organizations worldwide.
Industry: Cybersecurity
Company Size: Enterprise (1,001+ employees)
Revenue: USD 250M to 500M
Growth: Established Company
Funding: Private Equity Backed
Headquarters: Campbell, United States
Founded: 2003
WebsiteLinkedIn