Cloud Site Reliability Engineer II

Barracuda Networks
Ann Arbor
Workplace: HybridFull timeCAD 100,000 - 120,000 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 2-4 yearsSkills: ["Analytical troubleshooting","Communication","Collaboration"]

Build and run centralized observability for a multi-tenant Kubernetes platform, delivering reliable telemetry, alerting, and actionable Grafana dashboards. You’ll automate log/metrics/tracing pipelines across AWS EKS and Azure AKS using GitOps and Infrastructure as Code, and improve Kubernetes platform health through performance tuning and modernization. Collaborate with product and tenant teams on observability onboarding and distributed tracing, leveraging AI coding tools to accelerate operational diagnostics and automation.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Barracuda Networks
Barracuda Networks
4 days ago

Cloud Site Reliability Engineer II

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 13 hours agoStatus: Live
Reposted: similar role first listed 1 day ago

Job Summary

Build and run centralized observability for a multi-tenant Kubernetes platform, delivering reliable telemetry, alerting, and actionable Grafana dashboards. You’ll automate log/metrics/tracing pipelines across AWS EKS and Azure AKS using GitOps and Infrastructure as Code, and improve Kubernetes platform health through performance tuning and modernization. Collaborate with product and tenant teams on observability onboarding and distributed tracing, leveraging AI coding tools to accelerate operational diagnostics and automation.
Location: Ann Arbor
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Operate, scale, and automate centralized LGTM observability infrastructure (Loki, Mimir/Prometheus, Tempo, Grafana).
  • •Design high-impact Grafana dashboards and executive health overviews for platform services, Kubernetes clusters, and tenant workloads.
  • •Implement alerting strategies, SLO/SLI tracking (Sloth), and notification routing to detect and resolve degradation.
  • •Automate deployment of log collectors, metric exporters, and monitoring agents across multi-cluster EKS/AKS using GitOps (ArgoCD) and Terragrunt.
  • •Collaborate with product and tenant teams on observability onboarding, distributed tracing instrumentation, and performance troubleshooting.

Pay and Benefits

Salary: CAD 100,000 - 120,000 annually
Equity and Bonus:Equity
Perks:Health InsuranceRetirementPaid LeaveVolunteering

Key Requirements

  • •2–4 years of experience with public cloud infrastructure (AWS and/or Azure) focused on observability, monitoring, and systems reliability.
  • •1–2+ years deploying, operating, or troubleshooting containerized workloads in Kubernetes (EKS/AKS).
  • •Experience configuring and operating observability dashboards and systems (Grafana, ELK, Splunk, etc.).
  • •Working knowledge of Infrastructure as Code with Terraform and/or Terragrunt and GitOps delivery workflows (ArgoCD or Flux).
  • •Strong scripting skills in Python or Bash for automation and telemetry pipelines (Go is a plus).
Experience:2-4 yearsPublic cloudKubernetesObservabilityGitOps
Skills:Analytical troubleshootingCommunicationCollaboration
Tech Stack:GrafanaPrometheusMimirLokiTempoOpenTelemetryGrafana AlloySlothAlertmanagerKubernetesAWS EKSAzure AKSHelmKustomizeTerragruntTerraformArgoCDGitHub ActionsAWSMicrosoft Azure

Company Brief

Barracuda Networks
Provides cloud-enabled security and data protection solutions including email protection, network and application security, and backup and recovery services for businesses, service providers, and government organizations worldwide.
Industry: Cybersecurity
Company Size: Enterprise (1,001+ employees)
Revenue: USD 250M to 500M
Growth: Established Company
Funding: Private Equity Backed
Headquarters: Campbell, United States
Founded: 2003
WebsiteLinkedIn