Weekend Site Reliability Engineer

Sporty Group
Anywhere
Workplace: RemoteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 3+ yearsSkills: ["Incident response","Triaging","Root cause analysis","Troubleshooting","Mentoring"]

Work as a Weekend SRE covering Saturday through Monday, improving infrastructure stability and processes across multiple countries. You’ll enhance the Kubernetes platform using GitOps-first practices, own weekend on-call operations (triage, RCA, and post-incident reviews), and build alert pipelines with low-noise signal. Define SLIs/SLOs to drive reliability improvements, collaborate on architecture changes for rapid country rollouts, and mentor less experienced engineers.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Sporty Group
Sporty Group
7 months ago

Weekend Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 13 hours agoStatus: Live

Job Summary

Work as a Weekend SRE covering Saturday through Monday, improving infrastructure stability and processes across multiple countries. You’ll enhance the Kubernetes platform using GitOps-first practices, own weekend on-call operations (triage, RCA, and post-incident reviews), and build alert pipelines with low-noise signal. Define SLIs/SLOs to drive reliability improvements, collaborate on architecture changes for rapid country rollouts, and mentor less experienced engineers.
Location: Anywhere
Workplace: Remote
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Cover weekend on-call operations: triage and respond to production incidents, conduct root cause analysis, and drive post-incident reviews.
  • •Improve Kubernetes platform stability and efficiency to optimize resource utilization, reduce costs, and streamline environment provisioning with GitOps-first practices.
  • •Monitor and maintain cloud infrastructure using autoscaling, alerting pipelines, and Grafana dashboards (metrics, logs, traces, and RUM).
  • •Design and maintain alert pipelines to ensure actionable signal quality and prevent alert fatigue and notification flooding.
  • •Define and maintain SLIs and SLOs, support rapid architecture reconfiguration for new country deployments, and liaise on security audits and internal security sweeps.

Pay and Benefits

Perks:Paid LeaveAnnual Bonus

Key Requirements

  • •3+ years of DevOps / platform engineering experience.
  • •Based in Europe, Asia, or LatAm.
  • •Experience independently leading planning and deployment of a project.
  • •Strong Kubernetes and container orchestration experience, including EKS and GitOps tooling (ArgoCD and Helm valued).
  • •Hands-on experience with Infrastructure-as-Code, especially Terraform, plus observability and on-call incident response experience.
Experience:3+ years
Skills:Incident responseTriagingRoot cause analysisTroubleshootingMentoring
Tech Stack:AWSKubernetesEKSGitOpsArgoCDHelmTerraformBashPythonGolangRustPrometheusLokiTempoPyroscopeOpenTelemetryGrafana FaroGrafanaAlertmanagerOpenTelemetry SDK

Company Brief

Sporty Group
Operates an online retail platform offering sports apparel, equipment, and accessories across multiple sports categories, targeting consumers and amateur athletes with a range of brands and direct-to-consumer products.
Industry: E-commerce
Website