Site Reliability Engineer

Sporty Group
EMEA
Workplace: RemoteFull timeFunction: DevOps, Cloud & InfrastructureSkills: ["Incident response","Troubleshooting","Mentorship","Root cause analysis","Leading post-mortems"]

Own the reliability of cloud and Kubernetes platforms by improving infrastructure, streamlining provisioning with GitOps, and designing alerting frameworks that reduce noise. Monitor and maintain AWS-based systems with autoscaling and observability across metrics, logs, traces, and RUM. Lead incident response (on-call, triage, root cause, post-mortems), define SLIs/SLOs, and support rapid deployments to new countries while mentoring teammates.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Sporty Group
Sporty Group
2 months ago

Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 6 hours agoStatus: Live

Job Summary

Own the reliability of cloud and Kubernetes platforms by improving infrastructure, streamlining provisioning with GitOps, and designing alerting frameworks that reduce noise. Monitor and maintain AWS-based systems with autoscaling and observability across metrics, logs, traces, and RUM. Lead incident response (on-call, triage, root cause, post-mortems), define SLIs/SLOs, and support rapid deployments to new countries while mentoring teammates.
Location: EMEA
Workplace: Remote
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Improve infrastructure and deployment processes across deployed countries and enable streamlining for future country launches.
  • •Continuously enhance Kubernetes platform stability and efficiency, optimizing resource utilisation and cost with GitOps-first provisioning.
  • •Monitor and maintain cloud infrastructure using autoscaling, alerting pipelines, and observability dashboards across metrics, logs, traces, and RUM.
  • •Own weekend on-call operations: triage and respond to production incidents, perform root cause analysis, and drive post-incident reviews.
  • •Define and maintain SLIs and SLOs, and design alert pipelines/frameworks to prevent alert fatigue and avoid noisy alerting patterns.

Pay and Benefits

Perks:Remote WorkPaid Leave

Key Requirements

  • •3+ years of DevOps/SRE/platform engineering experience.
  • •Must be based in Europe.
  • •Independently lead planning and deployment of projects.
  • •Strong AWS experience and solid Kubernetes/container orchestration knowledge (EKS; GitOps with ArgoCD and Helm valued).
  • •Proficiency with Infrastructure-as-Code (Terraform) and automation scripting (Bash, Python, or Golang).
Skills:Incident responseTroubleshootingMentorshipRoot cause analysisLeading post-mortems
Languages:English
Tech Stack:AWSKubernetesEKSGitOpsArgoCDHelmTerraformBashPythonGolangRustPrometheusLokiTempoPyroscopeOpenTelemetryGrafana FaroOpenTelemetry SDKGrafanaAlertmanager

Company Brief

Sporty Group
Operates an online retail platform offering sports apparel, equipment, and accessories across multiple sports categories, targeting consumers and amateur athletes with a range of brands and direct-to-consumer products.
Industry: E-commerce
Website