Senior Site Reliability Engineer

Airalo
Spain, United Kingdom
Workplace: RemoteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 5+ yearsEducation: bachelorsSkills: ["Communication","Teamwork","Problem-solving","Collaboration","Initiative"]

Lead the design of scalable, fault-tolerant and self-healing systems across multi-region AWS deployments. Define and track SLOs/SLIs, conduct blameless postmortems, and drive automation and observability to minimize manual work. Develop runbooks, improve on-call experience, and collaborate with software engineers to build reliable, cost-efficient systems from the start of the SDLC.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Airalo
Airalo
2 months ago

Senior Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 13 hours agoStatus: Live

Job Summary

Lead the design of scalable, fault-tolerant and self-healing systems across multi-region AWS deployments. Define and track SLOs/SLIs, conduct blameless postmortems, and drive automation and observability to minimize manual work. Develop runbooks, improve on-call experience, and collaborate with software engineers to build reliable, cost-efficient systems from the start of the SDLC.
Location: Spain, United Kingdom
Workplace: Remote
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Manager level

Key Responsibilities

  • •Lead the design of scalable, fault-tolerant and self-healing systems in a multi-region AWS environment.
  • •Define and track Service Level Objectives (SLOs) and Service Level Indicators (SLIs) to drive architectural decisions and error budget policies.
  • •Conduct blameless post-incident reviews to uncover systemic root causes and implement long-term preventive measures.
  • •Identify patterns of manual work and lead the development of internal tools/automation to permanently eliminate them.
  • •Develop and maintain automated runbooks and playbooks for common operational tasks and complex incident response.

Key Requirements

  • •Bachelor’s degree in Computer Engineering or a similar discipline.
  • •5+ years of experience as a Site Reliability Engineer or in a similar role.
  • •3+ years of experience with AWS services including strong knowledge of container orchestration.
  • •2+ years of Kubernetes experience
  • •Proficiency in at least one programming language (Python, Go, Java, etc.) for building automation and internal tooling.
Experience:5+ yearsSRECloudObservabilityAutomation
Education:Bachelor's
Skills:CommunicationTeamworkProblem-solvingCollaborationInitiative
Certifications:AWS Certified DevOps EngineerCKA
Languages:English
Tech Stack:AWSKubernetesPrometheusDatadogOpenTelemetryTerraformGitHub ActionsPythonGoJavaSNSSQS

Company Brief

Airalo
Operates a global eSIM marketplace and mobile connectivity platform that enables travelers and global customers to buy and install eSIMs for data plans across countries without physical SIM cards.
Industry: Telecommunications
Company Size: Medium (51 to 250 employees)
Growth: Growth Stage Startup
Headquarters: Singapore, Singapore
Founded: 2019
WebsiteLinkedIn