Senior Site Reliability Engineer - Platform Reliability (Resilience)

Elastic
United Kingdom
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureSkills: ["Collaboration","Customer-first mindset","Inclusive communication","Operational excellence","Coaching & mentoring"]

Design, build, and mature a multi-cloud platform for hosting Elastic internal and external services, enabling rapid deployment across the company. Lead technical initiatives to automate system engineering efforts that guarantee reliability, grow global platform infrastructure for scaling, and prevent repeated customer impact through incident response and problem management. Partner with engineers to deliver resilient solutions and foster an inclusive, collaborative environment using a follow-the-sun on-call model.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Elastic
Elastic
1 day ago

Senior Site Reliability Engineer - Platform Reliability (Resilience)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 20 hours agoStatus: Live

Job Summary

Design, build, and mature a multi-cloud platform for hosting Elastic internal and external services, enabling rapid deployment across the company. Lead technical initiatives to automate system engineering efforts that guarantee reliability, grow global platform infrastructure for scaling, and prevent repeated customer impact through incident response and problem management. Partner with engineers to deliver resilient solutions and foster an inclusive, collaborative environment using a follow-the-sun on-call model.
Location: United Kingdom
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Lead technical initiatives to automate system engineering efforts that guarantee reliability for global infrastructure.
  • •Grow global platform infrastructure by developing and maintaining software, tooling, and automations for scaling.
  • •Respond to and prevent repeated customer impact through major incident response and prioritized problem management.
  • •Use a follow-the-sun on-call model participating in on-call during mostly your working hours.
  • •Champion an inclusive environment focused on collaboration and operational excellence while uplifting others.

Pay and Benefits

Perks:Health InsuranceParental Leave

Key Requirements

  • •Background in software engineering to collaborate with engineers and deliver operational solutions with an SRE perspective.
  • •Experience with public cloud and managed Kubernetes services (advantageous).
  • •Experience operating SaaS in public cloud using Infrastructure-as-Code tools such as Crossplane or Terraform (bonus).
  • •Built or operated Kubernetes-at-scale infrastructure across multiple cloud providers and the automation supporting it (bonus).
  • •Hands-on Linux system administration skills on distributed systems at scale (plus demonstrated incident/alerting and major incident management experience, e.g., Prometheus/Graphite/Influx).
Experience:SaaSPublic cloudKubernetesInfrastructure as codeDistributed systems
Skills:CollaborationCustomer-first mindsetInclusive communicationOperational excellenceCoaching & mentoring
Languages:English
Tech Stack:KubernetesCrossplaneTerraformGolangDockerLinuxElastic StackGraphitePrometheusInfluxInfrastructure-as-CodeElastic Cloud HostedServerless

Company Brief

Elastic
Builds the Elastic Stack (Elasticsearch, Kibana, Beats, Logstash) and provides search, observability, and security solutions that enable organizations to search, analyze, and protect data in real time across applications, infrastructure, and enterprises.
Industry: Data Infrastructure
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Amsterdam, Netherlands
Founded: 2012
WebsiteLinkedIn