Senior Site Reliability Engineer - Platform Reliability (Resilience)

Elastic
Spain
Full timeEUR 76,000 - 101,800 annuallyFunction: DevOps, Cloud & InfrastructureSkills: ["Collaboration","Operational excellence","Inclusive communication","Customer-first problem solving","Mentoring"]

Lead SRE technical initiatives to automate system engineering and ensure reliability across Elastic’s global, multi-cloud platform. Build and maintain software, tooling, and automations that scale hosting for services like Elastic Cloud Hosted and Serverless. Respond to and prevent customer-impacting incidents through prioritized problem management using a follow-the-sun on-call model. Collaborate with engineers to improve operational excellence and resilience across the platform.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Elastic
Elastic
1 day ago

Senior Site Reliability Engineer - Platform Reliability (Resilience)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 19 hours agoStatus: Live

Job Summary

Lead SRE technical initiatives to automate system engineering and ensure reliability across Elastic’s global, multi-cloud platform. Build and maintain software, tooling, and automations that scale hosting for services like Elastic Cloud Hosted and Serverless. Respond to and prevent customer-impacting incidents through prioritized problem management using a follow-the-sun on-call model. Collaborate with engineers to improve operational excellence and resilience across the platform.
Location: Spain
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Lead technical initiatives to automate system engineering efforts and guarantee reliability of Elastic’s global infrastructure.
  • •Grow global platform infrastructure by developing and maintaining software, tooling, and automations to meet scaling demands.
  • •Respond to and prevent repeated customer impact through major incident response and prioritized problem management.
  • •Contribute to an inclusive, collaboration-focused environment centered on operational excellence and uplifting others.
  • •Participate in a follow-the-sun on-call rotation (mostly during your working hours).

Pay and Benefits

Salary: EUR 76,000 - 101,800 annually
Perks:Health InsuranceParental LeavePaid Leave

Key Requirements

  • •Background in software engineering to collaborate on identifying, implementing, and delivering solutions.
  • •Experience with public cloud and managed Kubernetes services.
  • •Experience operating SaaS in a public cloud using Infrastructure-as-Code tooling such as Crossplane or Terraform.
  • •Built or operated Kubernetes-at-scale infrastructure across multiple cloud providers, including automation to support it.
  • •Experience with incident management and alerting/metrics systems (e.g., Elastic Stack, Graphite, Prometheus, Influx).
Experience:SaaSPublic cloudKubernetesInfrastructure-as-CodeDistributed systems
Skills:CollaborationOperational excellenceInclusive communicationCustomer-first problem solvingMentoring
Tech Stack:GolangCrossplaneTerraformKubernetesDockerLinuxElastic StackGraphitePrometheusInflux

Company Brief

Elastic
Builds the Elastic Stack (Elasticsearch, Kibana, Beats, Logstash) and provides search, observability, and security solutions that enable organizations to search, analyze, and protect data in real time across applications, infrastructure, and enterprises.
Industry: Data Infrastructure
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Amsterdam, Netherlands
Founded: 2012
WebsiteLinkedIn