Senior Site Reliability Engineer - Platform Reliability (Resilience)

Elastic
Portugal
Workplace: OnsiteFull timeEUR 62,800 - 84,300Function: DevOps, Cloud & InfrastructureSkills: ["Collaboration","Inclusive communication","Customer-first mindset","Operational excellence","Coaching/mentoring"]

Design, build, and mature a multi-cloud platform that hosts Elastic internal and external services. Lead initiatives to automate system engineering and ensure global infrastructure reliability, while growing platform infrastructure to meet scaling demands through software, tooling, and automations. Respond to and prevent customer impact via major incident response and problem management, participating in a mostly follow-the-sun on-call rotation, and foster an operationally excellent, collaborative culture.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Elastic
Elastic
1 day ago

Senior Site Reliability Engineer - Platform Reliability (Resilience)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 2 hours agoStatus: Live

Job Summary

Design, build, and mature a multi-cloud platform that hosts Elastic internal and external services. Lead initiatives to automate system engineering and ensure global infrastructure reliability, while growing platform infrastructure to meet scaling demands through software, tooling, and automations. Respond to and prevent customer impact via major incident response and problem management, participating in a mostly follow-the-sun on-call rotation, and foster an operationally excellent, collaborative culture.
Location: Portugal
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Lead technical initiatives to automate system engineering efforts to guarantee reliability of the global infrastructure.
  • •Develop and maintain software, tooling, and automations to grow global platform infrastructure and meet scaling demands.
  • •Champion an environment focused on collaboration, operational excellence, and uplifting others.
  • •Respond to and prevent repeated customer impact via major incident response and prioritized problem management.
  • •Participate in a mostly follow-the-sun on-call rotation for incident handling.

Pay and Benefits

Salary: EUR 62,800 - 84,300
Perks:Health InsurancePaid LeaveParental LeaveVolunteer TimeDonation Match

Key Requirements

  • •Take an SRE approach to solve operational problems with a customer-first mindset and continuous progress focus.
  • •Background in software engineering to identify, implement, and deliver reliability solutions in collaboration with engineers.
  • •Experience with public cloud and managed Kubernetes services (advantage).
  • •Experience writing non-trivial programs in Golang or other programming languages.
  • •System administration skills in Linux on distributed systems at scale.
Experience:SaaSPublic cloudKubernetesMulti-cloudDistributed systemsInfrastructure-as-CodeObservability
Skills:CollaborationInclusive communicationCustomer-first mindsetOperational excellenceCoaching/mentoring
Tech Stack:GolangCrossplaneTerraformKubernetesDockerLinuxElastic StackGraphitePrometheusInflux

Company Brief

Elastic
Builds the Elastic Stack (Elasticsearch, Kibana, Beats, Logstash) and provides search, observability, and security solutions that enable organizations to search, analyze, and protect data in real time across applications, infrastructure, and enterprises.
Industry: Data Infrastructure
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Amsterdam, Netherlands
Founded: 2012
WebsiteLinkedIn