Senior Site Reliability Engineer- Remote

ClickHouse
United Kingdom
Workplace: RemoteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 8+ yearsEducation: bachelorsSkills: ["Problem-solving","Production debugging","Ownership","Accountability","Communication","Interpersonal skills"]

Build and lead processes that keep ClickHouse Cloud reliable, available, scalable, and fast. Partner with engineering teams to design and implement secure, fault-tolerant distributed systems, set SLOs/SLAs, and ensure monitoring and alerting across cloud infrastructure. Own incident management, blameless postmortems, and continuous reliability improvements, including on-call and chaos initiatives. Use Go/Python and automation tooling to create operational platforms and tools that improve efficiency at ClickHouse’s high scale.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ClickHouse
ClickHouse
6 months ago

Senior Site Reliability Engineer- Remote

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 5 hours agoStatus: Live

Job Summary

Build and lead processes that keep ClickHouse Cloud reliable, available, scalable, and fast. Partner with engineering teams to design and implement secure, fault-tolerant distributed systems, set SLOs/SLAs, and ensure monitoring and alerting across cloud infrastructure. Own incident management, blameless postmortems, and continuous reliability improvements, including on-call and chaos initiatives. Use Go/Python and automation tooling to create operational platforms and tools that improve efficiency at ClickHouse’s high scale.
Location: United Kingdom
Workplace: Remote
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Collaborate with engineering teams to design and implement scalable, secure, highly available systems for ClickHouse.
  • •Establish and manage SLOs and SLAs for ClickHouse Cloud.
  • •Ensure monitoring and alerting are in place across ClickHouse Cloud infrastructure components to detect and resolve incidents quickly.
  • •Enhance incident response and run blameless post-mortems, including communication with impacted customers.
  • •Plan and drive chaos initiatives, manage on-call processes, and continuously improve reliability and performance.

Pay and Benefits

Equity and Bonus:Equity
Perks:Health InsuranceEquityHome Office

Key Requirements

  • •At least 8 years of experience in Site Reliability Engineering or a related field.
  • •Hands-on experience with Go and/or Python.
  • •Strong knowledge of cloud computing platforms such as AWS, Azure, or Google Cloud Platform.
  • •Hands-on experience with distributed databases and SQL, with ClickHouse experience a major plus.
  • •Hands-on experience with container orchestration (Kubernetes or Docker Swarm) and automation/config management tools (Ansible, Terraform, or Puppet).
Experience:8+ yearsCloud infrastructureDistributed systemsDistributed databasesAutomationIncident management
Education:Bachelor's in Computer Science or a related field
Skills:Problem-solvingProduction debuggingOwnershipAccountabilityCommunicationInterpersonal skills
Tech Stack:GoPythonAWSAzureGoogle Cloud PlatformClickHouseSQLKubernetesDocker SwarmAnsibleTerraformPuppet

Company Brief

ClickHouse
Develops ClickHouse, a high-performance open-source columnar database for real-time analytics, enabling fast querying and processing of large volumes of data for analytics, monitoring, and business intelligence workloads.
Industry: Data Infrastructure
Headquarters: Menlo Park, United States
Founded: 2016
WebsiteLinkedIn