Site Reliability Engineer - ClickHouse

PostHog
EU, United Kingdom
Workplace: RemoteFull timeFunction: DevOps, Cloud & InfrastructureSkills: ["Proactive ownership","Problem-solving","Collaboration","Low-ego leadership","Respectful communication"]

Own reliability for PostHog’s large, self-managed ClickHouse deployments on AWS at petabyte scale. Help turn a fast-growing stateful system into a predictable, well-automated platform—covering provisioning, scaling, rebalancing, recovery, and incident response. Build and improve operational tooling for deploys, schema changes, backups, restores, and on-call, while working with ClickHouse engineers to translate database needs into infrastructure solutions.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
PostHog
PostHog
3 weeks ago

Site Reliability Engineer - ClickHouse

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 3 hours agoStatus: Live

Job Summary

Own reliability for PostHog’s large, self-managed ClickHouse deployments on AWS at petabyte scale. Help turn a fast-growing stateful system into a predictable, well-automated platform—covering provisioning, scaling, rebalancing, recovery, and incident response. Build and improve operational tooling for deploys, schema changes, backups, restores, and on-call, while working with ClickHouse engineers to translate database needs into infrastructure solutions.
Location: EU, United Kingdom
Workplace: Remote
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure

Key Responsibilities

  • •Manage large fleets of EC2-based VMs, disks, and networking for data-intensive workloads.
  • •Improve operational tooling for deploys, schema changes, backups, restores, and incident response.
  • •Work with ClickHouse engineers to turn database-level needs into infra-level solutions.
  • •Reduce operational load by identifying repeat pain points and eliminating them via code and self-healing automation.
  • •Participate in on-call and incident response, focusing on making incidents rarer over time.

Key Requirements

  • •Prior experience with ClickHouse or other OLAP databases.
  • •Strong experience operating production infrastructure on AWS.
  • •Hands-on experience with VM-based systems (EC2), not just managed PaaS.
  • •Experience automating infrastructure using tools like Terraform, Ansible, or similar.
  • •Solid understanding of Linux systems and experience supporting stateful systems (databases, queues, storage).
Skills:Proactive ownershipProblem-solvingCollaborationLow-ego leadershipRespectful communication
Tech Stack:ClickHouseAWSEC2VMsTerraformAnsibleLinux

Company Brief

PostHog
PostHog provides an open-source product analytics and developer platform that combines analytics, session replay, feature flags, A/B testing, and data warehousing to help engineering and product teams build and optimize products.
Industry: Developer Tools
Company Size: Medium (51 to 250 employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series E+
Headquarters: San Francisco, United States
Founded: 2020
Glassdoor
Glassdoor: 4.3
WebsiteLinkedInGlassdoor