Database Reliability Engineer - Core Team

ClickHouse
United Kingdom, Germany, Netherlands
Workplace: RemoteFull timeFunction: Administration & Executive AssistanceExperience: 5+ yearsSkills: ["Problem-solving","Production debugging","Ownership","Communication","Collaboration"]

Build and lead reliability processes for the ClickHouse Core team, improving uptime, scalability, performance, and incident response. Create metrics, alerts, and production-debugging workflows to prevent customer impact, investigate outages with blameless postmortems, and drive continuous improvements. Collaborate across Control Plane, Dataplane, Security, Support, and Operations, and plan Chaos initiatives across Engineering. Operate on-call and manage escalations to resolve issues effectively.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ClickHouse
ClickHouse
5 months ago

Database Reliability Engineer - Core Team

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 5 hours agoStatus: Live

Job Summary

Build and lead reliability processes for the ClickHouse Core team, improving uptime, scalability, performance, and incident response. Create metrics, alerts, and production-debugging workflows to prevent customer impact, investigate outages with blameless postmortems, and drive continuous improvements. Collaborate across Control Plane, Dataplane, Security, Support, and Operations, and plan Chaos initiatives across Engineering. Operate on-call and manage escalations to resolve issues effectively.
Location: United Kingdom, Germany, Netherlands
Workplace: Remote
Employment Type: Full time
Job Function: Administration & Executive Assistance
Seniority: Mid level

Key Responsibilities

  • •Continuously improve reliability and performance of ClickHouse Core.
  • •Create metrics and alerts to identify and prevent production issues before they affect customers.
  • •Investigate common customer problems to identify root causes and submit bug fixes, issue reports, and improvements.
  • •Enhance incident response and post-mortem processes for core outages, coordinating communications with support and cloud teams.
  • •Manage on-call processes, escalation coordination, and drive Chaos initiatives across engineering teams.

Pay and Benefits

Perks:Health InsuranceEquityRemote WorkHome Office

Key Requirements

  • •Bachelor’s or Master’s degree in Computer Science or a related field.
  • •At least 5 years of experience in Reliability Engineering, QA, or customer-facing engineering.
  • •Previous experience operating ClickHouse or other SQL databases in production.
  • •Excellent understanding of distributed database internals and SQL (ClickHouse preferred).
  • •Scripting experience with Shell or Python, and the ability to read and understand C++ code.
Experience:5+ yearsReliability EngineeringQADistributed databasesSQL databasesProduction operations
Education:
Skills:Problem-solvingProduction debuggingOwnershipCommunicationCollaboration
Tech Stack:ClickHouseSQLShellPythonC++AWSAzureGoogle Cloud PlatformGCP

Company Brief

ClickHouse
Develops ClickHouse, a high-performance open-source columnar database for real-time analytics, enabling fast querying and processing of large volumes of data for analytics, monitoring, and business intelligence workloads.
Industry: Data Infrastructure
Headquarters: Menlo Park, United States
Founded: 2016
WebsiteLinkedIn