Database Reliability Engineer - Core Team

ClickHouse
Netherlands, United Kingdom, United States, Germany
Workplace: RemoteFull timeFunction: Administration & Executive AssistanceExperience: 5+ yearsEducation: bachelorsSkills: ["Problem-solving","Production debugging","Ownership","Accountability","Communication"]

Build and lead reliability engineering processes for ClickHouse Core to improve reliability, availability, scalability, and performance. Create and refine metrics, alerts, and incident response workflows, including blameless postmortems and continuous improvement. Investigate customer-impacting issues, drive Chaos initiatives, manage on-call and escalation practices, and collaborate across control plane, dataplane, security, support, and cloud teams to optimize ClickHouse Cloud.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ClickHouse
ClickHouse
5 months ago

Database Reliability Engineer - Core Team

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 5 hours agoStatus: Live

Job Summary

Build and lead reliability engineering processes for ClickHouse Core to improve reliability, availability, scalability, and performance. Create and refine metrics, alerts, and incident response workflows, including blameless postmortems and continuous improvement. Investigate customer-impacting issues, drive Chaos initiatives, manage on-call and escalation practices, and collaborate across control plane, dataplane, security, support, and cloud teams to optimize ClickHouse Cloud.
Location: Netherlands, United Kingdom, United States, Germany
Workplace: Remote
Employment Type: Full time
Job Function: Administration & Executive Assistance
Seniority: Mid level

Key Responsibilities

  • •Continuously improve the reliability and performance of ClickHouse core.
  • •Improve and create metrics and alerts to identify and prevent production problems.
  • •Investigate common customer issues to identify root causes and submit bug fixes, issue reports, and improvements.
  • •Enhance incident response processes and post-mortem analysis for outages, coordinating with support and cloud teams.
  • •Plan, enable, and drive Chaos initiatives and manage on-call processes and escalation best practices.

Pay and Benefits

Equity and Bonus:Equity
Perks:Health InsuranceEquityHome Office

Key Requirements

  • •Bachelor’s or Master’s degree in Computer Science or a related field.
  • •At least 5 years of experience in Reliability Engineering, QA, or customer-facing engineering.
  • •Previous experience operating ClickHouse or other SQL databases in production.
  • •Excellent understanding of distributed database internals and SQL (ClickHouse is a plus).
  • •Scripting experience with Shell or Python, plus ability to read and understand C++ code.
Experience:5+ yearsDistributed databasesReliability engineeringQACloud computing
Education:Bachelor's in Computer Science (or related field)
Skills:Problem-solvingProduction debuggingOwnershipAccountabilityCommunication
Tech Stack:ClickHouseSQLShellPythonC++AWSAzureGoogle Cloud Platform

Company Brief

ClickHouse
Develops ClickHouse, a high-performance open-source columnar database for real-time analytics, enabling fast querying and processing of large volumes of data for analytics, monitoring, and business intelligence workloads.
Industry: Data Infrastructure
Headquarters: Menlo Park, United States
Founded: 2016
WebsiteLinkedIn