Database Reliability Engineer - Core Team

ClickHouse
EMEA
Workplace: RemoteFull timeFunction: Administration & Executive AssistanceExperience: 5+ yearsEducation: mastersSkills: ["Problem-solving","Production debugging","Ownership","Responsibility","Communication"]

Build and lead reliability engineering processes for ClickHouse Core, improving reliability, availability, scalability, and performance. Create and refine metrics, alerts, and incident response practices, including blameless postmortems and continuous improvement. Partner across Control Plane, Dataplane, Security, Support, and Operations to guide production best practices and drive Chaos initiatives across engineering teams.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ClickHouse
ClickHouse
5 months ago

Database Reliability Engineer - Core Team

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 5 hours agoStatus: Live
Reposted: similar role first listed 1 year ago

Job Summary

Build and lead reliability engineering processes for ClickHouse Core, improving reliability, availability, scalability, and performance. Create and refine metrics, alerts, and incident response practices, including blameless postmortems and continuous improvement. Partner across Control Plane, Dataplane, Security, Support, and Operations to guide production best practices and drive Chaos initiatives across engineering teams.
Location: EMEA
Workplace: Remote
Employment Type: Full time
Job Function: Administration & Executive Assistance
Seniority: Mid level

Key Responsibilities

  • •Continuously improve reliability and performance of ClickHouse Core.
  • •Create and enhance metrics and alerts to identify and prevent production issues before impacting customers.
  • •Investigate recurring customer problems to identify root causes and drive bug fixes, issue reports, and improvements.
  • •Improve incident response and blameless post-mortem analysis for core outages with support and cloud teams.
  • •Manage on-call processes and drive best practices for escalation and response; plan Chaos initiatives across engineering teams.

Pay and Benefits

Equity and Bonus:Equity
Perks:Health InsuranceEquityRemote WorkHome OfficeFlexible Time

Key Requirements

  • •Bachelor’s or Master’s degree in Computer Science or a related field.
  • •At least 5 years of experience in Reliability Engineering, QA, or customer-facing engineering.
  • •Experience operating ClickHouse (or other SQL databases) in production.
  • •Strong distributed database fundamentals and SQL, with ClickHouse in production as a major plus.
  • •Scripting skills (Shell or Python) and ability to read/understand C++ code, plus knowledge of major cloud platforms (AWS, Azure, or GCP).
Experience:5+ yearsReliability engineeringQACustomer-facing engineeringDistributed databasesSQL databasesCloud computing
Education:Master's in Computer Science
Skills:Problem-solvingProduction debuggingOwnershipResponsibilityCommunication
Tech Stack:ClickHouseSQLShellPythonC++AWSAzureGoogle Cloud Platform

Company Brief

ClickHouse
Develops ClickHouse, a high-performance open-source columnar database for real-time analytics, enabling fast querying and processing of large volumes of data for analytics, monitoring, and business intelligence workloads.
Industry: Data Infrastructure
Headquarters: Menlo Park, United States
Founded: 2016
WebsiteLinkedIn