Senior Site Reliability Engineer - (US-Remote)

ClickHouse
United States
Workplace: RemoteFull timeUSD 150,000 - 230,000 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 8+ yearsEducation: bachelorsSkills: ["Problem-solving","Communication","Interpersonal skills","Ownership","Accountability"]

Join ClickHouse to build and lead its newly formed Site Reliability Engineering team. Own reliability for ClickHouse Cloud’s distributed infrastructure—driving monitoring, SLOs/SLAs, incident management, and blameless post-mortems. Collaborate across Control Plane, Dataplane, Core, Security, Support, and Operations to deliver scalable, fault-tolerant systems. Use Go/Python and automation to improve operational efficiency, and plan Chaos initiatives to raise resilience.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ClickHouse
ClickHouse
1 day ago

Senior Site Reliability Engineer - (US-Remote)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 14 hours agoStatus: Live

Job Summary

Join ClickHouse to build and lead its newly formed Site Reliability Engineering team. Own reliability for ClickHouse Cloud’s distributed infrastructure—driving monitoring, SLOs/SLAs, incident management, and blameless post-mortems. Collaborate across Control Plane, Dataplane, Core, Security, Support, and Operations to deliver scalable, fault-tolerant systems. Use Go/Python and automation to improve operational efficiency, and plan Chaos initiatives to raise resilience.
Location: United States
Workplace: Remote
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Collaborate with engineering teams to design and implement scalable, secure, highly available systems for ClickHouse.
  • •Establish and manage SLOs and SLAs for ClickHouse Cloud.
  • •Ensure monitoring and alerting are in place across infrastructure components to detect and resolve incidents.
  • •Enhance incident response and run blameless post-mortems, coordinating communications with impacted customers.
  • •Plan Chaos initiatives, manage on-call processes, and continuously improve reliability and performance.

Pay and Benefits

Salary: USD 150,000 - 230,000 annually
Perks:Health InsuranceEquityRemote WorkHome Office

Key Requirements

  • •Bachelor’s or Master’s degree in Computer Science or a related field.
  • •At least 8 years of experience in Site Reliability Engineering or a related field.
  • •Previous experience using ClickHouse in production.
  • •Hands-on experience with Go and/or Python.
  • •Strong knowledge of cloud platforms such as AWS, Azure, or Google Cloud Platform.
Experience:8+ years
Education:Bachelor's in Computer Science
Skills:Problem-solvingCommunicationInterpersonal skillsOwnershipAccountability
Tech Stack:GoPythonAWSAzureGoogle Cloud PlatformSQLClickHouseKubernetesDocker SwarmAnsibleTerraformPuppet

Company Brief

ClickHouse
Develops ClickHouse, a high-performance open-source columnar database for real-time analytics, enabling fast querying and processing of large volumes of data for analytics, monitoring, and business intelligence workloads.
Industry: Data Infrastructure
Headquarters: Menlo Park, United States
Founded: 2016
WebsiteLinkedIn