Senior Site Reliability Engineer- Remote

ClickHouse
Canada
Workplace: RemoteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 8+ yearsEducation: bachelorsSkills: ["Communication","Problem-solving","Ownership","Teamwork","Analytical thinking"]

Join ClickHouse’s central Site Reliability Engineering team to design, build, and operate scalable, secure, and highly available cloud infrastructure. You’ll own incident management, SLOs/SLAs, on-call processes, and drive improvements across Data Plane, Control Plane, and Core services while collaborating with multiple engineering teams to ensure reliability and performance at elastic scale.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ClickHouse
ClickHouse
4 months ago

Senior Site Reliability Engineer- Remote

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 20 hours agoStatus: Live

Job Summary

Join ClickHouse’s central Site Reliability Engineering team to design, build, and operate scalable, secure, and highly available cloud infrastructure. You’ll own incident management, SLOs/SLAs, on-call processes, and drive improvements across Data Plane, Control Plane, and Core services while collaborating with multiple engineering teams to ensure reliability and performance at elastic scale.
Location: Canada
Workplace: Remote
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure

Key Responsibilities

  • •Collaborate with various engineering teams to design and implement scalable, secure, and highly available systems for ClickHouse Cloud.
  • •Establish and manage service level objectives (SLOs) and service level agreements (SLAs) for ClickHouse Cloud.
  • •Ensure all infrastructure components have monitoring and alerting to enable timely incident detection and resolution.
  • •Enhance incident response processes and post-mortem analysis, including communicating with customers via the support team.
  • •Plan, enable, and drive chaos initiatives, manage on-call processes, and coordinate escalation to minimize downtime.

Pay and Benefits

Perks:Remote WorkHealth InsuranceEquityTime OffHome Office

Key Requirements

  • •Bachelor’s or Master’s degree in Computer Science or a related field; at least 8 years of experience in Site Reliability Engineering or a related field.
  • •Hands-on experience with Go and/or Python.
  • •Strong knowledge of cloud computing platforms such as AWS, Azure, or Google Cloud Platform.
  • •Excellent understanding of distributed databases and SQL, with ClickHouse experience as a major plus.
  • •Hands-on experience with container orchestration tools such as Kubernetes or Docker Swarm.
Experience:8+ years
Education:Bachelor's
Skills:CommunicationProblem-solvingOwnershipTeamworkAnalytical thinking
Languages:English
Tech Stack:GoPythonAWSAzureGCPKubernetesDockerAnsibleTerraformPuppetSQLClickHouse

Company Brief

ClickHouse
Develops ClickHouse, a high-performance open-source columnar database for real-time analytics, enabling fast querying and processing of large volumes of data for analytics, monitoring, and business intelligence workloads.
Industry: Data Infrastructure
Headquarters: Menlo Park, United States
Founded: 2016
WebsiteLinkedIn