Senior Site Reliability Engineer- Remote

ClickHouse
Canada
Workplace: RemoteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 8+ yearsEducation: bachelorsSkills: ["Problem-solving","Production debugging","Ownership","Communication","Interpersonal skills"]

Build and lead ClickHouse Cloud reliability practices across Data Plane, Control Plane, Core, Security, Support, and Operations. Own SLOs/SLAs, monitoring and alerting, incident response, and blameless post-mortems with continuous reliability improvements. Develop platforms and tools to optimize operational efficiency using software engineering skills, and drive chaos initiatives. Manage on-call processes, coordination, and escalations to minimize downtime for ClickHouse at scale.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ClickHouse
ClickHouse
6 months ago

Senior Site Reliability Engineer- Remote

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 5 hours agoStatus: Live
Reposted: similar role first listed 1 day ago

Job Summary

Build and lead ClickHouse Cloud reliability practices across Data Plane, Control Plane, Core, Security, Support, and Operations. Own SLOs/SLAs, monitoring and alerting, incident response, and blameless post-mortems with continuous reliability improvements. Develop platforms and tools to optimize operational efficiency using software engineering skills, and drive chaos initiatives. Manage on-call processes, coordination, and escalations to minimize downtime for ClickHouse at scale.
Location: Canada
Workplace: Remote
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Collaborate with engineering teams to design and implement scalable, secure, highly available systems for ClickHouse.
  • •Establish and manage SLOs and SLAs for ClickHouse Cloud.
  • •Ensure monitoring and alerting for infrastructure components to detect and resolve incidents.
  • •Enhance incident response and run blameless post-mortems, coordinating communication with impacted customers.
  • •Manage on-call processes and coordinate escalations to resolve performance and reliability issues and minimize downtime.

Pay and Benefits

Perks:Health InsuranceEquityHome Office

Key Requirements

  • •At least 8 years of experience in Site Reliability Engineering or a related field.
  • •Hands-on experience with Go and/or Python.
  • •Strong knowledge of cloud computing platforms such as AWS, Azure, or Google Cloud Platform.
  • •Hands-on experience with container orchestration tools such as Kubernetes or Docker Swarm.
  • •Strong experience with automation and configuration management tools such as Ansible, Terraform, or Puppet.
Experience:8+ yearsCloudDistributed systemsDistributed databases
Education:Bachelor's
Skills:Problem-solvingProduction debuggingOwnershipCommunicationInterpersonal skills
Tech Stack:GoPythonAWSAzureGoogle Cloud PlatformClickHouseSQLKubernetesDocker SwarmAnsibleTerraformPuppetDistributed databases

Company Brief

ClickHouse
Develops ClickHouse, a high-performance open-source columnar database for real-time analytics, enabling fast querying and processing of large volumes of data for analytics, monitoring, and business intelligence workloads.
Industry: Data Infrastructure
Headquarters: Menlo Park, United States
Founded: 2016
WebsiteLinkedIn