Senior Site Reliability Engineer- EMEA(Remote)

ClickHouse
Germany, EMEA, Canada, United States
Workplace: RemoteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 8+ yearsEducation: bachelorsSkills: ["Problem-solving","Production debugging","Ownership","Accountability","Communication"]

Build and lead reliability processes for ClickHouse Cloud, ensuring reliability, availability, scalability, and performance across its distributed infrastructure. Partner with Engineering teams to design highly available systems, define SLOs/SLAs, and implement monitoring/alerting. Own incident management, blameless postmortems, chaos initiatives, and continuous reliability improvements. Use software engineering to develop platforms and tools that increase operational and engineering efficiency at ClickHouse Cloud.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ClickHouse
ClickHouse
1 month ago

Senior Site Reliability Engineer- EMEA(Remote)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 4 days agoStatus: Live

Job Summary

Build and lead reliability processes for ClickHouse Cloud, ensuring reliability, availability, scalability, and performance across its distributed infrastructure. Partner with Engineering teams to design highly available systems, define SLOs/SLAs, and implement monitoring/alerting. Own incident management, blameless postmortems, chaos initiatives, and continuous reliability improvements. Use software engineering to develop platforms and tools that increase operational and engineering efficiency at ClickHouse Cloud.
Location: Germany, EMEA, Canada, United States
Workplace: Remote
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Collaborate with engineering teams to design and implement scalable, secure, highly available systems for ClickHouse.
  • •Establish and manage SLOs and SLAs for ClickHouse Cloud.
  • •Ensure monitoring and alerting are in place across infrastructure components to detect and resolve incidents.
  • •Enhance incident response and run blameless postmortem analysis for outages, including communicating with impacted customers.
  • •Manage on-call processes, drive chaos initiatives, and continuously improve reliability and performance.

Pay and Benefits

Perks:Remote WorkHealth InsuranceEquityTime OffHome Office

Key Requirements

  • •Bachelor’s or Master’s degree in Computer Science or a related field.
  • •At least 8 years of experience in Site Reliability Engineering or a related field.
  • •Previous experience using ClickHouse in production.
  • •Hands-on experience with Go and/or Python.
  • •Strong knowledge of cloud computing platforms such as AWS, Azure, or Google Cloud Platform.
Experience:8+ yearsCloud infrastructureDistributed systemsObservabilityData warehousingReal-time analytics
Education:Bachelor's in Computer Science or related field
Skills:Problem-solvingProduction debuggingOwnershipAccountabilityCommunication
Languages:English
Tech Stack:GoPythonAWSAzureGoogle Cloud PlatformSQLClickHouseKubernetesDocker SwarmAnsibleTerraformPuppetDistributed databasesServerless

Company Brief

ClickHouse
Develops ClickHouse, a high-performance open-source columnar database for real-time analytics, enabling fast querying and processing of large volumes of data for analytics, monitoring, and business intelligence workloads.
Industry: Data Infrastructure
Headquarters: Menlo Park, United States
Founded: 2016
WebsiteLinkedIn