Senior Site Reliability Engineer- Remote

ClickHouse
Singapore
Workplace: RemoteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 8+ yearsEducation: mastersSkills: ["Problem-solving","Production debugging","Ownership","Accountability","Communication"]

Build and lead ClickHouse’s Site Reliability Engineering practices to ensure reliability, availability, scalability, and performance of ClickHouse Cloud. Partner with engineering teams to design fault-tolerant distributed systems, set SLOs/SLAs, and implement monitoring/alerting. Own incident management, blameless postmortems, and continuous reliability improvements, including chaos initiatives and on-call best practices. Develop internal platforms and tools using Go/Python to improve operational efficiency.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ClickHouse
ClickHouse
2 months ago

Senior Site Reliability Engineer- Remote

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 5 hours agoStatus: Live

Job Summary

Build and lead ClickHouse’s Site Reliability Engineering practices to ensure reliability, availability, scalability, and performance of ClickHouse Cloud. Partner with engineering teams to design fault-tolerant distributed systems, set SLOs/SLAs, and implement monitoring/alerting. Own incident management, blameless postmortems, and continuous reliability improvements, including chaos initiatives and on-call best practices. Develop internal platforms and tools using Go/Python to improve operational efficiency.
Location: Singapore
Workplace: Remote
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Collaborate with engineering teams to design and implement scalable, secure, and highly available systems for ClickHouse.
  • •Establish and manage SLOs and SLAs for ClickHouse Cloud.
  • •Ensure monitoring and alerting are in place across infrastructure components (Dataplane, Control Plane, and ClickHouse Core).
  • •Own incident management and response, including blameless post-mortems and customer communication through Support.
  • •Plan, enable, and drive Chaos initiatives and manage on-call processes with escalation best practices to minimize downtime.

Pay and Benefits

Equity and Bonus:Equity
Perks:Health InsuranceEquityHome Office

Key Requirements

  • •At least 8 years of experience in Site Reliability Engineering or a related field.
  • •Bachelor’s or Master’s degree in Computer Science or a related field.
  • •Hands-on experience with Go and/or Python.
  • •Strong knowledge of cloud computing platforms such as AWS, Azure, or Google Cloud Platform.
  • •Hands-on experience with container orchestration tools such as Kubernetes or Docker Swarm, and automation/configuration tools such as Ansible, Terraform, or Puppet.
Experience:8+ yearsSite reliability engineeringCloud computingDistributed systemsDistributed databasesServerless
Education:Master's in Computer Science or a related field
Skills:Problem-solvingProduction debuggingOwnershipAccountabilityCommunication
Tech Stack:GoPythonAWSAzureGoogle Cloud PlatformClickHouseSQLKubernetesDocker SwarmAnsibleTerraformPuppetDistributed databasesServerless

Eligibility

Work Authorization:Authorization required. Sponsorship not provided.

Company Brief

ClickHouse
Develops ClickHouse, a high-performance open-source columnar database for real-time analytics, enabling fast querying and processing of large volumes of data for analytics, monitoring, and business intelligence workloads.
Industry: Data Infrastructure
Headquarters: Menlo Park, United States
Founded: 2016
WebsiteLinkedIn