Senior Site Reliability Engineer- Remote

ClickHouse
Australia
Workplace: RemoteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 8+ yearsEducation: bachelorsSkills: ["Software engineering","Infrastructure","Automation","Monitoring","Cloud platforms","Container orchestration","Post-mortem analysis","On-call management"]

Lead the reliability and performance of ClickHouse Cloud by building scalable, secure, and highly available distributed systems. Collaborate with Control Plane, Data Plane, Core, Security, Support and Operations; own incident management, post-mortems, and continuous improvement while leveraging Go/Python, Kubernetes, and major cloud platforms to optimize operational efficiency in a remote, globally distributed team.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ClickHouse
ClickHouse
4 months ago

Senior Site Reliability Engineer- Remote

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 21 hours agoStatus: Live

Job Summary

Lead the reliability and performance of ClickHouse Cloud by building scalable, secure, and highly available distributed systems. Collaborate with Control Plane, Data Plane, Core, Security, Support and Operations; own incident management, post-mortems, and continuous improvement while leveraging Go/Python, Kubernetes, and major cloud platforms to optimize operational efficiency in a remote, globally distributed team.
Location: Australia
Workplace: Remote
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure

Key Responsibilities

  • •Collaborate with various engineering teams to design and implement scalable, secure, and highly available systems for ClickHouse.
  • •Establish and manage service level objectives (SLOs) and service level agreements (SLAs) for ClickHouse Cloud.
  • •Ensure all infrastructure components have monitoring and alerting in place to ensure timely detection and resolution of incidents.
  • •Enhance and refine incident response processes and post-mortem analysis for outages, including communicating with impacted customers.
  • •Continuously improve the reliability and performance of ClickHouse services and drive Chaos initiatives across Engineering teams.

Pay and Benefits

Equity and Bonus:Equity
Perks:Remote WorkHealth InsuranceEquityTime OffHome OfficeGlobal Gatherings

Key Requirements

  • •Bachelor’s or Master’s degree in Computer Science or a related field.
  • •8+ years of experience in Site Reliability Engineering or a related field.
  • •Hands-on experience with Go and/or Python.
  • •Strong knowledge of cloud computing platforms such as AWS, Azure, or Google Cloud Platform.
  • •Hands-on experience with container orchestration tools such as Kubernetes or Docker Swarm; strong experience with automation and configuration management tools such as Ansible, Terraform, or Puppet.
Experience:8+ yearsCloud computingDistributed systemsSite reliability engineering
Education:Bachelor's
Skills:Software engineeringInfrastructureAutomationMonitoringCloud platformsContainer orchestrationPost-mortem analysisOn-call management
Languages:English
Tech Stack:GoPythonAWSAzureGoogle Cloud PlatformKubernetesDockerAnsibleTerraformPuppet

Company Brief

ClickHouse
Develops ClickHouse, a high-performance open-source columnar database for real-time analytics, enabling fast querying and processing of large volumes of data for analytics, monitoring, and business intelligence workloads.
Industry: Data Infrastructure
Headquarters: Menlo Park, United States
Founded: 2016
WebsiteLinkedIn