Senior Principal Site Reliability Engineer

Bybit
Kuala Lumpur
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 8+ yearsSkills: ["Technical documentation","Solution design","Cross-team collaboration","Risk control","Independent planning"]

Design and build an enterprise chaos engineering platform for multi-cluster, multi-region and multi-environment Kubernetes/EC2 deployments. Own production safety controls such as blast radius limiting, kill switch, rollback and real-time impact monitoring, then orchestrate and execute routine resilience validation experiments with closed-loop integration to monitoring, alerting and SLO systems. Select tooling, define standards, and mentor a small SRE engineering team.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Bybit
Bybit
1 day ago

Senior Principal Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 21 hours agoStatus: Live

Job Summary

Design and build an enterprise chaos engineering platform for multi-cluster, multi-region and multi-environment Kubernetes/EC2 deployments. Own production safety controls such as blast radius limiting, kill switch, rollback and real-time impact monitoring, then orchestrate and execute routine resilience validation experiments with closed-loop integration to monitoring, alerting and SLO systems. Select tooling, define standards, and mentor a small SRE engineering team.
Location: Kuala Lumpur
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Sr. Director level

Key Responsibilities

  • •Design and build an enterprise chaos engineering platform for multi-cluster (K8s + EC2 hybrid), multi-region and multi-environment deployments
  • •Create fault injection capabilities including pod/node/AZ-level simulation and production safety controls like blast radius control, kill switch, rollback and impact monitoring
  • •Define safety standards and approval workflows for mainnet fault injection and run routine patrol/periodic/large-scale resilience experiments
  • •Develop a resilience scoring system and deliver improvement recommendations to remediate weaknesses
  • •Select the technology foundation, develop best practices/playbooks, and mentor a team of 2–3 engineers in chaos engineering

Key Requirements

  • •8+ years in backend/infrastructure engineering with 3+ years focused on chaos engineering or stability engineering
  • •Hands-on large-scale fault injection in production with deep knowledge of production safety constraints
  • •Expert proficiency in Kubernetes fault injection (Chaos Mesh/Litmus/custom), including CRD/operator development
  • •Proficiency in at least one backend language (Go preferred) and ability to design platform-level architectures
  • •Deep understanding of distributed system failure modes and observability (Prometheus/Grafana/Thanos/OpenTelemetry)
Experience:8+ yearsDistributed systemsSREKubernetesAWSObservability
Skills:Technical documentationSolution designCross-team collaborationRisk controlIndependent planning
Tech Stack:KubernetesChaos MeshLitmusGoCRDOperatorPrometheusGrafanaThanosOpenTelemetryEC2EKSAWSMulti-AZMulti-RegionNetflix Chaos EngineeringAWS Fault Injection SimulatorGremlin

Company Brief

Bybit
Operates a cryptocurrency exchange and trading platform offering spot, derivatives, copy trading, and related digital asset services for retail and institutional users. The platform focuses on high-liquidity crypto markets, trading tools, and Web3-related products.
Industry: Trading Platforms
Company Size: Enterprise (1,001+ employees)
Growth: Established Company
Founded: 2018
WebsiteLinkedIn