Senior Principal Site Reliability Engineer

Bybit
Kuala Lumpur
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 8+ yearsSkills: ["Technical documentation","Solution design","Cross-team collaboration","Risk control","Independent planning"]

Design and build an enterprise chaos engineering platform that injects safe, production-ready faults across Kubernetes and EC2 hybrid, multi-cluster and multi-region environments. Own the closed-loop workflow from fault injection to monitoring-based pass/fail, define mainnet safety standards and approval workflows, and run resilience drills and scoring. Evaluate chaos engineering technology options, develop playbooks for SRE and app teams, and mentor 2–3 engineers.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Bybit
Bybit
1 day ago

Senior Principal Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 21 hours agoStatus: Live

Job Summary

Design and build an enterprise chaos engineering platform that injects safe, production-ready faults across Kubernetes and EC2 hybrid, multi-cluster and multi-region environments. Own the closed-loop workflow from fault injection to monitoring-based pass/fail, define mainnet safety standards and approval workflows, and run resilience drills and scoring. Evaluate chaos engineering technology options, develop playbooks for SRE and app teams, and mentor 2–3 engineers.
Location: Kuala Lumpur
Workplace: Onsite
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Design and build an enterprise chaos engineering platform for multi-cluster, multi-region, and multi-environment deployments.
  • •Implement fault injection capabilities (fault injection, network latency/packet loss/partition, dependency timeouts, and error injection) with production safety controls (blast radius, kill switch, rollback, impact monitoring).
  • •Define mainnet safety standards and approval workflows, and drive routine resilience experiments (daily patrol, periodic validations, cross-AZ/region drills).
  • •Integrate chaos experiments with monitoring/alerting/SLO systems to enable closed-loop inject→observe→pass/fail automation and establish resilience scoring.
  • •Evaluate and select chaos engineering technologies, develop best practices/playbooks, and mentor a small SRE team in chaos engineering capabilities.

Pay and Benefits

Perks:Learning Budget

Key Requirements

  • •8+ years of backend/infrastructure engineering experience, including 3+ years focused on chaos engineering or stability engineering.
  • •Hands-on production fault injection experience with deep knowledge of production safety constraints.
  • •Expert in Kubernetes fault injection (Chaos Mesh/Litmus/custom), including CRD/operator development.
  • •Proficient in a backend language (Go preferred) and able to design platform-level systems architecture.
  • •Strong understanding of distributed system failure modes and observability (Prometheus/Grafana/Thanos/OpenTelemetry).
Experience:8+ yearsSREChaos engineeringDistributed systemsKubernetesAWSObservabilityFintech/trading
Skills:Technical documentationSolution designCross-team collaborationRisk controlIndependent planning
Languages:English
Tech Stack:GoKubernetesK8sEC2EKSChaos MeshLitmusCRDOperatorPrometheusGrafanaThanosOpenTelemetryAWS Fault Injection SimulatorGremlinNetflix Chaos Engineering

Company Brief

Bybit
Operates a cryptocurrency exchange and trading platform offering spot, derivatives, copy trading, and related digital asset services for retail and institutional users. The platform focuses on high-liquidity crypto markets, trading tools, and Web3-related products.
Industry: Trading Platforms
Company Size: Enterprise (1,001+ employees)
Growth: Established Company
Founded: 2018
WebsiteLinkedIn