Senior Principal Site Reliability Engineer
Kuala Lumpur
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 8+ yearsSkills: ["Technical documentation","Solution design","Cross-team collaboration","Risk control","Independent planning"]Design and build an enterprise chaos engineering platform that injects safe, production-ready faults across Kubernetes and EC2 hybrid, multi-cluster and multi-region environments. Own the closed-loop workflow from fault injection to monitoring-based pass/fail, define mainnet safety standards and approval workflows, and run resilience drills and scoring. Evaluate chaos engineering technology options, develop playbooks for SRE and app teams, and mentor 2–3 engineers.
Loading
Loading job details...
Preparing the role view and application actions.

