Senior Principal Site Reliability Engineer
Kuala Lumpur
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 8+ yearsSkills: ["Technical documentation","Solution design","Cross-team collaboration","Risk control","Independent planning"]Design and build an enterprise chaos engineering platform for multi-cluster, multi-region and multi-environment Kubernetes/EC2 deployments. Own production safety controls such as blast radius limiting, kill switch, rollback and real-time impact monitoring, then orchestrate and execute routine resilience validation experiments with closed-loop integration to monitoring, alerting and SLO systems. Select tooling, define standards, and mentor a small SRE engineering team.
Loading
Loading job details...
Preparing the role view and application actions.

