Site Reliability Engineer
India
Workplace: RemoteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 3-6 yearsSkills: ["Incident response","Cross-functional coordination","Analytical thinking","Process improvement","Writing production code"]Engineer reliability for a highly distributed, hyper-growth platform by defining SLOs and error budgets, leading incident response, and building automation and self-healing mechanisms. Own on-call rotations and act as Incident Commander during major events. Improve observability with high-cardinality metrics and distributed traces, partner with product engineering to bake in reliability patterns from day one, and validate resilience through load testing and chaos engineering.
Loading
Loading job details...
Preparing the role view and application actions.

