SRE Reliability Engineer
Bengaluru
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 5+ yearsSkills: ["Incident coordination","Root-cause analysis","Problem management","Analytical troubleshooting","Stakeholder communication","Automation mindset"]Own the reliability, availability, and operational health of business-critical production systems. Provide L2/L3 production support, triaging incidents, troubleshooting Kubernetes-deployed Java applications and microservices, and driving RCA and post-incident reviews. Implement observability with Datadog/Prometheus, define SLOs/SLIs/error budgets, use SQL for incident investigation, and automate operational toil to improve MTTR, performance, capacity, and stability.
Loading
Loading job details...
Preparing the role view and application actions.

