Site Reliability Engineer
Pythian
Hyderabad, India
Workplace: HybridFull timeFunction: DevOps, Cloud & InfrastructureSkills: ["Problem-solving","Automation","Scalability","Reliability","Collaboration"]Build and run resilient, high-performing infrastructure for large-scale distributed systems across compute, storage, networking, and AI/ML environments. You’ll operate Kubernetes clusters and Istio service mesh, automate workflows with Go, Python, and Shell, and implement monitoring using Prometheus, Grafana, and Loki. Partner with clients and internal teams to troubleshoot complex issues, improve reliability through on-call and postmortems, and ensure infrastructure readiness for model training and data pipelines.

