Site Reliability Engineer

Pythian
Hyderabad, India
Workplace: HybridFull timeFunction: DevOps, Cloud & InfrastructureSkills: ["Problem-solving","Automation","Scalability","Reliability","Collaboration"]

Build and run resilient, high-performing infrastructure for large-scale distributed systems across compute, storage, networking, and AI/ML environments. You’ll operate Kubernetes clusters and Istio service mesh, automate workflows with Go, Python, and Shell, and implement monitoring using Prometheus, Grafana, and Loki. Partner with clients and internal teams to troubleshoot complex issues, improve reliability through on-call and postmortems, and ensure infrastructure readiness for model training and data pipelines.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Pythian
Pythian
2 months ago

Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 18 hours agoStatus: Live

Job Summary

Build and run resilient, high-performing infrastructure for large-scale distributed systems across compute, storage, networking, and AI/ML environments. You’ll operate Kubernetes clusters and Istio service mesh, automate workflows with Go, Python, and Shell, and implement monitoring using Prometheus, Grafana, and Loki. Partner with clients and internal teams to troubleshoot complex issues, improve reliability through on-call and postmortems, and ensure infrastructure readiness for model training and data pipelines.
Location: Hyderabad, India
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure

Key Responsibilities

  • •Operate and optimize Kubernetes clusters, Istio service mesh, and Linux-based systems.
  • •Automate workflows using Go, Python, and Shell scripting.
  • •Build monitoring and observability solutions with Prometheus, Grafana, and Loki.
  • •Troubleshoot complex networking, storage, and system performance issues.
  • •Participate in on-call rotations and postmortem reviews to improve system resilience, including partnering with AI/ML teams for infrastructure readiness.

Pay and Benefits

Perks:Remote WorkLearning BudgetHome OfficePaid LeaveWellness Stipend

Key Requirements

  • •Experience with Google Cloud and IaC tools such as Terraform.
  • •Strong knowledge of microservices and containers including Kubernetes and Docker.
  • •SRE mindset focused on automation, scalability, and reliability.
  • •Hands-on experience with PKI, service mesh, and Linux systems administration.
  • •Experience troubleshooting networking, storage, and system performance issues.
Experience:CloudMicroservicesContainersKubernetesAI/ML
Skills:Problem-solvingAutomationScalabilityReliabilityCollaboration
Tech Stack:KubernetesIstioLinuxGoPythonShellPrometheusGrafanaLokiTerraformGoogle CloudMicroservicesDockerPKINetworkingAI/MLAWSAzureOracleSAP

Company Brief

Pythian
Provides data, cloud, and managed services to help organizations migrate, operate, and optimize analytics, databases, and cloud infrastructure across hybrid and multi-cloud environments.
Industry: Consulting
Company Size: Large (251 to 1,000 employees)
Growth: Established Company
Headquarters: Ottawa, Canada
Founded: 1997
WebsiteLinkedIn