Site Reliability Engineer

Pythian
Hyderabad, India
Workplace: RemoteFull timeFunction: DevOps, Cloud & InfrastructureSkills: ["Problem-solving","Collaboration","Automation","Reliability","Scalability"]

Build and run resilient, high-performing infrastructure for large-scale distributed systems. You’ll operate and optimize Kubernetes clusters, Istio service mesh, and Linux-based environments, automate workflows with Go, Python, and Shell, and deliver observability with Prometheus, Grafana, and Loki. Collaborate with clients and AI/ML teams to ensure infrastructure readiness, while handling troubleshooting, on-call rotations, and postmortems to continuously improve reliability.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Pythian
Pythian
2 months ago

Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 7 hours agoStatus: Live
Reposted: similar role first listed 2 months ago

Job Summary

Build and run resilient, high-performing infrastructure for large-scale distributed systems. You’ll operate and optimize Kubernetes clusters, Istio service mesh, and Linux-based environments, automate workflows with Go, Python, and Shell, and deliver observability with Prometheus, Grafana, and Loki. Collaborate with clients and AI/ML teams to ensure infrastructure readiness, while handling troubleshooting, on-call rotations, and postmortems to continuously improve reliability.
Location: Hyderabad, India
Workplace: Remote
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure

Key Responsibilities

  • •Operate and optimize Kubernetes clusters, Istio service mesh, and Linux-based systems.
  • •Automate workflows using Go, Python, and Shell scripting.
  • •Build monitoring and observability solutions with Prometheus, Grafana, and Loki.
  • •Troubleshoot complex networking, storage, and system performance issues.
  • •Participate in on-call rotations and postmortem reviews to improve system resilience.

Pay and Benefits

Perks:Remote WorkLearning BudgetWellness StipendPaid LeaveSick Days

Key Requirements

  • •Experience with Google Cloud and Infrastructure-as-Code tools such as Terraform.
  • •Strong knowledge of microservices, containers (Kubernetes, Docker), and networking.
  • •SRE mindset focused on automation, scalability, and reliability.
  • •Hands-on experience with PKI, service mesh, and Linux systems administration.
  • •Ability to troubleshoot complex networking, storage, and system performance issues.
Skills:Problem-solvingCollaborationAutomationReliabilityScalability
Tech Stack:Google CloudTerraformKubernetesIstioLinuxGoPythonShell scriptingPrometheusGrafanaLokiMicroservicesDockerPKIService meshAI/MLObservability

Company Brief

Pythian
Provides data, cloud, and managed services to help organizations migrate, operate, and optimize analytics, databases, and cloud infrastructure across hybrid and multi-cloud environments.
Industry: Consulting
Company Size: Large (251 to 1,000 employees)
Growth: Established Company
Headquarters: Ottawa, Canada
Founded: 1997
WebsiteLinkedIn