Site Reliability Engineer

Pythian
Mexico, Costa Rica
Workplace: RemoteFull timeFunction: DevOps, Cloud & InfrastructureSkills: ["Problem-solving","Automation","Scalability","Reliability","Collaboration"]

Design, deploy, and operate large-scale distributed infrastructure as part of a new SRE team. Operate and optimize Kubernetes clusters and Istio service mesh on Linux, automate workflows with Go/Python/Shell, and build observability using Prometheus, Grafana, and Loki. Troubleshoot networking, storage, and performance issues, collaborate with AI/ML teams to support model training and data pipelines, and participate in on-call rotations and postmortems to improve resilience.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Pythian
Pythian
17 hours ago

Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 17 hours agoStatus: Live

Job Summary

Design, deploy, and operate large-scale distributed infrastructure as part of a new SRE team. Operate and optimize Kubernetes clusters and Istio service mesh on Linux, automate workflows with Go/Python/Shell, and build observability using Prometheus, Grafana, and Loki. Troubleshoot networking, storage, and performance issues, collaborate with AI/ML teams to support model training and data pipelines, and participate in on-call rotations and postmortems to improve resilience.
Location: Mexico, Costa Rica
Workplace: Remote
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure

Key Responsibilities

  • •Operate and optimize Kubernetes clusters, Istio service mesh, and Linux-based systems.
  • •Automate workflows using Go, Python, and Shell scripting.
  • •Build monitoring and observability solutions with Prometheus, Grafana, and Loki.
  • •Troubleshoot complex networking, storage, and system performance issues.
  • •Partner with AI/ML teams to ensure infrastructure readiness for model training and data pipelines, including on-call rotations and postmortems.

Pay and Benefits

Perks:Remote WorkLearning BudgetWellness StipendPaid LeaveGym Membership

Key Requirements

  • •Experience with Google Cloud and IaC tools such as Terraform.
  • •Strong knowledge of microservices and containers including Kubernetes and Docker.
  • •Hands-on experience with PKI, service mesh, and Linux systems administration.
  • •An SRE mindset focused on automation, scalability, and reliability.
  • •Experience with Golang is an asset.
Skills:Problem-solvingAutomationScalabilityReliabilityCollaboration
Tech Stack:Google CloudTerraformKubernetesIstioLinuxGoPythonShellPrometheusGrafanaLokiPKIDockerMicroservicesAI/ML

Company Brief

Pythian
Provides data, cloud, and managed services to help organizations migrate, operate, and optimize analytics, databases, and cloud infrastructure across hybrid and multi-cloud environments.
Industry: Consulting
Company Size: Large (251 to 1,000 employees)
Growth: Established Company
Headquarters: Ottawa, Canada
Founded: 1997
WebsiteLinkedIn