Senior Site Reliability Engineer

Pythian
Argentina, Brazil, Uruguay, Mexico, Costa Rica
Workplace: RemoteFull timeFunction: DevOps, Cloud & InfrastructureSkills: ["Automation","Scalability","Reliability","Problem-solving","Collaboration","Troubleshooting"]

Design, deploy, and operate large-scale distributed infrastructure as part of Pythian’s next-generation SRE team. You’ll run and optimize Kubernetes clusters with Istio service mesh, build monitoring and observability with Prometheus/Grafana/Loki, automate workflows using Go, Python, and Shell, and troubleshoot complex networking and storage performance. Partner with AI/ML teams and improve resilience through on-call rotations and postmortems.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Pythian
Pythian
17 hours ago

Senior Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 17 hours agoStatus: Live

Job Summary

Design, deploy, and operate large-scale distributed infrastructure as part of Pythian’s next-generation SRE team. You’ll run and optimize Kubernetes clusters with Istio service mesh, build monitoring and observability with Prometheus/Grafana/Loki, automate workflows using Go, Python, and Shell, and troubleshoot complex networking and storage performance. Partner with AI/ML teams and improve resilience through on-call rotations and postmortems.
Location: Argentina, Brazil, Uruguay, Mexico, Costa Rica
Workplace: Remote
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Operate and optimize Kubernetes clusters, Istio service mesh, and Linux-based systems.
  • •Automate workflows using Go, Python, and Shell scripting.
  • •Build monitoring and observability solutions with Prometheus, Grafana, and Loki.
  • •Troubleshoot complex networking, storage, and system performance issues.
  • •Participate in on-call rotations and postmortems to improve system resilience.

Pay and Benefits

Perks:Training AllowanceRemote WorkWellness StipendPaid LeaveSick Days

Key Requirements

  • •Experience with Google Cloud and IaC tools such as Terraform.
  • •Strong knowledge of microservices, containers (Kubernetes, Docker), and networking.
  • •Hands-on experience with PKI, service mesh, and Linux systems administration.
  • •An SRE mindset focused on automation, scalability, and reliability.
  • •Experience with Golang is an asset.
Skills:AutomationScalabilityReliabilityProblem-solvingCollaborationTroubleshooting
Tech Stack:Google CloudTerraformKubernetesIstioLinuxGoPythonShell scriptingPrometheusGrafanaLokiMicroservicesContainersDockerNetworkingPKIService meshAI/MLAWSMicrosoft

Company Brief

Pythian
Provides data, cloud, and managed services to help organizations migrate, operate, and optimize analytics, databases, and cloud infrastructure across hybrid and multi-cloud environments.
Industry: Consulting
Company Size: Large (251 to 1,000 employees)
Growth: Established Company
Headquarters: Ottawa, Canada
Founded: 1997
WebsiteLinkedIn