Site Reliability Engineer
Mexico, Costa Rica
Workplace: RemoteFull timeFunction: DevOps, Cloud & InfrastructureSkills: ["Problem-solving","Automation","Scalability","Reliability","Collaboration"]Design, deploy, and operate large-scale distributed infrastructure as part of a new SRE team. Operate and optimize Kubernetes clusters and Istio service mesh on Linux, automate workflows with Go/Python/Shell, and build observability using Prometheus, Grafana, and Loki. Troubleshoot networking, storage, and performance issues, collaborate with AI/ML teams to support model training and data pipelines, and participate in on-call rotations and postmortems to improve resilience.
Loading
Loading job details...
Preparing the role view and application actions.

