Site Reliability Engineer
NTT
Guadalajara
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 6-8 yearsSkills: ["Communication","Incident management","Troubleshooting","Proactive problem-solving","Cross-functional collaboration"]Own and improve the reliability of cloud-based, microservice platforms by leading incident response, driving blameless postmortems, and improving service availability and scalability. You’ll manage production support through on-call, enhance Kubernetes-based workloads, and reduce operational toil with automation. Build Infrastructure as Code with Terraform, implement GitOps and CI/CD workflows with ArgoCD and related tooling, and strengthen monitoring, observability, performance, and disaster recovery readiness across AWS and Azure.

