Site Reliability Engineer
NTT
Guadalajara
Workplace: OnsiteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 6-8 yearsSkills: ["Communication","Incident response","Automation","Observability","Troubleshooting"]Own and improve reliability for cloud-based, microservice platforms running on AWS and Azure. You’ll lead incident response and blameless postmortems, participate in on-call rotation, and drive automation to reduce operational toil. Build infrastructure with Terraform and GitOps/deployment workflows using Atlantis and ArgoCD/CI-CD tools. Manage Kubernetes workloads, observability, performance/capacity trends, and support disaster recovery readiness, runbooks, and engineering documentation.

