Site Reliability Engineer – AI-first Platform
Belgrade
Workplace: HybridFull timeFunction: DevOps, Cloud & InfrastructureExperience: 5+ yearsSkills: ["Automation","Incident management","Root cause analysis","Observability","Reliability focus"]Build and operate an AI-first platform for reliable, secure, observable, and scalable production workloads. You’ll run and scale Kubernetes (EKS) clusters, support AWS AgentCore productionization, and create CI/CD and deployment automation with GitLab CI/CD. Define SRE practices (SLIs/SLOs, alerting, incident response, postmortems), maintain observability, and implement infrastructure using Terraform while partnering with platform and development teams to improve reliability and delivery speed.
Loading
Loading job details...
Preparing the role view and application actions.

