Site Reliability Engineer – AI-first Platform
BlueCat Networks
Belgrade
Workplace: HybridFull timeFunction: DevOps, Cloud & InfrastructureExperience: 5+ yearsSkills: ["Automation","Incident management","Root cause analysis","Observability","Reliability focus"]Build and operate an AI-first platform for reliable, secure, observable, and scalable production workloads. You’ll run and scale Kubernetes (EKS) clusters, support AWS AgentCore productionization, and create CI/CD and deployment automation with GitLab CI/CD. Define SRE practices (SLIs/SLOs, alerting, incident response, postmortems), maintain observability, and implement infrastructure using Terraform while partnering with platform and development teams to improve reliability and delivery speed.

