Site Reliability Engineer – AI-first Platform

BlueCat Networks
Belgrade
Workplace: HybridFull timeFunction: DevOps, Cloud & InfrastructureExperience: 5+ yearsSkills: ["Automation","Incident management","Root cause analysis","Observability","Reliability focus"]

Build and operate an AI-first platform for reliable, secure, observable, and scalable production workloads. You’ll run and scale Kubernetes (EKS) clusters, support AWS AgentCore productionization, and create CI/CD and deployment automation with GitLab CI/CD. Define SRE practices (SLIs/SLOs, alerting, incident response, postmortems), maintain observability, and implement infrastructure using Terraform while partnering with platform and development teams to improve reliability and delivery speed.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
BlueCat Networks
BlueCat Networks
1 month ago

Site Reliability Engineer – AI-first Platform

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 6 hours agoStatus: Live

Job Summary

Build and operate an AI-first platform for reliable, secure, observable, and scalable production workloads. You’ll run and scale Kubernetes (EKS) clusters, support AWS AgentCore productionization, and create CI/CD and deployment automation with GitLab CI/CD. Define SRE practices (SLIs/SLOs, alerting, incident response, postmortems), maintain observability, and implement infrastructure using Terraform while partnering with platform and development teams to improve reliability and delivery speed.
Location: Belgrade
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Operate and scale Kubernetes (EKS) clusters running AI and cloud-native workloads.
  • •Design and maintain CI/CD pipelines using GitLab CI/CD for applications and platform services.
  • •Implement observability (metrics, logs, traces) and enforce SRE practices including SLIs/SLOs, alerting, incident response, and postmortems.
  • •Participate in on-call rotation and support incident resolution and root cause analysis.
  • •Build repeatable infrastructure using Terraform and improve cost efficiency, performance, and reliability.

Pay and Benefits

Perks:Learning BudgetWellness Stipend

Key Requirements

  • •5+ years of experience in SRE, DevOps, or Platform Engineering.
  • •Hands-on experience with AWS and Kubernetes (EKS).
  • •Experience operating or supporting AWS AgentCore or similar AI/agent platforms.
  • •Strong proficiency in Python for automation.
  • •Knowledge of AWS identity and access management (IAM) and Infrastructure as Code with Terraform preferred.
Experience:5+ yearsSREDevOpsPlatform Engineering
Skills:AutomationIncident managementRoot cause analysisObservabilityReliability focus
Tech Stack:AWSKubernetesEKSAWS AgentCorePythonTerraformGitLab CI/CDIAMSSOIdentity CenterSLIsSLOsMetricsLogsTracesCI/CDPostmortems

Company Brief

BlueCat Networks
Provides enterprise DNS, DHCP and IP address management (DDI) solutions that enable centralized network control, automation, and security for large organizations and service providers across cloud, hybrid, and on-premises environments.
Industry: Cybersecurity
Company Size: Large (251 to 1,000 employees)
Growth: Established Company
Headquarters: Toronto, Canada
Founded: 2001
WebsiteLinkedIn