Site Reliability Engineer (Guardicore AI Platform) - Remote

Akamai Technologies
Madrid
Workplace: RemoteFull timeFunction: DevOps, Cloud & InfrastructureSkills: ["Problem-solving","Troubleshooting","Technical leadership","Cross-team collaboration","Ownership"]

Own the reliability, availability, performance, and operational readiness of a cloud-native Data and AI Platform for Akamai Guardicore Segmentation. Operate secure Kubernetes infrastructure for microservices, data pipelines, observability, and internal tooling, improving reliability, security, performance, and cost efficiency. Lead production investigations, participate in on-call rotations, and use LLM-driven automation to auto-remediate incidents while partnering across DevOps, Software, Data, AI, and Security teams.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Akamai Technologies
Akamai Technologies
1 month ago

Site Reliability Engineer (Guardicore AI Platform) - Remote

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 6 hours agoStatus: Live

Job Summary

Own the reliability, availability, performance, and operational readiness of a cloud-native Data and AI Platform for Akamai Guardicore Segmentation. Operate secure Kubernetes infrastructure for microservices, data pipelines, observability, and internal tooling, improving reliability, security, performance, and cost efficiency. Lead production investigations, participate in on-call rotations, and use LLM-driven automation to auto-remediate incidents while partnering across DevOps, Software, Data, AI, and Security teams.
Location: Madrid
Workplace: Remote
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Operate secure, highly available Kubernetes infrastructure for core microservices, data pipelines, observability, and internal tooling.
  • •Enhance platform reliability, observability, security, performance, and cost efficiency.
  • •Provide guidance to engineers and developers to ensure services perform as expected.
  • •Lead complex production investigations and drive long-term improvements.
  • •Use LLMs and AI-driven automation to auto-remediate incidents and streamline operations, including on-call rotations and cross-team troubleshooting.

Pay and Benefits

Perks:Remote Work

Key Requirements

  • •3+ years of experience in SRE, DevOps, or Platform Engineering, with a track record of mastering and troubleshooting complex system architectures.
  • •Design and implement monitoring and observability strategies using Prometheus and Grafana.
  • •Production experience with Kubernetes, Docker, Helm, and third-party clouds (GCP, Azure, Linode, AWS) on Linux-based infrastructure.
  • •Exceptional troubleshooting and problem-solving skills across network, system, application, and database layers.
  • •Experience with GitOps, CI/CD, and Infrastructure as Code; scripting/programming proficiency in Python, Go, and Bash.
Experience:SREDevOpsPlatform engineeringCloud-nativeCybersecurityData & AI
Skills:Problem-solvingTroubleshootingTechnical leadershipCross-team collaborationOwnership
Tech Stack:KubernetesMicroservicesData pipelinesObservabilityPrometheusGrafanaDockerHelmGCPAzureLinodeAWSLinuxGitOpsCI/CDInfrastructure as CodePythonGoBashLLMs

Company Brief

Akamai Technologies
Provides a global content delivery network (CDN) and cloud services to improve web and application performance, security, and delivery for enterprises, media companies, and cloud providers.
Industry: Cloud Computing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Cambridge, United States
Founded: 1998
WebsiteLinkedIn