Senior Site Reliability Engineer

Akamai Technologies
Poland
Workplace: RemoteFull timeFunction: DevOps, Cloud & InfrastructureSkills: ["Problem-solving","Collaboration","Ownership","Communication"]

Senior SRE responsible for owning reliability workstreams for Akamai's serverless inference platform, building automation and tooling, and guiding architecture and operational decisions. You’ll own critical reliability problems end-to-end, partner with product engineering, and develop expertise in GPU infrastructure, Kubernetes at scale, and AI inference workloads, while improving observability and deployment safety.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Akamai Technologies
Akamai Technologies
5 months ago

Senior Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 6 hours agoStatus: Live

Job Summary

Senior SRE responsible for owning reliability workstreams for Akamai's serverless inference platform, building automation and tooling, and guiding architecture and operational decisions. You’ll own critical reliability problems end-to-end, partner with product engineering, and develop expertise in GPU infrastructure, Kubernetes at scale, and AI inference workloads, while improving observability and deployment safety.
Location: Poland
Workplace: Remote
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Manager level

Key Responsibilities

  • •Build and maintain observability for AI workloads, including telemetry, dashboards, alerts, SLO/SLI tracking, and drive improvements when targets are missed
  • •Write automation and tooling to reduce operational toil, improve deployment safety, and accelerate incident response
  • •Integrate AI workloads into existing incident management processes, build runbooks, participate in on-call rotations, and conduct blameless post-mortems
  • •Build and maintain CI/CD integrations, deployment safety checks, and rollback automation
  • •Collaborate with product engineering teams to improve reliability, contribute to architecture decisions, and ensure operational readiness for product releases

Key Requirements

  • •SRE, infrastructure, or platform engineering with experience managing large-scale distributed systems
Experience:AI/MLCloudDistributed systems
Skills:Problem-solvingCollaborationOwnershipCommunication
Tech Stack:PrometheusGrafanaDistributed tracingKubernetesPythonGoTerraform

Company Brief

Akamai Technologies
Provides a global content delivery network (CDN) and cloud services to improve web and application performance, security, and delivery for enterprises, media companies, and cloud providers.
Industry: Cloud Computing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Cambridge, United States
Founded: 1998
WebsiteLinkedIn