Senior Site Reliability Engineer

Akamai Technologies
Krakow
Workplace: RemoteFull timeFunction: DevOps, Cloud & InfrastructureSkills: ["Reliability","Scalability","Continuous improvement","Operational excellence","Collaboration"]

Design, develop, and operate reliable systems that support Akamai Compute products at internet scale. Partner with engineers to automate deployments, improve monitoring, and speed incident detection and remediation through continuous improvements. Lead on-call rotations, write tooling to reduce operational toil, and help with capacity planning, autoscaling, and workload scheduling for AI compute infrastructure. Work with Kubernetes, Linux, and infrastructure-as-code to deliver scalability, efficiency, and operational excellence.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Akamai Technologies
Akamai Technologies
3 days ago

Senior Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 2 hours agoStatus: Live
Reposted: similar role first listed 5 months ago

Job Summary

Design, develop, and operate reliable systems that support Akamai Compute products at internet scale. Partner with engineers to automate deployments, improve monitoring, and speed incident detection and remediation through continuous improvements. Lead on-call rotations, write tooling to reduce operational toil, and help with capacity planning, autoscaling, and workload scheduling for AI compute infrastructure. Work with Kubernetes, Linux, and infrastructure-as-code to deliver scalability, efficiency, and operational excellence.
Location: Krakow
Workplace: Remote
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Support and mentor engineers in the department.
  • •Develop automated tools and scripts to improve system reliability, deployment processes, and incident response efficiency.
  • •Improve system monitoring for faster error detection and remediation, strengthening performance and reliability of the virtualization platform.
  • •Participate in on-call rotations to guide restoration and repair of service-impacting issues.
  • •Write automation and tooling to reduce operational toil, improve deployment safety, and accelerate incident response; contribute to capacity planning, autoscaling configuration, and workload scheduling for AI compute infrastructure.

Key Requirements

  • •Expert experience in a SysAdmin (Linux/Unix Administration), DevOps, or SRE role supporting large-scale distributed systems.
  • •Expertise in Kubernetes and large-scale containerization systems.
  • •At least one programming language (Python or Golang) and configuration management with Terraform and either SaltStack or Ansible.
  • •Define SLOs and use observability tools such as Prometheus, Grafana, and distributed tracing to improve monitoring and reliability.
  • •Experience architecting software and infrastructure at scale, collaborating effectively with engineering teams new to SRE practices.
Experience:Distributed systemsKubernetesDevOpsSRE
Skills:ReliabilityScalabilityContinuous improvementOperational excellenceCollaboration
Tech Stack:LinuxUnixKubernetesPythonGolangTerraformSaltStackAnsiblePrometheusGrafanaDistributed tracingAutoscalingVirtualizationAI compute infrastructure

Company Brief

Akamai Technologies
Provides a global content delivery network (CDN) and cloud services to improve web and application performance, security, and delivery for enterprises, media companies, and cloud providers.
Industry: Cloud Computing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Cambridge, United States
Founded: 1998
WebsiteLinkedIn