Senior Site Reliability Engineer

Akamai Technologies
Krakow
Workplace: RemoteFull timeFunction: DevOps, Cloud & InfrastructureSkills: ["Collaboration","Analytical thinking","Ownership","Incident response","Problem-solving"]

Ensure reliability and operational readiness for AI hardware and software systems across regional data centers. Partner with product teams to improve scalability, performance, and uptime by defining KPIs, monitoring breaches, and driving incident response. Build and scale Python infrastructure-as-code tooling, automate workflows across JIRA/Siebel/PagerDuty, and design observability pipelines with Prometheus/Grafana and OpenTelemetry/Loki. Join 24x7x365 on-call and coordinate with vendors and field technicians to maintain service availability.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Akamai Technologies
Akamai Technologies
1 month ago

Senior Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 11 hours agoStatus: Live
Reposted: similar role first listed 5 months ago

Job Summary

Ensure reliability and operational readiness for AI hardware and software systems across regional data centers. Partner with product teams to improve scalability, performance, and uptime by defining KPIs, monitoring breaches, and driving incident response. Build and scale Python infrastructure-as-code tooling, automate workflows across JIRA/Siebel/PagerDuty, and design observability pipelines with Prometheus/Grafana and OpenTelemetry/Loki. Join 24x7x365 on-call and coordinate with vendors and field technicians to maintain service availability.
Location: Krakow
Workplace: Remote
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Develop and scale programmatic tooling and infrastructure-as-code utilities in Python to reduce operational toil and automate fleet provisioning.
  • •Integrate automated workflows across ticketing platforms (JIRA, Siebel, PagerDuty) to improve resolution for hardware and network issues.
  • •Leverage AI utilities and LLM-assisted development to enhance technical execution, script creation, and system evaluation.
  • •Design and implement telemetry pipelines, monitoring dashboards (Prometheus/Grafana), and AI-based anomaly detection for bare-metal and virtualized environments.
  • •Participate in 24x7x365 on-call, manage real-time incidents, and run high-severity service disruption protocols using PagerDuty and Slack workflows.

Key Requirements

  • •Solid Computer Science foundation via formal education or equivalent practical experience in large-scale SRE or Production Engineering roles.
  • •Proficient in Python for scalable operational tooling, API integrations, and automation frameworks.
  • •Hands-on experience with observability stacks and time-series tools such as Prometheus, Grafana, OpenTelemetry, and Loki.
  • •Working understanding of networking topologies and high-bandwidth routing/switching, including BGP and dual-stack IPv4/IPv6.
  • •Expertise designing service rollouts with operational readiness criteria, telemetry baselines, and effective alerting thresholds.
Experience:SREProduction engineering
Skills:CollaborationAnalytical thinkingOwnershipIncident responseProblem-solving
Tech Stack:PythonInfrastructure-as-codeJIRASiebelPagerDutySlackLLMPrivate cloudComputeTelemetry pipelinesPrometheusGrafanaOpenTelemetryLokiAPI integrationsBGPIPv4IPv6

Company Brief

Akamai Technologies
Provides a global content delivery network (CDN) and cloud services to improve web and application performance, security, and delivery for enterprises, media companies, and cloud providers.
Industry: Cloud Computing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Cambridge, United States
Founded: 1998
WebsiteLinkedIn