Site Reliability Engineer

Akamai Technologies
Cambridge
Workplace: RemoteFull timeUSD 75,700 - 136,300 annuallyFunction: DevOps, Cloud & InfrastructureEducation: bachelorsSkills: []

Build and operate resilient compute platform services by troubleshooting complex Linux, networking, and distributed-system issues. Create software and automation to reduce operational toil, improve efficiency, and prevent recurring incidents. Develop AI-assisted tooling to speed incident investigation, establish monitoring with SLIs/SLOs, and drive root-cause analysis through post-incident reviews. Partner with engineering teams on deployment safety and operational readiness while participating in on-call incident response.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Akamai Technologies
Akamai Technologies
4 days ago

Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 6 hours agoStatus: Live
Reposted: similar role first listed 5 months ago

Job Summary

Build and operate resilient compute platform services by troubleshooting complex Linux, networking, and distributed-system issues. Create software and automation to reduce operational toil, improve efficiency, and prevent recurring incidents. Develop AI-assisted tooling to speed incident investigation, establish monitoring with SLIs/SLOs, and drive root-cause analysis through post-incident reviews. Partner with engineering teams on deployment safety and operational readiness while participating in on-call incident response.
Location: Cambridge
Workplace: Remote
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure

Key Responsibilities

  • •Troubleshoot complex production issues across Linux systems, networking, and distributed services.
  • •Build software and automation to reduce operational toil and improve efficiency and reliability.
  • •Develop AI-assisted tooling to accelerate incident investigation and identify operational patterns.
  • •Establish and improve monitoring, alerting, SLIs, and SLOs for critical services.
  • •Participate in on-call rotation and lead incident response with timely restoration and post-incident improvements.

Pay and Benefits

Salary: USD 75,700 - 136,300 annually
Equity and Bonus:Equity
Perks:Health Insurance401kPaid LeaveSick TimeParental LeaveEmployee Assistance

Key Requirements

  • •Bachelor's degree in Computer Engineering, Computer Science, or equivalent.
  • •Experience supporting large-scale distributed systems.
  • •Strong Linux and networking knowledge, including routing, DNS, firewalls, TCP/IP, and L7 traffic management.
  • •Proficiency in a programming language such as Python or Go.
  • •Experience with observability tools and infrastructure automation/configuration management (e.g., Prometheus, Grafana, Terraform, Ansible).
Experience:Distributed systems
Education:Bachelor's in Computer Engineering or Computer Science
Tech Stack:LinuxPythonGoPrometheusGrafanaLokiELKOpenSearchTerraformAnsibleSaltDockerPodmanKubernetesNomadDNSFirewallsTCP/IPL7 traffic managementSLIs

Company Brief

Akamai Technologies
Provides a global content delivery network (CDN) and cloud services to improve web and application performance, security, and delivery for enterprises, media companies, and cloud providers.
Industry: Cloud Computing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Cambridge, United States
Founded: 1998
WebsiteLinkedIn