Senior Site Reliability Engineer

Akamai Technologies
Cambridge
Workplace: RemoteFull timeUSD 121,400 - 218,600 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 5+ yearsEducation: bachelorsSkills: ["Collaboration","Ownership","Problem-solving","Cross-functional coordination","Incident management"]

Build and scale SRE tooling for Akamai’s next-generation dedicated AI hardware infrastructure. Partner with product teams to ensure reliability, scalability, and performance across globe-spanning systems by defining and defending key performance indicators. Develop infrastructure-as-code utilities in Python, automate provisioning and break-fix workflows, implement telemetry and Prometheus/Grafana dashboards with anomaly detection, and lead 24x7 incident response using PagerDuty and Slack while coordinating vendors and field technicians.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Akamai Technologies
Akamai Technologies
1 week ago

Senior Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 21 hours agoStatus: Live
Reposted: similar role first listed 5 months ago

Job Summary

Build and scale SRE tooling for Akamai’s next-generation dedicated AI hardware infrastructure. Partner with product teams to ensure reliability, scalability, and performance across globe-spanning systems by defining and defending key performance indicators. Develop infrastructure-as-code utilities in Python, automate provisioning and break-fix workflows, implement telemetry and Prometheus/Grafana dashboards with anomaly detection, and lead 24x7 incident response using PagerDuty and Slack while coordinating vendors and field technicians.
Location: Cambridge
Workplace: Remote
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Develop and scale programmatic tooling and infrastructure-as-code utilities in Python to reduce operational toil and automate fleet-wide provisioning.
  • •Integrate automated workflows across disconnected corporate ticketing systems to improve time-to-mitigate for hardware and network break-fix events.
  • •Work with private cloud and compute technologies to improve availability, latency, and systemic health for high-density hardware environments.
  • •Design and implement telemetry pipelines, custom Prometheus/Grafana monitoring dashboards, and AI-based anomaly detection for bare-metal and virtualized environments.
  • •Participate in 24x7x365 on-call rotations, manage high-severity incident protocols with PagerDuty and Slack workflows, and coordinate with vendors and field technicians for uptime activities.

Pay and Benefits

Salary: USD 121,400 - 218,600 annually
Equity and Bonus:Equity
Perks:Health Insurance401kPaid LeaveSick TimeParental Leave

Key Requirements

  • •5+ years of relevant experience and a Bachelor’s degree in Computer Engineering, Computer Science, or equivalent.
  • •Strong Python tooling/coding ability to build scalable operational tools, automation frameworks, and API integrations.
  • •Hands-on experience with observability stacks and time series engines such as Prometheus, Grafana, OpenTelemetry, and Loki.
  • •Working knowledge of advanced networking topologies, high-bandwidth routing/switching, BGP, and dual-stack IPv4/IPv6 networks.
  • •Experience designing service rollouts with operational readiness criteria, telemetry baselines, alerting thresholds, and leading runbooks and blameless post-mortems.
Experience:5+ yearsDistributed systemsAIObservabilityInfrastructureOn-call
Education:Bachelor's in Computer Engineering or Computer Science (or equivalent)
Skills:CollaborationOwnershipProblem-solvingCross-functional coordinationIncident management
Tech Stack:PythonInfrastructure-as-codeAPI integrationsAutomation frameworksPrometheusGrafanaOpenTelemetryLokiTelemetry pipelinesAnomaly detectionPagerDutySlackAI utilitiesLLM-assisted developmentPrivate cloudRoutingSwitchingBGPIPv4IPv6

Company Brief

Akamai Technologies
Provides a global content delivery network (CDN) and cloud services to improve web and application performance, security, and delivery for enterprises, media companies, and cloud providers.
Industry: Cloud Computing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Cambridge, United States
Founded: 1998
WebsiteLinkedIn