Site Reliability Engineer II

Akamai Technologies
Cambridge, United States
Workplace: RemoteFull timeUSD 95,000 - 171,000 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 2+ yearsEducation: bachelorsSkills: ["Collaboration","Troubleshooting","On-call readiness","Willingness to learn","Curiosity"]

Build and run reliable AI inference infrastructure by automating monitoring and incident response. Improve dashboards, alerts, and SLO tracking for inference workloads using Akamai observability tooling. Develop automation and runbooks, participate in CI/CD safety and rollback processes, and support on-call rotations with blameless post-mortems. Collaborate with product engineering teams to troubleshoot issues across the stack, with opportunities to focus on GPU infrastructure and Kubernetes for serverless inference workloads.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Akamai Technologies
Akamai Technologies
5 months ago

Site Reliability Engineer II

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 11 hours agoStatus: Live

Job Summary

Build and run reliable AI inference infrastructure by automating monitoring and incident response. Improve dashboards, alerts, and SLO tracking for inference workloads using Akamai observability tooling. Develop automation and runbooks, participate in CI/CD safety and rollback processes, and support on-call rotations with blameless post-mortems. Collaborate with product engineering teams to troubleshoot issues across the stack, with opportunities to focus on GPU infrastructure and Kubernetes for serverless inference workloads.
Location: Cambridge, United States
Workplace: Remote
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Build and maintain dashboards, alerts, and monitoring for inference workloads using Akamai observability.
  • •Write automation and tooling in Python or Go to reduce operational toil and improve reliability.
  • •Build and improve inference-specific runbooks and integrate them into incident management processes.
  • •Contribute to SLO tracking and reporting, identifying trends and improvement areas.
  • •Support CI/CD pipeline maintenance, deployment safety checks, rollback procedures, and participate in on-call rotations with incident response and blameless post-mortems.

Pay and Benefits

Salary: USD 95,000 - 171,000 annually
Equity and Bonus:Equity
Perks:Health Insurance401kPaid LeaveParental Leave

Key Requirements

  • •2+ years of experience in Site Reliability Engineering and a Bachelor's Degree or equivalent experience.
  • •Coding ability in at least one language (Python or Go), with experience writing automation.
  • •Experience with Linux systems administration and troubleshooting complex infrastructure issues.
  • •Familiarity with Kubernetes and containerization concepts.
  • •Experience with monitoring/observability tools such as Prometheus and Grafana, plus exposure to CI/CD and infrastructure-as-code tools (Terraform, SaltStack, or equivalent).
Experience:2+ yearsAI infrastructureCloud computingDistributed systemsServerless
Education:Bachelor's
Skills:CollaborationTroubleshootingOn-call readinessWillingness to learnCuriosity
Tech Stack:LinuxPythonGoGPU infrastructureKubernetesContainerizationPrometheusGrafanaObservability platformDashboardsAlertsSLO trackingCI/CDInfrastructure-as-codeTerraformSaltStackRunbooksIncident management

Company Brief

Akamai Technologies
Provides a global content delivery network (CDN) and cloud services to improve web and application performance, security, and delivery for enterprises, media companies, and cloud providers.
Industry: Cloud Computing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Cambridge, United States
Founded: 1998
WebsiteLinkedIn