Senior Site Reliability Engineer

Akamai Technologies
Indiana, India
Workplace: RemoteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 5+ yearsEducation: bachelorsSkills: ["Collaboration","Automation","Incident response","Cross-functional coordination","Ownership"]

Join the AI Hardware SRE team to oversee, scale, and optimize next-generation dedicated AI hardware infrastructure. Define KPIs, drive proactive monitoring, automation, and rapid incident resolution across high-density hardware and regional data centers. Build Python-based infrastructure-as-code and operational tooling, implement observability with Prometheus/Grafana and OpenTelemetry/Loki, and lead telemetry/telemetry baselines for reliable service rollouts while participating in 24x7x365 on-call.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Akamai Technologies
Akamai Technologies
1 day ago

Senior Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 4 hours agoStatus: Live

Job Summary

Join the AI Hardware SRE team to oversee, scale, and optimize next-generation dedicated AI hardware infrastructure. Define KPIs, drive proactive monitoring, automation, and rapid incident resolution across high-density hardware and regional data centers. Build Python-based infrastructure-as-code and operational tooling, implement observability with Prometheus/Grafana and OpenTelemetry/Loki, and lead telemetry/telemetry baselines for reliable service rollouts while participating in 24x7x365 on-call.
Location: Indiana, India
Workplace: Remote
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Develop and scale infrastructure-as-code utilities and programmatic tooling in Python to reduce operational toil and automate fleet-wide provisioning
  • •Integrate automated workflows with corporate ticketing systems to improve resolution times for hardware and network break-fix incidents
  • •Design and implement telemetry pipelines, Prometheus/Grafana dashboards, and AI-based anomaly detection for bare-metal and virtualized environments
  • •Work on private cloud and compute technologies to improve availability, latency, and systemic health of high-density hardware environments
  • •Own real-time incident management during 24x7x365 on-call, including high-severity service disruption protocols using PagerDuty and Slack workflows

Key Requirements

  • •5+ years of relevant experience and a Bachelor's degree in Computer Science or related field
  • •Strong Python proficiency to build scalable operational tools, API integrations, and automation frameworks
  • •Hands-on experience with observability stacks and time-series monitoring tools like Prometheus, Grafana, OpenTelemetry, and Loki
  • •Working knowledge of advanced networking topologies, high-bandwidth routing/switching, BGP, and dual-stack IPv4/IPv6 networks
  • •Experience designing runbooks, leading incident response, and driving blameless post-mortems across complex incidents
Experience:5+ yearsAI infrastructureDistributed systemsPrivate cloudData centersHigh-density hardware
Education:Bachelor's in Computer Science
Skills:CollaborationAutomationIncident responseCross-functional coordinationOwnership
Tech Stack:PythonInfrastructure-as-codePrometheusGrafanaOpenTelemetryLokiPagerDutySlackTelemetry pipelinesAILLMAPI integrationsBGPIPv4IPv6TelemetryTimeseries

Company Brief

Akamai Technologies
Provides a global content delivery network (CDN) and cloud services to improve web and application performance, security, and delivery for enterprises, media companies, and cloud providers.
Industry: Cloud Computing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Cambridge, United States
Founded: 1998
WebsiteLinkedIn