Senior Site Reliability Engineer

Akamai Technologies
Bengaluru
Workplace: RemoteFull timeFunction: DevOps, Cloud & InfrastructureEducation: bachelorsSkills: ["Collaboration","Mentorship","Troubleshooting","Problem-solving","Automation"]

Ensure the operation and uptime of Compute services and infrastructure, supervising critical systems and partnering across teams to improve availability, reliability, scalability, and usability. Define requirements during the product lifecycle, deploy and maintain internal platforms and tools, and modernize Compute cloud interfaces to speed error detection and remediation. Build automation to reduce toil and participate in on-call rotations to guide restoration and resolve escalations and incidents.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Akamai Technologies
Akamai Technologies
2 days ago

Senior Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 13 hours agoStatus: Live
Reposted: similar role first listed 6 months ago

Job Summary

Ensure the operation and uptime of Compute services and infrastructure, supervising critical systems and partnering across teams to improve availability, reliability, scalability, and usability. Define requirements during the product lifecycle, deploy and maintain internal platforms and tools, and modernize Compute cloud interfaces to speed error detection and remediation. Build automation to reduce toil and participate in on-call rotations to guide restoration and resolve escalations and incidents.
Location: Bengaluru
Workplace: Remote
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Ensure operation and uptime of Compute services and infrastructure, supervising critical systems.
  • •Collaborate across teams to create tooling and software that monitors and improves reliability of systems.
  • •Deploy and maintain internal platforms and tools, and influence new designs and standards via requirements definition.
  • •Improve the Compute Cloud Interface platform to speed error detection and remediation, enhancing performance and reliability.
  • •Develop automation to reduce toil, participate in on-call rotations, and troubleshoot/escalate incident resolution for customers.

Key Requirements

  • •5+ years of relevant experience and a Bachelor's degree in Computer Science (or equivalent).
  • •Experience automating with Python and/or Golang, plus scripting with bash.
  • •Knowledge of systems reliability practices, including observability/monitoring and adherence to SLOs.
  • •Experience with configuration management tools such as SaltStack, Terraform, and Ansible, plus CI/CD solutions like Jenkins.
  • •Hands-on Linux administration and container platforms such as Docker.
Education:Bachelor's in Computer Science
Skills:CollaborationMentorshipTroubleshootingProblem-solvingAutomation
Tech Stack:PythonGolangBashSaltStackTerraformAnsibleJenkinsLinuxDockerPrometheusGrafanaLokiNginxEnvoyHaproxyRedis

Company Brief

Akamai Technologies
Provides a global content delivery network (CDN) and cloud services to improve web and application performance, security, and delivery for enterprises, media companies, and cloud providers.
Industry: Cloud Computing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Cambridge, United States
Founded: 1998
WebsiteLinkedIn