Site Reliability Engineer, Infrastructure - ThousandEyes

Cisco
London
Workplace: HybridFull timeFunction: DevOps, Cloud & InfrastructureEducation: bachelorsSkills: ["Communication","Documentation","Ownership","Technical judgment","Attention to operational details"]

Design and operate highly available distributed systems for the ThousandEyes platform, enabling reliable processing of large volumes of telemetry data. Use AI tooling to write production-quality code and automate reliable releases, while evaluating scalability, resiliency, performance, and security. Troubleshoot complex infrastructure and platform issues, participate in on-call and incident management, and improve reliability through root-cause findings. Collaborate with engineering teams to meet service objectives and customer SLAs.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Cisco
Cisco
1 month ago

Site Reliability Engineer, Infrastructure - ThousandEyes

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 2 hours agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Design and operate highly available distributed systems for the ThousandEyes platform, enabling reliable processing of large volumes of telemetry data. Use AI tooling to write production-quality code and automate reliable releases, while evaluating scalability, resiliency, performance, and security. Troubleshoot complex infrastructure and platform issues, participate in on-call and incident management, and improve reliability through root-cause findings. Collaborate with engineering teams to meet service objectives and customer SLAs.
Location: London
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Design and operate large-scale, highly available distributed systems for ThousandEyes.
  • •Use AI tooling to write high-quality code and automate solutions for fast, reliable releases and reduced operational expense.
  • •Evaluate scalability, resiliency, performance, and security of production services, including high-availability and disaster-recovery testing.
  • •Troubleshoot complex issues across infrastructure and platform services, participate in on-call and incident management, and drive reliability improvements from root-cause findings.
  • •Collaborate with application development teams and stakeholders to meet service level objectives and customer-facing service level agreements while improving lifecycle tooling and automation.

Key Requirements

  • •Proficiency writing production-quality code in at least one language such as Python or Go.
  • •Hands-on experience designing, operating, or troubleshooting production systems in AWS and Kubernetes, including on-call rotation.
  • •Hands-on experience with infrastructure-as-code tooling and codebases, preferably Terraform.
  • •Hands-on experience leveraging AI to support SRE activities, such as automating toil and improving operational efficiency.
  • •Professional experience administering and troubleshooting GNU/Linux systems, including system libraries, file systems, networking, and client-server protocols.
Education:Bachelor's
Skills:CommunicationDocumentationOwnershipTechnical judgmentAttention to operational details
Tech Stack:PythonGoAWSKubernetesTerraformGNU/LinuxLinuxAIInfrastructure as code

Company Brief

Cisco
Global technology company that designs, manufactures, and sells networking hardware, telecommunications equipment, and high-technology services and products for enterprises, service providers, and governments worldwide.
Industry: Networking Equipment
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: San Jose, United States
Founded: 1984
Glassdoor
Glassdoor: 4.0
WebsiteLinkedInGlassdoor