Senior Site Reliability Engineer, Production Engineer - ThousandEyes

Cisco
San Francisco
Workplace: HybridFull timeUSD 165,000 - 241,400 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 5+ yearsSkills: ["Ownership","Communication","Documentation","Attention to detail"]

Design and manage large-scale, highly available distributed systems for the ThousandEyes SaaS platform. Collaborate with application teams to improve reliability, performance, and security, while building scalable operations tooling and incident response practices. Implement and maintain AWS cloud-native services, expand CNCF-based reliability stack (Kubernetes, Service Mesh, Prometheus, OpenTelemetry, ArgoCD), and automate production operations with testing, automation, and infrastructure-as-code guardrails.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Cisco
Cisco
4 days ago

Senior Site Reliability Engineer, Production Engineer - ThousandEyes

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 12 hours agoStatus: Live

Job Summary

Design and manage large-scale, highly available distributed systems for the ThousandEyes SaaS platform. Collaborate with application teams to improve reliability, performance, and security, while building scalable operations tooling and incident response practices. Implement and maintain AWS cloud-native services, expand CNCF-based reliability stack (Kubernetes, Service Mesh, Prometheus, OpenTelemetry, ArgoCD), and automate production operations with testing, automation, and infrastructure-as-code guardrails.
Location: San Francisco
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Collaborate with software engineers to optimize architecture and services for availability, latency, performance, and reliability using cloud-native tools.
  • •Design and implement scalable operations tooling to support platform growth and scaling across multiple regions.
  • •Design, deploy, and maintain AWS cloud-native services that are elastic and resilient to failure.
  • •Participate in and improve 24x7 incident response and on-call rotation.
  • •Use and expand CNCF solutions like Kubernetes, Service Mesh, Prometheus, OpenTelemetry, and ArgoCD to increase platform reliability; automate production operations with guardrails and continuous operation.

Pay and Benefits

Salary: USD 165,000 - 241,400 annually
Perks:Health InsuranceDentalVision401kPaid ParentalLong-term DisabilityBasic LifePaid HolidaysFloating HolidayPaid LeaveSick TimeEquity

Key Requirements

  • •5+ years of experience in a related role.
  • •Proficiency in software development with languages such as Python or Go.
  • •Build and implement scalable, well-tested, security-focused solutions integrated across development and deployment.
  • •Strong understanding of Unix/Linux systems including kernel, system libraries, file systems, and client-server protocols.
  • •Knowledge of Site Reliability principles including incident response, change management, distributed systems, deployment strategies, and SLOs.
Experience:5+ yearsSaaSCloudDistributed systemsOperationsSecurity-focused development
Skills:OwnershipCommunicationDocumentationAttention to detail
Tech Stack:AWSPythonGoUnixLinuxKubernetesService MeshPrometheusOpenTelemetryArgoCDCNCFMicroservicesSLOsIncident responseChaos testingScale testingService mesh

Company Brief

Cisco
Global technology company that designs, manufactures, and sells networking hardware, telecommunications equipment, and high-technology services and products for enterprises, service providers, and governments worldwide.
Industry: Networking Equipment
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: San Jose, United States
Founded: 1984
Glassdoor
Glassdoor: 4.0
WebsiteLinkedInGlassdoor