Senior Site Reliability Engineer, Production Engineer - ThousandEyes(Hybrid)

Cisco
San Francisco, Seattle, Austin, New York
Workplace: HybridFull timeUSD 167,700 - 245,200 annuallyFunction: DevOps, Cloud & InfrastructureEducation: bachelorsSkills: ["Communication","Documentation","Ownership","Attention to detail"]

Lead the design and management of large-scale, highly available distributed systems for a SaaS platform. Deploy resilient AWS cloud-native services and standardize operations across a multi-region, microservice architecture using Kubernetes, Prometheus, and service mesh. Build automation for service operations, including deployment and chaos testing, and apply “everything-as-code” to create guardrails. Tackle operational obstacles to keep systems scalable under high daily data volumes.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Cisco
Cisco
1 month ago

Senior Site Reliability Engineer, Production Engineer - ThousandEyes(Hybrid)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 8 hours agoStatus: Live

Job Summary

Lead the design and management of large-scale, highly available distributed systems for a SaaS platform. Deploy resilient AWS cloud-native services and standardize operations across a multi-region, microservice architecture using Kubernetes, Prometheus, and service mesh. Build automation for service operations, including deployment and chaos testing, and apply “everything-as-code” to create guardrails. Tackle operational obstacles to keep systems scalable under high daily data volumes.
Location: San Francisco, Seattle, Austin, New York
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Lead design and management of large-scale, highly available distributed systems for the ThousandEyes SaaS platform.
  • •Collaborate with application teams to ensure performance, reliability, and security.
  • •Deploy and operate resilient AWS cloud-native services across a multi-region, microservice-based architecture.
  • •Develop automation for service operations, including deployment, chaos testing, and “everything-as-code” strategies.
  • •Identify and resolve operational obstacles to keep systems scalable under high daily data volumes.

Pay and Benefits

Salary: USD 167,700 - 245,200 annually
Equity and Bonus:Equity
Perks:Health InsuranceDentalVision401kPaid ParentalPaid HolidaysPaid LeaveSick TimeLife InsuranceLong-term DisabilityEquity

Key Requirements

  • •Bachelors + 7 years relevant experience, or Masters + 4 years, or PhD + 1 year, or equivalent related work experience.
  • •Proficiency in software development using Python or Go.
  • •Ability to build and implement scalable, well-tested, security-focused solutions integrated throughout development and deployment.
  • •Strong understanding of Unix/Linux systems, including kernel, system libraries, file systems, and client-server protocols.
  • •Knowledge of Site Reliability principles including Incident Response, Change Management, Distributed Systems, Deployment Strategies, and SLOs.
Experience:SaaSCloudDistributed systemsMicroservicesEnterprise platforms
Education:Bachelor's
Skills:CommunicationDocumentationOwnershipAttention to detail
Tech Stack:PythonGoAWSKubernetesPrometheusService MeshUnix/Linux

Company Brief

Cisco
Global technology company that designs, manufactures, and sells networking hardware, telecommunications equipment, and high-technology services and products for enterprises, service providers, and governments worldwide.
Industry: Networking Equipment
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: San Jose, United States
Founded: 1984
Glassdoor
Glassdoor: 4.0
WebsiteLinkedInGlassdoor