Senior Site Reliability Engineer

Cisco
Bengaluru
Workplace: HybridFull timeFunction: DevOps, Cloud & InfrastructureExperience: 8-10 yearsEducation: bachelorsSkills: ["Communication","Collaboration","Technical leadership","Independent ownership","End-to-end initiative driving"]

Own and operate scalable, reliable, and cost-efficient infrastructure for the ThousandEyes Network Assurance Data Platform. Manage large-scale data workflows with Apache Airflow, AWS EMR, Spark, and Hadoop, and operate Amazon EKS/Kubernetes environments for containerized services and ML workloads. Build automation with Terraform and Python tooling, drive FinOps (visibility, allocation, forecasting, anomaly detection), and improve observability, incident response, and operational readiness across data and ML platforms.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Cisco
Cisco
1 week ago

Senior Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 11 hours agoStatus: Live
Reposted: similar role first listed 3 weeks ago

Job Summary

Own and operate scalable, reliable, and cost-efficient infrastructure for the ThousandEyes Network Assurance Data Platform. Manage large-scale data workflows with Apache Airflow, AWS EMR, Spark, and Hadoop, and operate Amazon EKS/Kubernetes environments for containerized services and ML workloads. Build automation with Terraform and Python tooling, drive FinOps (visibility, allocation, forecasting, anomaly detection), and improve observability, incident response, and operational readiness across data and ML platforms.
Location: Bengaluru
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Own and operate scalable, reliable, cost-efficient infrastructure for the Network Assurance Data Platform.
  • •Manage and optimize large-scale data workflows using Apache Airflow, AWS EMR, Spark, and Hadoop-based processing.
  • •Operate and improve Amazon EKS environments for containerized services and ML workloads.
  • •Build and maintain infrastructure automation using Terraform and infrastructure-as-code practices.
  • •Drive FinOps and improve observability, alerting, incident response, and operational readiness, including identifying performance bottlenecks and cost optimization opportunities.

Key Requirements

  • •Bachelor’s degree or higher in Engineering, Computer Science, or equivalent practical experience.
  • •8–10 years of relevant experience in Site Reliability Engineering, DevOps, cloud/platform engineering, data infrastructure, or production engineering.
  • •Strong experience operating production infrastructure on AWS, including production troubleshooting and reliability improvements.
  • •Strong experience with Apache Airflow for workflow orchestration and monitoring/troubleshooting.
  • •Strong experience with Terraform (infrastructure-as-code) and Python-based automation/integration tooling.
Experience:8-10 yearsSite reliability engineeringDevopsCloud infrastructureData infrastructureProduction engineeringMachine learning infrastructure
Education:Bachelor's
Skills:CommunicationCollaborationTechnical leadershipIndependent ownershipEnd-to-end initiative driving
Tech Stack:AWSAmazon EKSKubernetesApache AirflowAmazon EMRSparkHadoopTerraformPythonLinuxCloudWatchPrometheusGrafanaSplunkOpenSearchDatadogEBSEC2S3RDS

Company Brief

Cisco
Global technology company that designs, manufactures, and sells networking hardware, telecommunications equipment, and high-technology services and products for enterprises, service providers, and governments worldwide.
Industry: Networking Equipment
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: San Jose, United States
Founded: 1984
Glassdoor
Glassdoor: 4.0
WebsiteLinkedInGlassdoor