Software Engineering Technical Leader - SRE + Kubernetes + Cloud + Automation + Incident Management + AI-first (10-14 Years)

Cisco
Bengaluru
Workplace: HybridFull timeFunction: DevOps, Cloud & InfrastructureExperience: 7-14 yearsEducation: bachelorsSkills: ["Ownership","Collaboration","Innovation","Written communication"]

Lead reliability engineering for critical collaboration services across cloud and hybrid environments, ensuring scalability, resiliency, performance, and security. Own deployment and on-call operations, monitoring and alerting against SLOs/SLAs, and manage incident response with root-cause analysis and remediation. Improve reliability through automation, CI/CD pipeline and tooling enhancements, and observability-driven capacity planning and operational best practices.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Cisco
Cisco
3 days ago

Software Engineering Technical Leader - SRE + Kubernetes + Cloud + Automation + Incident Management + AI-first (10-14 Years)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 5 hours agoStatus: Live

Job Summary

Lead reliability engineering for critical collaboration services across cloud and hybrid environments, ensuring scalability, resiliency, performance, and security. Own deployment and on-call operations, monitoring and alerting against SLOs/SLAs, and manage incident response with root-cause analysis and remediation. Improve reliability through automation, CI/CD pipeline and tooling enhancements, and observability-driven capacity planning and operational best practices.
Location: Bengaluru
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure

Key Responsibilities

  • •Own the deployment and operation of critical collaboration services across cloud and hybrid environments, driving reliability and scalability.
  • •Design, evolve, and optimize CI/CD pipelines and automation, including AI-first tooling for deployment, monitoring, and incident response.
  • •Lead incident response for complex production issues, perform root cause analysis, and drive systemic reliability and performance improvements.
  • •Use observability data to guide capacity planning, scaling strategies, and resource optimization across services.
  • •Define and champion operational best practices, documentation standards, and a culture of reliability and operational excellence.

Key Requirements

  • •Bachelor’s degree in Computer Science, Engineering, or related field (or equivalent experience) with 7–13 years in Site Reliability Engineering, Cloud Operations, or Systems Engineering.
  • •Hands-on experience operating production services using Docker and Kubernetes in cloud or hybrid environments.
  • •Proficiency in one or more programming/scripting languages (Python, Go, Bash) to build automation and operational tooling.
  • •Experience with monitoring, observability, and incident response, including on-call participation and post-incident reviews.
  • •Working knowledge of Linux systems, networking, distributed systems, CI/CD pipelines, infrastructure-as-code, and Git-based workflows.
Experience:7-14 yearsSaaSCloud operationsSite reliability engineeringDistributed systems
Education:Bachelor's in Computer Science, Engineering, or related field
Skills:OwnershipCollaborationInnovationWritten communication
Tech Stack:DockerKubernetesCloudCI/CDAI-first toolingPythonGoBashLinuxNetworkingDistributed systemsInfrastructure-as-codeGitObservabilityMonitoringIncident responseDisaster recovery

Company Brief

Cisco
Global technology company that designs, manufactures, and sells networking hardware, telecommunications equipment, and high-technology services and products for enterprises, service providers, and governments worldwide.
Industry: Networking Equipment
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: San Jose, United States
Founded: 1984
Glassdoor
Glassdoor: 4.0
WebsiteLinkedInGlassdoor