Software Engineering Tech Lead (SRE + AI)

Cisco
London
Workplace: HybridFull timeFunction: Software EngineeringExperience: 3+ yearsEducation: bachelorsSkills: ["Technical leadership","Mentorship","Collaboration","Communication","Reliability focus"]

Drive the technical roadmap for an AI-powered Production Intelligence platform by combining Site Reliability Engineering with agentic AI. Lead architecture and implementation for AI-assisted observability, automated incident response, and self-healing SaaS infrastructure. Design telemetry ingestion and correlation using OpenTelemetry traces, logs, and runbooks to improve MTTD/MTTR, while building safe, guardrailed remediation workflows and mentoring engineers across global teams.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Cisco
Cisco
1 day ago

Software Engineering Tech Lead (SRE + AI)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 17 hours agoStatus: Live

Job Summary

Drive the technical roadmap for an AI-powered Production Intelligence platform by combining Site Reliability Engineering with agentic AI. Lead architecture and implementation for AI-assisted observability, automated incident response, and self-healing SaaS infrastructure. Design telemetry ingestion and correlation using OpenTelemetry traces, logs, and runbooks to improve MTTD/MTTR, while building safe, guardrailed remediation workflows and mentoring engineers across global teams.
Location: London
Workplace: Hybrid
Employment Type: Full time
Job Function: Software Engineering
Seniority: Sr. Manager level

Key Responsibilities

  • •Define the technical roadmap and architecture for AI-assisted observability, automated incident response, and self-healing cloud infrastructure.
  • •Design and build production-grade AI agents, MCP tool integrations, and deterministic evaluation pipelines for operational decision support.
  • •Architect ingestion and correlation pipelines across distributed logs, metrics, OpenTelemetry traces, change events, and runbooks to accelerate MTTD and MTTR.
  • •Develop proactive anomaly detection and Human-in-the-Loop remediation workflows with safety, security, and quality guardrails.
  • •Partner with application and infrastructure teams on SLIs/SLOs, error budgets, and lead deep-dive post-incident reviews (PIRs).

Key Requirements

  • •Bachelor’s degree + 8 years of related experience (or Master’s + 6 years, or PhD + 3 years) in Computer Science, Software Engineering, or a related technical field.
  • •Proven record as a Technical Lead or Lead SRE/Software Engineer delivering distributed, high-availability SaaS platforms at scale.
  • •Strong proficiency in Python, Go, Java, or C++ with experience designing microservices, APIs, and production automation.
  • •Deep experience with Kubernetes, Docker, and container orchestration in large-scale multi-cluster environments.
  • •SRE practices experience: SLI/SLO design, observability platforms (metrics/logs/traces), incident management, and automated RCA.
Experience:3+ yearsSaaSMicroservicesCloud infrastructureObservabilitySRE
Education:Bachelor's in Computer Science
Skills:Technical leadershipMentorshipCollaborationCommunicationReliability focus
Tech Stack:PythonGoJavaC++MicroservicesAPIsProduction automationKubernetesDockerContainer orchestrationSLI/SLOObservabilityMetricsLogsTracesIncident managementAutomated RCAOpenTelemetryPrometheusGrafana

Company Brief

Cisco
Global technology company that designs, manufactures, and sells networking hardware, telecommunications equipment, and high-technology services and products for enterprises, service providers, and governments worldwide.
Industry: Networking Equipment
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: San Jose, United States
Founded: 1984
Glassdoor
Glassdoor: 4.0
WebsiteLinkedInGlassdoor