Lead Site Reliability Engineer

Lumen
United States
Workplace: RemoteFull timeUSD 105,786 - 155,152 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 5+ yearsEducation: bachelorsSkills: ["Automation","Incident management","Blameless postmortems","Communication","Troubleshooting"]

Own reliability for the NaaS platform by driving observability, incident management, and automation. Build and tune dashboards, proactive alerting, and SLI/SLO/error budget practices, then lead blameless postmortems and follow-through to prevent recurrence. Automate CI/CD and cloud infrastructure with infrastructure-as-code, develop utilities to reduce toil, and apply AI/agentic workflows to deployment, administration, and investigations.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Lumen
Lumen
14 hours ago

Lead Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 2 hours agoStatus: Live
Reposted: similar role first listed 2 months ago

Job Summary

Own reliability for the NaaS platform by driving observability, incident management, and automation. Build and tune dashboards, proactive alerting, and SLI/SLO/error budget practices, then lead blameless postmortems and follow-through to prevent recurrence. Automate CI/CD and cloud infrastructure with infrastructure-as-code, develop utilities to reduce toil, and apply AI/agentic workflows to deployment, administration, and investigations.
Location: United States
Workplace: Remote
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure

Key Responsibilities

  • •Serve as subject matter expert for network automation platform applications, services, and hosting environments.
  • •Build and maintain the observability stack: instrument services, collect and curate metrics, and create dashboards/visualizations.
  • •Define and tune proactive alerting and champion SRE principles (SLIs, SLOs, error budgets) for resilient, fault-tolerant architecture.
  • •Participate in an on-call rotation; lead incident response, drive blameless postmortems, and ensure follow-up actions are completed.
  • •Automate deployment pipelines and cloud infrastructure provisioning/scaling/configuration using infrastructure-as-code; develop utilities and apply AI-assisted agentic workflows to reduce toil.

Pay and Benefits

Salary: USD 105,786 - 155,152 annually
Perks:Health InsuranceDentalVision401kGym MembershipLearning Budget

Key Requirements

  • •Bachelor’s degree (or equivalent) in engineering, computer science, or related field.
  • •8+ years in software development, systems engineering, and/or networking; 5+ years of related experience.
  • •Hands-on experience with at least one major cloud platform (AWS, Azure, or GCP), including compute, networking, and identity services.
  • •Strong automation and infrastructure-as-code skills with Terraform, Ansible, and Python.
  • •Experience with Kubernetes and modern observability/monitoring tools (e.g., Datadog, CloudWatch, Grafana, Prometheus), including dashboards and alerting.
Experience:5+ yearsNetworkingCloudInfrastructure-as-codeObservabilitySite reliability engineering
Education:Bachelor's
Skills:AutomationIncident managementBlameless postmortemsCommunicationTroubleshooting
Tech Stack:AWSAzureGCPTerraformAnsiblePythonKubernetesCI/CDInfrastructure as codeDatadogCloudWatchGrafanaPrometheusSLIsSLOsError budgetsNetwork automationMetrics dashboardsAlerting

Company Brief

Lumen
Provides integrated communications, network, edge cloud, and security services to enterprise, government, and carrier customers, operating a global fiber network and delivering connectivity and managed IT solutions.
Industry: Telecommunications
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Monroe, United States
WebsiteLinkedIn