Engineering Manager, Site Reliability Engineering

Replit
United States
Workplace: HybridFull timeUSD 250,000 - 325,000 annuallyFunction: DevOps, Cloud & InfrastructureSkills: ["Leadership","Coaching","Debugging","Cross-team collaboration","Performance judgment"]

Lead SRE across observability, incident management, load testing, performance engineering, and rollout infrastructure for Replit’s agentic software platform. Manage and grow a high-ownership distributed team while staying hands-on to debug production failures, improve reliability and performance, and reduce recovery time. Build telemetry and SLO capabilities, drive measurable reliability outcomes, and partner with service owners to implement fixes—using rigorous verification including AI-assisted coding.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Replit
Replit
2 days ago

Engineering Manager, Site Reliability Engineering

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 19 hours agoStatus: Live

Job Summary

Lead SRE across observability, incident management, load testing, performance engineering, and rollout infrastructure for Replit’s agentic software platform. Manage and grow a high-ownership distributed team while staying hands-on to debug production failures, improve reliability and performance, and reduce recovery time. Build telemetry and SLO capabilities, drive measurable reliability outcomes, and partner with service owners to implement fixes—using rigorous verification including AI-assisted coding.
Location: United States
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Manager level

Key Responsibilities

  • •Build and operate observability capabilities, including metrics, logs, traces, alerting, and meaningful SLOs for diagnosing issues and verifying improvements.
  • •Own incident management tooling and practices, coordinate cross-team response, and convert incident reviews into engineering improvements that reduce recovery time and repeat failures.
  • •Develop and maintain load/failure testing capabilities to validate critical paths, quantify headroom, and test recovery and production readiness with service owners.
  • •Lead deep performance engineering engagements using profiling, telemetry, and load tests to identify bottlenecks and deliver improvements with service owners.
  • •Lead and grow a high-ownership engineering team by coaching engineers, managing performance, hiring, and making distributed collaboration and coverage deliberate.

Pay and Benefits

Salary: USD 250,000 - 325,000 annually
Equity and Bonus:Equity
Perks:Health InsuranceDentalVisionLife Insurance401kPaid ParentalFlexible TimeWellness Stipend

Key Requirements

  • •Demonstrated engineering management experience leading and developing engineers, setting priorities and performance decisions, and hiring thoughtfully.
  • •Built and operated distributed systems or reliability platforms; can reason across deployment behavior, Kubernetes, telemetry, service dependencies, and recovery mechanisms.
  • •Proven ability to lead safe changes and performance investigations, using measurement to diagnose reliability/performance issues and validate fixes realistically.
  • •Ability to provide platform-level capabilities other teams can adopt; leads hands-on engagements while balancing reliability, performance, effort, and cost tradeoffs.
  • •Comfort coordinating incident responses and turning incident reviews into engineering improvements that prevent repeat failures.
Experience:Distributed systemsSRECloudObservability
Skills:LeadershipCoachingDebuggingCross-team collaborationPerformance judgment
Tech Stack:KubernetesTelemetryMetricsLogsTracesAlertingSLOsGitOpsHarnessArgoCDKargoOpenTelemetryGCPLoad testingProfilingAI coding

Company Brief

Replit
Provides a browser-based integrated development environment (IDE) and collaborative coding platform that lets developers write, run, and deploy code instantly across many languages and frameworks.
Industry: Developer Tools
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Headquarters: San Francisco, United States
Founded: 2016
WebsiteLinkedIn