Staff Site Reliability Engineer

Replit
United States
Workplace: RemoteFull timeUSD 250,000 - 325,000 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 8-10 yearsSkills: ["Communication","Mentoring","Incident response","Critical thinking","Collaboration"]

Join the Site Reliability Engineering team to ensure Replit’s infrastructure is reliable, scalable, and high-performing for millions of developers. You’ll design observability (metrics, logging, tracing), set and track SLOs/SLIs, lead incident response and blameless post-mortems, and automate operations with infrastructure-as-code and CI/CD. You’ll also optimize Kubernetes/GCP performance, debug distributed systems, and mentor engineers to make reliability a core value.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Replit
Replit
21 hours ago

Staff Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 16 hours agoStatus: Live

Job Summary

Join the Site Reliability Engineering team to ensure Replit’s infrastructure is reliable, scalable, and high-performing for millions of developers. You’ll design observability (metrics, logging, tracing), set and track SLOs/SLIs, lead incident response and blameless post-mortems, and automate operations with infrastructure-as-code and CI/CD. You’ll also optimize Kubernetes/GCP performance, debug distributed systems, and mentor engineers to make reliability a core value.
Location: United States
Workplace: Remote
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Sr. Manager level

Key Responsibilities

  • •Design, build, and lead comprehensive monitoring, logging, and tracing with dashboards for real-time system visibility.
  • •Define and drive reliability standards by implementing SLOs/SLIs and systems to track and report them.
  • •Lead incident management and response, run blameless post-mortems, and improve runbooks and automation to reduce MTTR.
  • •Drive automation and infrastructure-as-code efforts, including Terraform/Pulumi and CI/CD pipelines, to eliminate toil.
  • •Optimize performance on large-scale Kubernetes/GCP deployments and debug distributed-system issues to implement long-term fixes.

Pay and Benefits

Salary: USD 250,000 - 325,000 annually
Equity and Bonus:Equity
Perks:Health InsuranceDentalVisionLife Insurance401kPaid LeaveDisabilityWellness Stipend

Key Requirements

  • •8-10 years of experience in Site Reliability Engineering or similar roles (e.g., DevOps, Systems Engineering, Infrastructure Engineering).
  • •Strong programming skills in Python or Go, with a track record of writing high-quality, well-tested code.
  • •Deep understanding of distributed systems and production service design, scaling, and maintenance.
  • •Deep experience with Kubernetes and cloud-native technologies, plus production observability (metrics, logging, tracing).
  • •Experience leading incident response for complex systems, with infrastructure-as-code knowledge (e.g., Terraform, Pulumi) and strong communication/mentoring skills.
Experience:8-10 yearsSREDevOpsKubernetesCloud-nativeDistributed systemsObservability
Skills:CommunicationMentoringIncident responseCritical thinkingCollaboration
Tech Stack:PythonGoKubernetesDockerGCPTerraformPulumiCI/CDObservabilityMetricsLoggingTracingPrometheusGrafanaDatadogOpenTelemetrySLOSLIDashboardsRunbooks

Company Brief

Replit
Provides a browser-based integrated development environment (IDE) and collaborative coding platform that lets developers write, run, and deploy code instantly across many languages and frameworks.
Industry: Developer Tools
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Headquarters: San Francisco, United States
Founded: 2016
WebsiteLinkedIn