Senior Site Reliability Engineer

Replit
United States
Workplace: RemoteFull timeUSD 210,000 - 275,000 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 4-8 yearsSkills: ["Problem-solving","Self-direction","Communication","Continuous learning","Automation"]

Own reliability for a platform used by millions of developers, partnering with engineering to improve scalability, performance, and availability. Build observability with monitoring, alerting, dashboards, and logging. Automate operations using infrastructure as code and CI/CD, and implement self-healing workflows. Define SLOs/SLIs, lead incident response with post-mortems and runbooks, and drive continuous performance optimization across global regions.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Replit
Replit
21 hours ago

Senior Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 16 hours agoStatus: Live

Job Summary

Own reliability for a platform used by millions of developers, partnering with engineering to improve scalability, performance, and availability. Build observability with monitoring, alerting, dashboards, and logging. Automate operations using infrastructure as code and CI/CD, and implement self-healing workflows. Define SLOs/SLIs, lead incident response with post-mortems and runbooks, and drive continuous performance optimization across global regions.
Location: United States
Workplace: Remote
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Design and implement observability by building monitoring and alerting systems, dashboards, metrics, and logging for real-time system health visibility.
  • •Drive automation and infrastructure as code using Terraform, Ansible, or Pulumi, and build CI/CD pipelines for consistent deployments and self-healing responses.
  • •Establish SLOs and SLIs with product and engineering teams, track and report these metrics to maintain reliability while supporting faster innovation.
  • •Lead incident management and response, conduct post-mortems, maintain runbooks, and implement process/tooling to reduce MTTR.
  • •Optimize performance by identifying bottlenecks, improving capacity planning and resource utilization, reducing latency, and enhancing efficiency across global regions.

Pay and Benefits

Salary: USD 210,000 - 275,000 annually
Equity and Bonus:Equity
Perks:Health InsuranceDentalVisionLife Insurance401kPaid LeaveWellness Stipend

Key Requirements

  • •4-8 years of experience in Site Reliability Engineering or similar roles (DevOps, Systems Engineering, Infrastructure Engineering).
  • •Strong programming skills for automation (Python, Go, or similar).
  • •Deep understanding of distributed systems and operational reliability at scale.
  • •Experience with Kubernetes and cloud-native technologies.
  • •Experience implementing monitoring/observability and leading incident response.
Experience:4-8 years
Skills:Problem-solvingSelf-directionCommunicationContinuous learningAutomation
Tech Stack:TerraformAnsiblePulumiCI/CDObservabilitySLOsSLIsRunbooksMTTRPythonGoKubernetesPrometheusGrafanaDatadogGoogle Cloud Platform (GCP)

Company Brief

Replit
Provides a browser-based integrated development environment (IDE) and collaborative coding platform that lets developers write, run, and deploy code instantly across many languages and frameworks.
Industry: Developer Tools
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Headquarters: San Francisco, United States
Founded: 2016
WebsiteLinkedIn