Site Reliability Engineer

Supabase
Anywhere
Workplace: RemoteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 7+ yearsSkills: ["Communication","Influencing without authority","Async collaboration","System thinking"]

Build reliability practices across Supabase by partnering with service teams to define SLOs/SLIs, error budgets, and operational readiness. Own and evolve the Operational Readiness Review (ORR) process, strengthen the incident-to-improvement pipeline, and act as the reliability expert for architecture reviews and resilience design. Reduce operational toil with automation, improve on-call practices, and track org-wide operational maturity to drive systemic remediation.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Supabase
Supabase
3 months ago

Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 2 hours agoStatus: Live

Job Summary

Build reliability practices across Supabase by partnering with service teams to define SLOs/SLIs, error budgets, and operational readiness. Own and evolve the Operational Readiness Review (ORR) process, strengthen the incident-to-improvement pipeline, and act as the reliability expert for architecture reviews and resilience design. Reduce operational toil with automation, improve on-call practices, and track org-wide operational maturity to drive systemic remediation.
Location: Anywhere
Workplace: Remote
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Partner with service teams to define SLIs and SLOs grounded in customer experience, and build error budget policies that guide engineering decisions.
  • •Own and evolve the Operational Readiness Review (ORR) process across observability, alerting, runbooks, capacity, and graceful degradation.
  • •Strengthen the incident-to-improvement pipeline by connecting postmortem learnings to operational readiness gaps and driving systemic fixes.
  • •Act as a reliability expert for architecture reviews, failure mode analysis, dependency mapping, and resilience design.
  • •Identify and quantify operational toil, advocate for automation, and help teams design sustainable on-call practices and reduce noise.

Pay and Benefits

Perks:Health InsuranceEquityLearning BudgetRemote Work

Key Requirements

  • •7+ years of experience in SRE, production engineering, or reliability-focused roles, including shaping SRE practices and driving adoption across engineering teams.
  • •A software engineering mindset—writing code and building tools, not only configuring systems.
  • •Hands-on experience defining and operationalizing SLOs/SLIs at scale, including error budget policies that influence engineering decisions.
  • •Deep experience with incident response, postmortem facilitation, and turning incident learnings into systemic improvements.
  • •Proficiency with cloud infrastructure (AWS preferred) and infrastructure-as-code (Pulumi preferred; Terraform/CDK acceptable).
Experience:7+ yearsMulti-tenant systems
Skills:CommunicationInfluencing without authorityAsync collaborationSystem thinking
Tech Stack:SLOSLIAWSPulumiTerraformCDKKubernetesOpenTelemetryVictoriaMetricsGrafanaPostgresError budgetOperational Readiness Review (ORR)PostmortemsDORA metrics

Company Brief

Supabase
Provides an open-source backend as a service built on PostgreSQL, offering real-time APIs, authentication, storage, and edge functions to help developers quickly build and scale applications.
Industry: Cloud Computing
Company Size: Large (251 to 1,000 employees)
Growth: Scaleup
Funding: Series B
Headquarters: New York, United States
Founded: 2020
WebsiteLinkedIn