Site Reliability Engineer

WorkOS
United States, Canada
Workplace: RemoteFull timeUSD 175,000 - 275,000 annuallyFunction: DevOps, Cloud & InfrastructureSkills: ["Systems thinking","Debugging","Independent work","Ownership"]

Join the Site Reliability Engineering team to keep the WorkOS platform fast, reliable, and resilient at scale. You’ll design and evolve reliability systems and tooling, define SLIs/SLOs, and collaborate with product and infrastructure teams on production readiness and observability. Build and optimize TypeScript backend systems, lead incident response and postmortems, and create automation to operate and scale services—while participating in an on-call rotation.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
WorkOS
WorkOS
7 months ago

Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 11 hours agoStatus: Live

Job Summary

Join the Site Reliability Engineering team to keep the WorkOS platform fast, reliable, and resilient at scale. You’ll design and evolve reliability systems and tooling, define SLIs/SLOs, and collaborate with product and infrastructure teams on production readiness and observability. Build and optimize TypeScript backend systems, lead incident response and postmortems, and create automation to operate and scale services—while participating in an on-call rotation.
Location: United States, Canada
Workplace: Remote
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure

Key Responsibilities

  • •Design and evolve systems, tooling, and processes to improve reliability and performance.
  • •Collaborate with product and infrastructure teams to ensure services are production-ready, observable, and resilient to failure.
  • •Define and measure SLIs and SLOs to guide reliability improvements.
  • •Write and optimize backend systems in TypeScript with a focus on performance, maintainability, and graceful degradation.
  • •Improve incident response, lead postmortems, drive follow-through on reliability risks, and participate in on-call.

Pay and Benefits

Salary: USD 175,000 - 275,000 annually
Equity and Bonus:Equity
Perks:Health InsuranceDentalVisionParental LeavePaid Leave401k MatchingRemote WorkWellness StipendPaid Holidays

Key Requirements

  • •Experience operating and scaling production systems in cloud environments.
  • •Familiarity with service reliability practices including monitoring, alerting, incident response, and root cause analysis.
  • •Comfort working across infrastructure layers such as compute, networking, storage, and observability tooling.
  • •Strong debugging and systems thinking skills to trace issues across services and layers.
  • •Ability to work independently, take ownership, and drive projects from problem discovery through resolution.
Experience:CloudProduction systems
Skills:Systems thinkingDebuggingIndependent workOwnership
Tech Stack:AWSTypeScriptKubernetesPrometheusGrafanaDatadogOpenTelemetry

Company Brief

WorkOS
Provides developer-facing enterprise-ready identity and authentication infrastructure (SSO, SCIM, Audit Logs, and Directory Sync) via APIs to help apps integrate enterprise features quickly and securely.
Industry: Cybersecurity
Company Size: Medium (51 to 250 employees)
Growth: Growth Stage Startup
Funding: Series B
Headquarters: San Francisco, United States
Founded: 2018
WebsiteLinkedIn