Staff Site Reliability Engineer

bolt.new
Anywhere
Workplace: RemoteFull timeFunction: DevOps, Cloud & InfrastructureSkills: ["Technical leadership","Strategic execution","Systems thinking","Data-driven leadership","Verbal and written communication"]

Embed with product and platform teams from early design through launch readiness, ensuring systems are observable, scalable, and operable before production. Define production-readiness standards, create SLIs/SLOs and error budgets, and build reliable “golden paths” across AWS, GCP, and Azure using Terraform. Lead reliability improvements through incident management influence, blameless postmortems, and cross-team technical leadership while sharing an on-call rotation (about one week per month).

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
bolt.new
bolt.new
2 months ago

Staff Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 11 hours agoStatus: Live

Job Summary

Embed with product and platform teams from early design through launch readiness, ensuring systems are observable, scalable, and operable before production. Define production-readiness standards, create SLIs/SLOs and error budgets, and build reliable “golden paths” across AWS, GCP, and Azure using Terraform. Lead reliability improvements through incident management influence, blameless postmortems, and cross-team technical leadership while sharing an on-call rotation (about one week per month).
Location: Anywhere
Workplace: Remote
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Partner with development teams early in the project lifecycle to embed an SRE perspective into design, architecture reviews, and launch readiness.
  • •Establish and evolve production-readiness standards including design reviews, launch checklists, and operational acceptance criteria.
  • •Define measurable reliability goals using SLIs, SLOs, and error budgets to drive prioritization decisions.
  • •Create frameworks, tooling, and reliable “golden paths” across AWS, GCP, and Azure with Terraform.
  • •Influence incident management and blameless postmortems to turn operational signals into systematic improvements, while contributing to the shared on-call rotation.

Key Requirements

  • •Fluency across AWS, GCP, and Azure, with Terraform as an infrastructure-as-code layer.
  • •Experience as an SRE, production/platform engineer, or software engineer with deep reliability focus operating at scale.
  • •Comfort supporting and contributing to TypeScript services and Ruby on Rails backend services.
  • •Strong software engineering fundamentals and ability to write production-quality code with long-term maintainability.
  • •Demonstrated technical leadership and influence across team boundaries without formal authority.
Skills:Technical leadershipStrategic executionSystems thinkingData-driven leadershipVerbal and written communication
Languages:English
Tech Stack:AWSGCPAzureTerraformTypeScriptRuby on Rails

Company Brief

bolt.new
Provides a one‑click checkout platform, fraud prevention, and payment infrastructure to help online merchants increase conversion and securely process transactions across web and mobile channels.
Industry: Payments
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series E+
Headquarters: San Francisco, United States
Founded: 2014
WebsiteLinkedIn