Staff Site Reliability Engineer - Release Engineering

Plaid
New York, San Francisco
Workplace: HybridFull timeUSD 207,600 - 273,600 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 8+ yearsSkills: ["Technical leadership","Organizational change","Cross-team influence","Sound technical judgment","High-stakes decision-making"]

Own and scale Plaid’s reliability practices for the Release Engineering path from merge to production. Architect SLO and error-budget programs, drive progressive delivery and automated safety gates, and ensure new products are production-ready. Partner with product, platform, and infrastructure teams to turn complex production needs into self-service tooling. Lead incident response and post-mortem-driven improvements while scaling safety nets for faster, AI-assisted development.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Plaid
Plaid
2 months ago

Staff Site Reliability Engineer - Release Engineering

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 11 hours agoStatus: Live

Job Summary

Own and scale Plaid’s reliability practices for the Release Engineering path from merge to production. Architect SLO and error-budget programs, drive progressive delivery and automated safety gates, and ensure new products are production-ready. Partner with product, platform, and infrastructure teams to turn complex production needs into self-service tooling. Lead incident response and post-mortem-driven improvements while scaling safety nets for faster, AI-assisted development.
Location: New York, San Francisco
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Expand reliability standards across product engineering and turn infrastructure foundations into lasting operational habits and tooling.
  • •Architect and manage SLO and error-budget frameworks so teams can use reliability data for product and release decisions.
  • •Promote progressive delivery and automated safety gates to maintain high velocity while protecting production stability.
  • •Guide teams toward production readiness using observability, incident response, and scalable deployment health.
  • •Lead critical incident response and ensure post-mortem actions result in permanent platform improvements.

Pay and Benefits

Salary: USD 207,600 - 273,600 annually
Equity and Bonus:Equity
Perks:Health InsuranceDentalVision401k

Key Requirements

  • •Over 8 years of professional experience in backend systems, SRE, or platform engineering roles.
  • •Proven experience designing reliability programs (e.g., service maturity models or SLI frameworks) with cross-team adoption.
  • •Direct experience building or operating canary rollout systems, metric-gated analysis, or automated rollback infrastructure.
  • •Strong software development proficiency, with a preference for Go or similar systems languages.
  • •Experience with Kubernetes, service mesh technologies, Prometheus, or ArgoCD is a strong asset.
Experience:8+ yearsBackend systemsSREPlatform engineering
Skills:Technical leadershipOrganizational changeCross-team influenceSound technical judgmentHigh-stakes decision-making
Tech Stack:GoKubernetesService meshPrometheusArgoCDSLOError-budgetCanary rolloutMetric-gated analysisAutomated rollbackZero-touch deploymentProgressive rolloutsObservabilityIncident response

Company Brief

Plaid
Provides APIs that enable applications to connect with users’ bank accounts, verify financial data, and power payment and account verification workflows for fintechs and financial services companies.
Industry: Fintech Infrastructure
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series D
Headquarters: San Francisco, United States
Founded: 2013
WebsiteLinkedIn