Senior Infrastructure Engineer, SRE

Rocket Money
San Francisco, Washington, New York
Workplace: RemoteFull timeUSD 150,000 - 185,000 annuallyFunction: DevOps, Cloud & InfrastructureSkills: ["Reliability","Incident management","Automation","Problem-solving","Collaboration"]

Lead reliability and operational evolution of a platform running hundreds of production services that process billions of transactions. Own and improve SLIs/SLOs and error budgets, disaster recovery strategy, and incident practices. Evolve observability (metrics, tracing, logs) and strengthen standards with instrumentation, while partnering with product engineering teams to help them own and operate their services. Participate in infrastructure build-outs and on-call rotation.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Rocket Money
Rocket Money
3 days ago

Senior Infrastructure Engineer, SRE

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 1 hour agoStatus: Live

Job Summary

Lead reliability and operational evolution of a platform running hundreds of production services that process billions of transactions. Own and improve SLIs/SLOs and error budgets, disaster recovery strategy, and incident practices. Evolve observability (metrics, tracing, logs) and strengthen standards with instrumentation, while partnering with product engineering teams to help them own and operate their services. Participate in infrastructure build-outs and on-call rotation.
Location: San Francisco, Washington, New York
Workplace: Remote
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Build and improve reliability and resiliency of systems and services running in production.
  • •Establish and maintain SLIs, SLOs, and error budgets, reviewing them regularly with service owners.
  • •Own and evolve disaster recovery strategy, including recovery objectives, failover/restore paths, and regular exercises.
  • •Evolve observability platform and standards across metrics, tracing, and logs, including alert quality and instrumentation.
  • •Strengthen incident practice (paging thresholds, runbooks, postmortem follow-through) and contribute to day-to-day cloud infrastructure work with a shared on-call rotation.

Pay and Benefits

Salary: USD 150,000 - 185,000 annually
Perks:Health InsuranceDentalVision401kPaid LeaveCommuter Benefits

Key Requirements

  • •5+ years of hands-on cloud or infrastructure engineering with substantial time on reliability and production operations at scale.
  • •Defined SLIs and SLOs for real production services, with examples of improvements and outcomes.
  • •Hands-on observability experience in production; Datadog is strongly preferred.
  • •Comfort writing production code and internal tooling/automation (Python, Go, TypeScript, or similar).
  • •Experience with disaster recovery plans, including failover/restore steps and drills that prove they work.
Experience:Production operationsCloud infrastructureReliability engineeringDisaster recoveryObservability
Skills:ReliabilityIncident managementAutomationProblem-solvingCollaboration
Languages:English
Tech Stack:PythonGoTypeScriptTerraformAWSDatadog

Company Brief

Rocket Money
Provides a personal finance app that helps users track spending, manage subscriptions, automate savings, and lower bills through budgeting tools and financial insights to simplify everyday money management.
Industry: Neobanking
Company Size: Medium (51 to 250 employees)
Growth: Growth Stage Startup
Headquarters: Los Angeles, United States
Founded: 2015
WebsiteLinkedIn