Senior Software Engineer, Robinhood Command Center

Robinhood
New York, Menlo Park
Workplace: OnsiteFull timeFunction: Software EngineeringExperience: 5+ yearsSkills: ["OpenTelemetry","Prometheus","Grafana","Observability","Fault-tolerant architecture","Multi-region","Multi-cluster","Capacity planning","Failover"]

Senior Engineering role building and leading Robinhood Command Center reliability initiatives. You’ll drive incident leadership, observability strategy, and tooling across Robinhood’s distributed infrastructure, coordinate cross-functional responders, maintain global dashboards, and improve MTTR/MTTD. This is a founding RCC position focused on incident response processes, governance, and learning to reduce customer impact.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Robinhood
Robinhood
2 months ago

Senior Software Engineer, Robinhood Command Center

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 8 hours agoStatus: Live

Job Summary

Senior Engineering role building and leading Robinhood Command Center reliability initiatives. You’ll drive incident leadership, observability strategy, and tooling across Robinhood’s distributed infrastructure, coordinate cross-functional responders, maintain global dashboards, and improve MTTR/MTTD. This is a founding RCC position focused on incident response processes, governance, and learning to reduce customer impact.
Location: New York, Menlo Park
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering
Seniority: Manager level

Key Responsibilities

  • •Serve as a senior technical leader driving the long-term reliability and observability strategy across Robinhood’s infrastructure
  • •Partner closely across many different types of engineers to raise the bar for operational excellence and incident response
  • •Lead incident mitigation efforts by coordinating service owners, facilitating time-sensitive decisions like rollbacks, traffic shifts, and maintaining a clear source of truth during active incidents
  • •Develop and maintain incident management processes and procedures to ensure timely resolution and minimize customer impact
  • •Own incident discovery at the company level by defining and maintaining global dashboards and alerts tied to critical user journeys (CUJs), availability, and business-impact metrics

Pay and Benefits

Equity and Bonus:Equity
Perks:Health Insurance401kEquityAnnual BonusLife InsuranceDisability InsuranceParential LeaveWellness Stipend

Key Requirements

  • •5+ years of software engineering experience, including significant experience operating production systems
  • •2+ years focused on reliability engineering, infrastructure, distributed systems, or production operations
  • •Hands-on experience serving in incident leadership roles (e.g., IMOC, incident commander, primary oncall)
  • •Strong communication and cross-functional collaboration skills, especially during high-severity incidents
  • •Deep knowledge of systems reliability, observability frameworks, and fault-tolerant architecture design
Experience:5+ yearsReliabilityObservabilityDistributed systemsProduction operations
Skills:OpenTelemetryPrometheusGrafanaObservabilityFault-tolerant architectureMulti-regionMulti-clusterCapacity planningFailover
Languages:English
Tech Stack:OpenTelemetryPrometheusGrafanaDistributed systemsObservabilityMulti-regionFailoverCapacity planning

Company Brief

Robinhood
Commission-free trading platform offering stocks, ETFs, options, and cryptocurrencies, plus cash management and retirement products. Aims to democratize finance through easy-to-use mobile and web apps for retail investors.
Industry: Trading Platforms
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Menlo Park, United States
Founded: 2013
WebsiteLinkedIn