Incident Operations Lead (EMEA/AMER)

Alpaca
United States
Workplace: RemoteFull timeFunction: Business OperationsExperience: 5+ yearsSkills: ["Leadership","Communication","Writing","Decision-making","Process discipline"]

Lead the incident command function for critical outages across Alpaca’s brokerage and API services. Build the 24x7, follow-the-sun coverage and roster, recruit and certify Incident Commanders, and drive a blameless review culture. Own severity maturity, escalation and communication paths, KPIs for response/mitigation, and the incident-to-learning loop. Build the process as productized, versioned artifacts and automate incident workflows with AI agents.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Alpaca
Alpaca
1 day ago

Incident Operations Lead (EMEA/AMER)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 3 hours agoStatus: Live

Job Summary

Lead the incident command function for critical outages across Alpaca’s brokerage and API services. Build the 24x7, follow-the-sun coverage and roster, recruit and certify Incident Commanders, and drive a blameless review culture. Own severity maturity, escalation and communication paths, KPIs for response/mitigation, and the incident-to-learning loop. Build the process as productized, versioned artifacts and automate incident workflows with AI agents.
Location: United States
Workplace: Remote
Employment Type: Full time
Job Function: Business Operations
Seniority: Manager level

Key Responsibilities

  • •Build the incident command function and stand up 24x7 command, including recruiting and certifying Incident Commanders and maintaining a follow-the-sun roster.
  • •Own incident response design: severity maturity, escalation paths, unanswered-page thresholds, and service ownership catalog accuracy.
  • •Connect engineering/technical support and partner communications with uninterrupted focus and a continuous, accurate update cadence.
  • •Run blameless reviews and post-incident packages with SRE and reliability programme management, then publish concise learnings across engineering.
  • •Define and manage end-to-end incident KPIs (time to respond/mitigate), including data hygiene and overdue review reporting; automate the operating model with AI workflows and agents.

Pay and Benefits

Equity and Bonus:Equity
Perks:Health InsuranceEquityHome OfficeStipend

Key Requirements

  • •Have stood up an incident command or major-incident function, including owning the severity model and driving adoption across teams.
  • •5+ years in production engineering, SRE, or technical operations, including hands-on command of high-severity incidents.
  • •Have led a distributed team across time zones and run a 24x7 rotation.
  • •Defend a severity call and build reliability metrics people trust, improving both numbers and reality.
  • •Write effectively enough that process documents are used, and can brief executives mid-incident while bridging calm under pressure.
Experience:5+ yearsFintechSaaSSRETechnical operationsRegulated financial services
Skills:LeadershipCommunicationWritingDecision-makingProcess discipline
Languages:English
Tech Stack:AIAgentic automation

Company Brief

Alpaca
Provides developer-first APIs and brokerage infrastructure for algorithmic and automated stock trading, enabling fintechs and developers to build trading applications with commission-free market access and custody services.
Industry: Fintech Infrastructure
Company Size: Medium (51 to 250 employees)
Growth: Scaleup
Funding: Series C
Headquarters: San Mateo, United States
Founded: 2015
WebsiteLinkedIn