Senior Site Reliability Engineer

Alpaca
North America
Workplace: RemoteFull timeFunction: DevOps, Cloud & InfrastructureSkills: ["Incident response","Structured debugging","Written communication","Verbal communication","Mentoring"]

Build and operate Alpaca’s brokerage infrastructure as a Site Reliability Engineer, ensuring the platform is reliable, observable, and scalable. Own day-to-day production operations including on-call, incident response, postmortems, and reliability follow-ups. Define SLIs/SLOs and error budgets, improve observability across metrics/logs/traces/alerting, and ship infrastructure via GitOps on cloud and Kubernetes. Take meaningful ownership of PostgreSQL reliability, including tuning, migrations, HA/DR, and CDC.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Alpaca
Alpaca
6 days ago

Senior Site Reliability Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 4 hours agoStatus: Live

Job Summary

Build and operate Alpaca’s brokerage infrastructure as a Site Reliability Engineer, ensuring the platform is reliable, observable, and scalable. Own day-to-day production operations including on-call, incident response, postmortems, and reliability follow-ups. Define SLIs/SLOs and error budgets, improve observability across metrics/logs/traces/alerting, and ship infrastructure via GitOps on cloud and Kubernetes. Take meaningful ownership of PostgreSQL reliability, including tuning, migrations, HA/DR, and CDC.
Location: North America
Workplace: Remote
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Operate production day-to-day including on-call, incident response, postmortems, and follow-ups to close the loop.
  • •Own reliability practice by defining/refining SLIs, SLOs, and error budgets and ensuring product teams live within them.
  • •Strengthen observability across metrics, logs, traces, and alerting.
  • •Ship infrastructure through code using a GitOps workflow for cloud resources and Kubernetes workloads.
  • •Own PostgreSQL reliability, including performance tuning, schema/migration review, online migrations, HA/DR, and CDC pipelines.

Pay and Benefits

Perks:Health InsuranceEquity

Key Requirements

  • •4+ years in SRE, DevOps, Platform/Infrastructure, or backend engineering with significant production operations ownership.
  • •Hands-on production experience operating services on Kubernetes and shipping infrastructure as code in a GitOps workflow.
  • •Strong PostgreSQL production knowledge, including query plans, pg_stat_*, indexing/schema trade-offs, and safe online migrations.
  • •Comfort with cloud networking fundamentals (VPCs, routing, L4/L7 load balancing, DNS, TLS) and debugging cross-service connectivity.
  • •Proficiency with Go or Python and strong written and verbal communication; practiced incident response with postmortems that drive change.
Skills:Incident responseStructured debuggingWritten communicationVerbal communicationMentoring
Languages:English
Tech Stack:KubernetesGitOpsPostgreSQLLinuxGoPythonObservabilityMetricsLogsTracesAlertingVPCsDNSTLSRabbitMQKafkaRedpandaPgxGormSqlc

Company Brief

Alpaca
Provides developer-first APIs and brokerage infrastructure for algorithmic and automated stock trading, enabling fintechs and developers to build trading applications with commission-free market access and custody services.
Industry: Fintech Infrastructure
Company Size: Medium (51 to 250 employees)
Growth: Scaleup
Funding: Series C
Headquarters: San Mateo, United States
Founded: 2015
WebsiteLinkedIn