Senior Site Reliability Engineer
North America
Workplace: RemoteFull timeFunction: DevOps, Cloud & InfrastructureSkills: ["Incident response","Structured debugging","Written communication","Verbal communication","Mentoring"]Build and operate Alpaca’s brokerage infrastructure as a Site Reliability Engineer, ensuring the platform is reliable, observable, and scalable. Own day-to-day production operations including on-call, incident response, postmortems, and reliability follow-ups. Define SLIs/SLOs and error budgets, improve observability across metrics/logs/traces/alerting, and ship infrastructure via GitOps on cloud and Kubernetes. Take meaningful ownership of PostgreSQL reliability, including tuning, migrations, HA/DR, and CDC.
Loading
Loading job details...
Preparing the role view and application actions.

