Senior DevOps Engineer

Alpaca
United States
Workplace: RemoteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 5+ yearsSkills: ["Incident management","Structured debugging","Written communication","Blameless post-mortems","Accountability"]

Design, build, and operate cloud infrastructure on GCP that powers Alpaca’s global, trading-critical systems. Own Infrastructure-as-Code with Terraform and follow GitOps, building CI/CD pipelines for safe IaC changes. Improve observability across Prometheus, Thanos, Grafana, Loki, Tempo, and Alertmanager, operate GKE clusters with Helm, and run operator-level database and message broker operations. Participate in follow-the-sun on-call and embed SRE practices.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Alpaca
Alpaca
1 day ago

Senior DevOps Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 4 hours agoStatus: Live

Job Summary

Design, build, and operate cloud infrastructure on GCP that powers Alpaca’s global, trading-critical systems. Own Infrastructure-as-Code with Terraform and follow GitOps, building CI/CD pipelines for safe IaC changes. Improve observability across Prometheus, Thanos, Grafana, Loki, Tempo, and Alertmanager, operate GKE clusters with Helm, and run operator-level database and message broker operations. Participate in follow-the-sun on-call and embed SRE practices.
Location: United States
Workplace: Remote
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Design and evolve Alpaca’s cloud architecture on GCP (networking, interconnects, IAM, and high-availability topology) expressed as code with Terraform using GitOps.
  • •Build and own CI/CD pipelines for IaC changes, including policy guardrails, drift detection, and progressive rollout for safe infrastructure shipping.
  • •Advance Platform-as-a-Product by building self-serve capabilities so engineers can provision via a “golden path.”
  • •Strengthen observability across metrics, logs, traces, and alerting using Prometheus, Thanos, Grafana, Loki, Tempo, and Alertmanager.
  • •Operate GKE clusters and infrastructure services running on them, including Helm-packaged workloads, message brokers, and data stores, and participate in follow-the-sun on-call.

Pay and Benefits

Equity and Bonus:Equity
Perks:EquityHealth InsuranceHome OfficeMonthly Stipend

Key Requirements

  • •5+ years in DevOps, Platform/Infrastructure, or SRE operating large-scale, high-availability systems in production.
  • •Deep hands-on cloud architecture on Google Cloud Platform (GCP): landing zones, networking, IAM, and high-availability topology.
  • •Strong Infrastructure-as-Code with Terraform, structuring large codebases across environments with GitOps and least-privilege.
  • •Proven CI/CD pipelines for IaC (automated plan/apply, code review, Policy-as-Code, drift detection, and safe rollout).
  • •Production experience with Kubernetes (ideally GKE) and deploying workloads with Helm.
Experience:5+ years
Skills:Incident managementStructured debuggingWritten communicationBlameless post-mortemsAccountability
Languages:English
Tech Stack:GCPGoogle Cloud PlatformTerraformGitOpsPolicy-as-CodeCI/CDKubernetesGKEHelmPrometheusThanosGrafanaLokiTempoAlertmanagerRabbitMQIBM MQPostgreSQLMessage BrokersRedPanda

Company Brief

Alpaca
Provides developer-first APIs and brokerage infrastructure for algorithmic and automated stock trading, enabling fintechs and developers to build trading applications with commission-free market access and custody services.
Industry: Fintech Infrastructure
Company Size: Medium (51 to 250 employees)
Growth: Scaleup
Funding: Series C
Headquarters: San Mateo, United States
Founded: 2015
WebsiteLinkedIn