Senior DevOps Engineer

Alpaca
Japan
Workplace: RemoteFull timeFunction: DevOps, Cloud & InfrastructureExperience: 5+ yearsSkills: ["Written communication","Incident management","Debugging under pressure","Autonomy","Accountability"]

Design, build, and operate the infrastructure that enables global scaling and trading-critical reliability. Own cloud architecture on GCP as code using Terraform and GitOps, create IaC-focused CI/CD with policy and safe rollouts, and develop self-serve platform capabilities. Strengthen observability with Prometheus/Thanos/Grafana/Loki/Tempo, operate GKE workloads with Helm, and lead incident management with SRE practices and follow-the-sun on-call from APAC hours.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Alpaca
Alpaca
1 month ago

Senior DevOps Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 18 hours agoStatus: Live

Job Summary

Design, build, and operate the infrastructure that enables global scaling and trading-critical reliability. Own cloud architecture on GCP as code using Terraform and GitOps, create IaC-focused CI/CD with policy and safe rollouts, and develop self-serve platform capabilities. Strengthen observability with Prometheus/Thanos/Grafana/Loki/Tempo, operate GKE workloads with Helm, and lead incident management with SRE practices and follow-the-sun on-call from APAC hours.
Location: Japan
Workplace: Remote
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Design and evolve GCP cloud architecture (networking, interconnects, IAM, high-availability topology) expressed entirely as code with Terraform using GitOps.
  • •Build and own CI/CD pipelines for IaC changes with policy guardrails, drift detection, and progressive rollout for safe infrastructure delivery.
  • •Advance the platform-as-a-product by building self-serve capabilities and golden paths so engineers can provision without hand-offs.
  • •Strengthen observability across metrics, logs, traces, and alerting using Prometheus, Thanos, Grafana, Loki, Tempo, and Alertmanager.
  • •Operate GKE clusters and infrastructure services (Helm workloads, message brokers, data stores) and lead incident management and blameless post-mortems under follow-the-sun on-call.

Pay and Benefits

Equity and Bonus:Equity
Perks:Health InsuranceEquityHome Office

Key Requirements

  • •5+ years in DevOps, Platform/Infrastructure, or SRE operating large-scale, high-availability systems in production.
  • •Deep hands-on experience designing cloud architecture on Google Cloud Platform (GCP), including landing zones, networking, IAM, and HA topology.
  • •Strong Infrastructure-as-Code skills with Terraform, using GitOps as a first principle and least-privilege as a default.
  • •Proven experience building CI/CD pipelines for IaC (automated plan/apply, code review, Policy-as-Code, drift detection, safe rollout).
  • •Significant production experience with Kubernetes (ideally GKE) and deploying workloads with Helm.
Experience:5+ yearsFintechTradingBrokerage infrastructureOpen sourceHigh-availability systems
Skills:Written communicationIncident managementDebugging under pressureAutonomyAccountability
Languages:English
Tech Stack:GCPGoogle Cloud PlatformTerraformGitOpsInfrastructure-as-CodePolicy-as-CodeIaCCI/CDKubernetesGKEHelmPrometheusThanosGrafanaLokiTempoAlertmanagerPostgreSQLRabbitMQIBM MQ

Company Brief

Alpaca
Provides developer-first APIs and brokerage infrastructure for algorithmic and automated stock trading, enabling fintechs and developers to build trading applications with commission-free market access and custody services.
Industry: Fintech Infrastructure
Company Size: Medium (51 to 250 employees)
Growth: Scaleup
Funding: Series C
Headquarters: San Mateo, United States
Founded: 2015
WebsiteLinkedIn