Site Reliability Engineer - Low-Latency Trading Systems
United States
Workplace: RemoteFull timeFunction: DevOps, Cloud & InfrastructureSkills: ["Incident response","Debugging","Observability","Judgment under pressure","Clear communication"]Own production reliability for real-time trading services, including trading engines, execution gateways, market data ingestion, and PnL/reconciliation pipelines. Operate and evolve multi-region Kubernetes clusters on AWS (EKS) with GitOps (Flux) and encrypted secrets. Build observability with Prometheus and Datadog, create SLOs and alerts, and improve deploy safety through progressive rollouts and guardrails. Debug incidents end to end and participate in on-call across US equity and 24/7 crypto venues.
Loading
Loading job details...
Preparing the role view and application actions.

