Site Reliability Engineer - Telemetry

Kraken
Brazil, Uruguay, Peru, Colombia, Argentina, Paraguay, Canada, Chile
Full timeFunction: DevOps, Cloud & InfrastructureExperience: 3+ yearsSkills: ["Incident response","Documentation","Collaboration","Troubleshooting"]

Build and operate Kraken’s shared telemetry platform, ensuring metrics, logs, traces, alerting, dashboards, and profiling systems are reliable and scalable. You’ll maintain collection, storage, querying, and alerting with Prometheus-compatible tooling, run log and tracing pipelines, and use infrastructure-as-code to deploy telemetry services. You’ll troubleshoot production issues, create runbooks, support incident response and on-call, and automate observability workflows.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Kraken
Kraken
9 hours ago

Site Reliability Engineer - Telemetry

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 7 hours agoStatus: Live

Job Summary

Build and operate Kraken’s shared telemetry platform, ensuring metrics, logs, traces, alerting, dashboards, and profiling systems are reliable and scalable. You’ll maintain collection, storage, querying, and alerting with Prometheus-compatible tooling, run log and tracing pipelines, and use infrastructure-as-code to deploy telemetry services. You’ll troubleshoot production issues, create runbooks, support incident response and on-call, and automate observability workflows.
Location: Brazil, Uruguay, Peru, Colombia, Argentina, Paraguay, Canada, Chile
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Operate and improve the shared telemetry platform for metrics, logs, traces, alerting, dashboards, and profiling.
  • •Maintain metrics collection, long-term storage, querying, dashboards, and alerting using Prometheus-compatible systems and modern alerting tools.
  • •Operate log pipelines and distributed tracing/profiling using Vector, Splunk, Loki, Grafana Alloy, Tempo, OpenTelemetry, and Pyroscope.
  • •Deploy and manage telemetry services using Terraform, Terragrunt, and container orchestration across multiple environments.
  • •Troubleshoot missing data, slow queries, broken alerts, backpressure, and capacity issues; participate in incident response and on-call, write runbooks, and improve the platform.

Key Requirements

  • •3+ years of experience as a Site Reliability Engineer, Platform/Infrastructure Engineer, Observability Engineer, or similar production engineering role.
  • •Experience managing production telemetry systems (metrics, logs, traces, or profiles) at scale.
  • •Experience with Prometheus or Prometheus-compatible monitoring stacks, including collection, querying, and alerting.
  • •Experience troubleshooting distributed production systems across availability, latency, data flow, and capacity issues.
  • •Experience with Infrastructure as Code, especially Terraform, along with CI/CD.
Experience:3+ yearsProduction engineeringDistributed systemsObservability
Skills:Incident responseDocumentationCollaborationTroubleshooting
Tech Stack:PrometheusVictoriaMetricsGrafanaPromQLVectorSplunkLokiGrafana AlloyTempoOpenTelemetryPyroscopeTerraformTerragruntNomadKubernetesCI/CDAlertmanagerLogQL

Company Brief

Kraken
Kraken (Payward, Inc.) is a global cryptocurrency exchange and financial infrastructure provider offering spot and derivatives trading, staking, custody, tokenized assets and institutional services to retail and institutional clients.
Industry: Trading Platforms
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Scaleup
Valuation: Decacorn (USD 10B+)
Funding: Series E+
Headquarters: San Francisco, United States
Founded: 2011
Glassdoor
Glassdoor: 4.2
WebsiteLinkedInGlassdoor