Site Reliability Engineer - Telemetry
Brazil, Uruguay, Peru, Colombia, Argentina, Paraguay, Canada, Chile
Full timeFunction: DevOps, Cloud & InfrastructureExperience: 3+ yearsSkills: ["Incident response","Documentation","Collaboration","Troubleshooting"]Build and operate Kraken’s shared telemetry platform, ensuring metrics, logs, traces, alerting, dashboards, and profiling systems are reliable and scalable. You’ll maintain collection, storage, querying, and alerting with Prometheus-compatible tooling, run log and tracing pipelines, and use infrastructure-as-code to deploy telemetry services. You’ll troubleshoot production issues, create runbooks, support incident response and on-call, and automate observability workflows.
Loading
Loading job details...
Preparing the role view and application actions.

