Staff Software Engineer - Adaptive Telemetry, Databases | USA | Remote

Grafana Labs
United States
Workplace: RemoteFull timeUSD 174,986 - 209,983Function: Software EngineeringSkills: ["Technical leadership","Mentorship","Communication","Stakeholder alignment","Bias to action"]

Lead backend engineering for Adaptive Telemetry systems in Grafana Cloud, defining technical strategy and driving large cross-functional initiatives from planning through long-term operations. Own architecture, reliability, performance, and cost for critical telemetry databases, set SLOs/SLIs, and run incident response with blameless post-mortems. Improve observability and automation to reduce toil and MTTR while mentoring engineers and aligning stakeholders across teams.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Grafana Labs
Grafana Labs
2 months ago

Staff Software Engineer - Adaptive Telemetry, Databases | USA | Remote

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 16 minutes agoStatus: Live

Job Summary

Lead backend engineering for Adaptive Telemetry systems in Grafana Cloud, defining technical strategy and driving large cross-functional initiatives from planning through long-term operations. Own architecture, reliability, performance, and cost for critical telemetry databases, set SLOs/SLIs, and run incident response with blameless post-mortems. Improve observability and automation to reduce toil and MTTR while mentoring engineers and aligning stakeholders across teams.
Location: United States
Workplace: Remote
Employment Type: Full time
Job Function: Software Engineering
Seniority: Sr. Manager level

Key Responsibilities

  • •Drive technical strategy and roadmap by defining architectural vision and influencing product/engineering decisions.
  • •Lead end-to-end delivery of large cross-functional initiatives, including planning, design, execution, rollout, and long-term operations.
  • •Own architecture, reliability, performance, and cost for critical systems, balancing scalability, availability, latency, and maintainability.
  • •Define SLOs/SLIs and lead incident response, including high-severity incident handling and blameless post-mortems with systemic fixes.
  • •Improve observability, automation, and operational readiness through telemetry, alerting, runbooks, capacity planning, and efforts that reduce toil and MTTR.

Pay and Benefits

Salary: USD 174,986 - 209,983
Equity and Bonus:Equity
Perks:Equity

Key Requirements

  • •Proven delivery and operation of large distributed systems across multiple teams with evidence of technical leadership and impact.
  • •Deep systems-design skills with strong tradeoff judgment around latency, consistency, availability, scaling, and cost.
  • •Hands-on experience with cloud-native architectures, including microservices, containers/Kubernetes, and IaC, plus operational practices to keep them healthy.
  • •Ability to define SLOs/SLIs, run capacity planning, tune performance, and drive reliability work end-to-end.
  • •Excellent coding and design skills; experience leading technical designs, with Go or transferable languages such as Python, C, C++, or Rust.
Experience:ObservabilityTelemetryDistributed systemsCloud-nativeOpen source
Skills:Technical leadershipMentorshipCommunicationStakeholder alignmentBias to action
Languages:English
Tech Stack:GoPythonCC++RustKafkaPrometheusGrafanaKubernetesMicroservicesIaCMimirLokiTempoPyroscope

Company Brief

Grafana Labs
Develops open-source and commercial observability tools, including Grafana for metrics visualization, Loki for log aggregation, and Tempo for tracing, enabling organizations to monitor, analyze, and visualize system performance and telemetry data.
Industry: Developer Tools
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series D
Headquarters: New York, United States
Founded: 2014
WebsiteLinkedIn