Staff Software Engineer, Cloud Monitoring Service (Distributed Systems)

Crusoe
San Francisco
Workplace: OnsiteFull timeUSD 215,000 - 260,000 annuallyFunction: Software EngineeringExperience: 8+ yearsSkills: ["Problem-solving","Technical leadership","Mentorship","Cross-functional communication","Operational excellence"]

Lead distributed systems design for the Cloud Monitoring Service, building the observability backbone that collects, processes, stores, and serves metrics and logs for AI workloads. Own scalable multi-tenant telemetry pipelines—from edge collection through ingestion, storage, and low-latency query—while improving reliability, data freshness, and cost efficiency. Provide technical direction, drive design reviews, and mentor engineers across code reviews and incident response in a cross-team environment.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Crusoe
Crusoe
2 months ago

Staff Software Engineer, Cloud Monitoring Service (Distributed Systems)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 15 hours agoStatus: Live

Job Summary

Lead distributed systems design for the Cloud Monitoring Service, building the observability backbone that collects, processes, stores, and serves metrics and logs for AI workloads. Own scalable multi-tenant telemetry pipelines—from edge collection through ingestion, storage, and low-latency query—while improving reliability, data freshness, and cost efficiency. Provide technical direction, drive design reviews, and mentor engineers across code reviews and incident response in a cross-team environment.
Location: San Francisco
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Own architecture and evolution of large-scale telemetry pipelines, including high-volume ingestion, stream processing, time-series and log storage, and low-latency query paths.
  • •Design highly scalable, durable, and fair multi-tenant services while addressing hot shards, high-cardinality data, backpressure, retention/compaction at scale, and graceful degradation.
  • •Build for operability from day one by improving pipeline reliability and data freshness and participating in a customer-facing on-call rotation.
  • •Set technical direction for the team’s distributed systems work, driving design reviews and raising the bar for scoping, building, and operating systems.
  • •Collaborate with product, compute, networking, and platform teams to ensure observability decisions are made with full context, and mentor engineers through design work and incident response.

Pay and Benefits

Salary: USD 215,000 - 260,000 annually
Equity and Bonus:Equity
Perks:EquityHealth InsuranceDentalVision401kPaid LeaveParental LeaveHsaVolunteer Time

Key Requirements

  • •8+ years of software development experience with sustained ownership of production distributed systems.
  • •Deep, hands-on experience designing and operating distributed systems at scale, including sharding, replication, consistency, load balancing, and concurrency.
  • •Experience building or operating large-scale observability/data infrastructure such as time-series databases, log aggregation, streaming pipelines, or distributed tracing backends.
  • •Strong programming fundamentals in Go (Go strongly preferred) or another modern compiled language, with comfort building systems in Kubernetes, microservices, and CI/CD.
  • •On-call experience on a customer-facing service and ability to treat incidents as design feedback to systematically eliminate root causes.
Experience:8+ yearsDistributed systemsObservabilityCloud infrastructureData infrastructureAI infrastructure
Skills:Problem-solvingTechnical leadershipMentorshipCross-functional communicationOperational excellence
Tech Stack:GoKubernetesMicroservicesCI/CDPrometheusVictoriaMetricsLokiOpenTelemetryKafkaVectorStream processingTime-series databasesLog aggregationDistributed tracing

Company Brief

Crusoe
Builds vertically integrated, energy-first AI infrastructure and purpose-built AI data centers (Crusoe Cloud), leveraging clean/stranded energy to power large-scale GPU compute for AI training and inference.
Industry: Data Centers
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Decacorn (USD 10B+)
Funding: Series E+
Headquarters: Denver, United States
Founded: 2018
Glassdoor
Glassdoor: 3.7
WebsiteLinkedInGlassdoor