Principal Engineer, Conductor Platform (CAPE)

Crusoe
San Francisco
Workplace: OnsiteFull timeUSD 285,000 - 335,000Function: Communications, PR & CommunityExperience: 10+ yearsSkills: ["Problem-solving","Opportunity-finding","Sense of urgency","Comfort operating in ambiguity","Judgment"]

Build Crusoe’s self-driving Conductor control plane by designing the systems that predict, decide, and remediate failures in a fleet of tens of thousands of accelerators. You’ll own unified observability, fleet-wide scheduling and health models, closed-loop autonomy, and goodput optimization—linking GPU/HPC telemetry, network fabrics, storage, and multi-tenant security to real-time power and cost controls.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Crusoe
Crusoe
1 month ago

Principal Engineer, Conductor Platform (CAPE)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 11 hours agoStatus: Live

Job Summary

Build Crusoe’s self-driving Conductor control plane by designing the systems that predict, decide, and remediate failures in a fleet of tens of thousands of accelerators. You’ll own unified observability, fleet-wide scheduling and health models, closed-loop autonomy, and goodput optimization—linking GPU/HPC telemetry, network fabrics, storage, and multi-tenant security to real-time power and cost controls.
Location: San Francisco
Workplace: Onsite
Employment Type: Full time
Job Function: Communications, PR & Community
Seniority: Mid level

Key Responsibilities

  • •Build a unified observability plane correlating GPU, networking (InfiniBand/RoCE), storage, orchestration, and workload signals for fast diagnosis and recovery.
  • •Create fleet-wide autonomy so the system predicts, decides, and remediates with no human in the loop (drain, checkpoint, replace, resume).
  • •Optimize goodput by continuously trading scheduling, placement, and maintenance decisions against a goodput objective function.
  • •Forecast failures hours ahead using GPU and hardware telemetry and pre-emptively migrate work.
  • •Implement energy-aware compute and fleet digital-twin approaches to ensure safe, zero-trust, fully auditable multi-tenancy operations.

Pay and Benefits

Salary: USD 285,000 - 335,000
Equity and Bonus:Equity
Perks:Health InsuranceDentalVisionEquityParental Leave401kPaid LeaveLife InsuranceMeal AllowanceHsa

Key Requirements

  • •10+ years building infrastructure-layer systems at scale (fleet management, distributed control planes, scheduler internals, or hardware lifecycle automation).
  • •Deep experience with distributed systems design including consensus and state reconciliation, with closed-loop automation operating against live production infrastructure.
  • •Hands-on fluency with GPU/HPC infrastructure: GPU health telemetry, NVLink/InfiniBand/RoCE fabrics, plus thermal and power behavior.
  • •Experience designing and shipping large-scale observability or telemetry platforms that correlate signals across compute, network, and storage layers.
  • •Strong software engineering fundamentals in at least one systems language (Go, Rust, C++ or similar) and experience applying ML/statistical methods to noisy operational telemetry (plus).
Experience:10+ years
Skills:Problem-solvingOpportunity-findingSense of urgencyComfort operating in ambiguityJudgment
Tech Stack:GoRustC++InfiniBandRoCENVLinkHPCGPUMLStatistical methods

Company Brief

Crusoe
Builds vertically integrated, energy-first AI infrastructure and purpose-built AI data centers (Crusoe Cloud), leveraging clean/stranded energy to power large-scale GPU compute for AI training and inference.
Industry: Data Centers
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Decacorn (USD 10B+)
Funding: Series E+
Headquarters: Denver, United States
Founded: 2018
Glassdoor
Glassdoor: 3.7
WebsiteLinkedInGlassdoor