Principal Production Engineer

Crusoe
San Francisco, Sunnyvale
Workplace: OnsiteFull timeFunction: Manufacturing & Production OperationsExperience: 15+ yearsSkills: ["Problem-solving","Leadership","Communication","Analytical","Mentorship"]

Lead the reliability, scalability, and operational excellence of Crusoe’s cloud infrastructure. Own on-call readiness, incident response, and observability tooling across compute, storage, and networking; partner with software, hardware, and network teams to influence architecture; mentor engineers and set production engineering standards to grow with the company.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Crusoe
Crusoe
4 months ago

Principal Production Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 6 hours agoStatus: Live

Job Summary

Lead the reliability, scalability, and operational excellence of Crusoe’s cloud infrastructure. Own on-call readiness, incident response, and observability tooling across compute, storage, and networking; partner with software, hardware, and network teams to influence architecture; mentor engineers and set production engineering standards to grow with the company.
Location: San Francisco, Sunnyvale
Workplace: Onsite
Employment Type: Full time
Job Function: Manufacturing & Production Operations

Key Responsibilities

  • •Own the reliability and scalability of Crusoe's cloud infrastructure — compute, storage, and networking — defining SLOs, leading incident response, and driving systemic improvements across the platform
  • •Build and mature observability and tooling from network telemetry to control plane instrumentation and on-call tooling to diagnose issues faster than customers notice them
  • •Drive platform reliability improvements across the full cloud stack by partnering with software, hardware, and network teams to influence early architecture decisions
  • •Act as a trusted advisor to senior leadership, advocating for long-term technology investments in observability and reliability
  • •Set technical standards for the production engineering organization, defining on-call culture, incident frameworks, and reliability practices as the company grows

Pay and Benefits

Perks:Health InsuranceVisionDental401kRsusParental LeaveLife InsuranceDisabilityTeladocCommuter BenefitsLearning Budget

Key Requirements

  • •15+ years of experience in infrastructure, networking, or production engineering with internet-scale exposure (cloud providers, CDNs, or similar)
  • •Strong systems fundamentals: Linux, distributed systems, storage, compute scheduling
  • •Hands-on data center experience with physical infra, power and thermal constraints
  • •Ability to write code to automate, instrument, and build tooling for the team
  • •Excellent analytical and problem-solving skills for synthesizing signals into clear statements and designs
  • •Strong incident command: calm leadership during outages and blameless retrospectives
Experience:15+ yearsCloudInfrastructureData centers
Skills:Problem-solvingLeadershipCommunicationAnalyticalMentorship
Tech Stack:LinuxDistributed systemsStorageCompute schedulingTelemetryControl planeOn-call toolingBGPOSPFECMPKubernetesSlurmInfiniBandRoCEPCIeCloudHardwareNetworking

Company Brief

Crusoe
Builds vertically integrated, energy-first AI infrastructure and purpose-built AI data centers (Crusoe Cloud), leveraging clean/stranded energy to power large-scale GPU compute for AI training and inference.
Industry: Data Centers
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Decacorn (USD 10B+)
Funding: Series E+
Headquarters: Denver, United States
Founded: 2018
Glassdoor
Glassdoor: 3.7
WebsiteLinkedInGlassdoor