Senior Staff Data Center Operations Engineer, GPU Hardware Architecture

Crusoe
San Francisco
Workplace: OnsiteFull timeFunction: Solutions Engineering & Sales EngineeringExperience: 10+ yearsSkills: ["Communication","Leadership","Problem-solving","Stakeholder management"]

Senior Staff Engineer for Data Center Operations acting as the technical architect and strategic partner to the Director of Data Center Operations. You’ll bridge high-level hardware engineering with ground-level execution across AI fleet hardware (H200, Blackwell GB200, GB300, Rubin), leading operational governance, tooling, power topology, resilience, and documentation to improve reliability and scale of Crusoe’s AI infrastructure.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Crusoe
Crusoe
4 months ago

Senior Staff Data Center Operations Engineer, GPU Hardware Architecture

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 6 hours agoStatus: Live

Job Summary

Senior Staff Engineer for Data Center Operations acting as the technical architect and strategic partner to the Director of Data Center Operations. You’ll bridge high-level hardware engineering with ground-level execution across AI fleet hardware (H200, Blackwell GB200, GB300, Rubin), leading operational governance, tooling, power topology, resilience, and documentation to improve reliability and scale of Crusoe’s AI infrastructure.
Location: San Francisco
Workplace: Onsite
Employment Type: Full time
Job Function: Solutions Engineering & Sales Engineering
Seniority: Sr. Manager level

Key Responsibilities

  • •Oversee operational governance and metrics for global ticket queue, and develop real-time dashboards tracking KPIs/SLAs (MTTR, fleet availability, sparing accuracy)
  • •Partner with Fleet Engineering to define software access, diagnostic hooks, and tooling for maximum repair efficiency and serviceability in the white space
  • •Lead power topology strategy, mapping end-to-end Power Strings and performing Build vs. Buy analyses for internal tools vs third-party solutions
  • •Architect the framework for Business Continuity Planning and Disaster Recovery, defining hardware recovery protocols and site failovers
  • •Provide technical advisory and sign-off to Documentation Committee, ensuring accurate break-fix SOPs and playbooks for global scale
  • •Mentor senior technicians and site leads, serving as the final technical authority for systemic hardware failures

Pay and Benefits

Equity and Bonus:Equity
Perks:Health InsuranceDentalVisionRsus401kLife InsuranceParental LeaveDisabilityCommuter Benefits

Key Requirements

  • •10+ years in Data Center Operations, Systems Engineering, or HPC hardware
  • •expert-level understanding of x86/GPU server architecture and electrical distribution
  • •experience in hardware maintenance at scale and translating field challenges into technical requirements
  • •ability to define operational KPIs and build dashboards (e.g., Tableau, Grafana)
  • •strong communication and leadership to distill risks and infrastructure hurdles for senior leadership
Experience:10+ yearsAI infrastructureData center operationsHPC hardwareGPU servers
Skills:CommunicationLeadershipProblem-solvingStakeholder management
Tech Stack:TableauGrafanaX86GPUHPCServer architectureElectrical distribution

Company Brief

Crusoe
Builds vertically integrated, energy-first AI infrastructure and purpose-built AI data centers (Crusoe Cloud), leveraging clean/stranded energy to power large-scale GPU compute for AI training and inference.
Industry: Data Centers
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Decacorn (USD 10B+)
Funding: Series E+
Headquarters: Denver, United States
Founded: 2018
Glassdoor
Glassdoor: 3.7
WebsiteLinkedInGlassdoor