Senior Staff Software Engineer, DC Infrastructure

Crusoe
San Francisco, Sunnyvale
Workplace: OnsiteFull timeUSD 250,000 - 300,000 annuallyFunction: Software EngineeringSkills: ["Problem-solving","Analytical thinking","Communication","Collaboration","Independent execution"]

Build and maintain software that manages Crusoe’s fleet of GPU servers and the data centers that house them. Own development of deep diagnostics, troubleshooting, observability, and automation tooling to execute component-level diagnosis and remediation for failed or degraded hardware. Partner with data center operations on critical environment tooling and facilities management for power and direct liquid cooling systems. Deploy, monitor, and support these tools to maximize GPU fleet availability, stability, and performance.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Crusoe
Crusoe
1 day ago

Senior Staff Software Engineer, DC Infrastructure

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 6 hours agoStatus: Live

Job Summary

Build and maintain software that manages Crusoe’s fleet of GPU servers and the data centers that house them. Own development of deep diagnostics, troubleshooting, observability, and automation tooling to execute component-level diagnosis and remediation for failed or degraded hardware. Partner with data center operations on critical environment tooling and facilities management for power and direct liquid cooling systems. Deploy, monitor, and support these tools to maximize GPU fleet availability, stability, and performance.
Location: San Francisco, Sunnyvale
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering
Seniority: Sr. Manager level

Key Responsibilities

  • •Develop deep-level diagnostics and troubleshooting for hardware faults within GPU racks and high-density compute systems.
  • •Build troubleshooting, automation, and AI/agent-driven remediation tooling for GPU platforms and degraded hardware.
  • •Create automation and observability/diagnostic tooling for managing critical environment alongside data center operations.
  • •Develop post-repair validation and testing tooling (e.g., burn-in, Pytorch, NVIDIA NCCL) to ensure stability and performance.
  • •Own deployment, monitoring, and operational support of tooling to maximize GPU fleet availability and performance.

Pay and Benefits

Salary: USD 250,000 - 300,000 annually
Equity and Bonus:Equity
Perks:Health InsuranceRsus401kHsaPaid ParentalPaid LeaveLife Insurance

Key Requirements

  • •Software engineering experience and ability to identify problems, develop scalable solutions, and ship them.
  • •Expertise in distributed systems, reliability, and cloud platforms such as Kubernetes and IaC.
  • •Strength in at least one programming language: Go, Python, Java, or Rust.
  • •Strong analytical and problem-solving skills, with excellent communication and collaboration.
  • •Ability to work independently and set the technical direction for a project.
Experience:AI infrastructureGPU fleet operationsHigh-performance computingData center operationsCloud services
Skills:Problem-solvingAnalytical thinkingCommunicationCollaborationIndependent execution
Tech Stack:KubernetesIaCGCPGoPythonJavaRustTemporalPytorchNVIDIA NCCLAI agentsObservability tooling

Company Brief

Crusoe
Builds vertically integrated, energy-first AI infrastructure and purpose-built AI data centers (Crusoe Cloud), leveraging clean/stranded energy to power large-scale GPU compute for AI training and inference.
Industry: Data Centers
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Decacorn (USD 10B+)
Funding: Series E+
Headquarters: Denver, United States
Founded: 2018
Glassdoor
Glassdoor: 3.7
WebsiteLinkedInGlassdoor