Software Engineer I (DCIE)

Crusoe
San Francisco
Workplace: OnsiteFull timeUSD 140,000 - 165,000 annuallyFunction: Software EngineeringExperience: 2-3 yearsSkills: ["Problem-solving","Analytical thinking","Communication","Collaboration"]

Build and deploy software that keeps Crusoe’s GPU fleet and data center infrastructure running reliably. You’ll develop diagnostics, observability, automation, and repair tooling to troubleshoot GPU rack and compute-system hardware faults, including post-repair validation and performance testing. Partner with data center operations to create AI-driven agents and operational tooling for critical environments, power, and direct liquid cooling systems. Own monitoring and operational support to maximize fleet availability and performance.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Crusoe
Crusoe
3 days ago

Software Engineer I (DCIE)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 3 hours agoStatus: Live

Job Summary

Build and deploy software that keeps Crusoe’s GPU fleet and data center infrastructure running reliably. You’ll develop diagnostics, observability, automation, and repair tooling to troubleshoot GPU rack and compute-system hardware faults, including post-repair validation and performance testing. Partner with data center operations to create AI-driven agents and operational tooling for critical environments, power, and direct liquid cooling systems. Own monitoring and operational support to maximize fleet availability and performance.
Location: San Francisco
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Develop and implement deep-level diagnostics and troubleshooting for hardware faults in GPU racks and high-density compute systems.
  • •Build troubleshooting and automation tooling for GPU platforms including NVIDIA A100/H200/GB200/B200 and AMD 350X/355X.
  • •Create automation and AI agents for component-level diagnosis and remediation of failed or degraded hardware.
  • •Develop tooling for post-repair validation and testing (e.g., burn-in, PyTorch, NVIDIA NCCL) to ensure stability and performance.
  • •Own deployment, monitoring, and operational support of tooling, maximizing GPU fleet availability and performance in partnership with data center operations.

Pay and Benefits

Salary: USD 140,000 - 165,000 annually
Equity and Bonus:Equity
Perks:Health InsuranceVisionDental401kHsa AccountsPaid ParentalLife InsuranceDisabilityPaid LeaveCommuter Benefits

Key Requirements

  • •2-3 years of software engineering experience.
  • •Ability to identify problems, rapidly develop scalable solutions, and ship them.
  • •Comfort working independently while also supporting team members on critical initiatives.
  • •Expertise in distributed systems, reliability, and cloud platforms (e.g., Kubernetes, IaC, GCP).
  • •Strength in at least one programming language: Go, Python, Java, or Rust.
Experience:2-3 years
Skills:Problem-solvingAnalytical thinkingCommunicationCollaboration
Tech Stack:KubernetesIaCGCPGoPythonJavaRustTemporalPytorchNVIDIA NCCL

Company Brief

Crusoe
Builds vertically integrated, energy-first AI infrastructure and purpose-built AI data centers (Crusoe Cloud), leveraging clean/stranded energy to power large-scale GPU compute for AI training and inference.
Industry: Data Centers
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Decacorn (USD 10B+)
Funding: Series E+
Headquarters: Denver, United States
Founded: 2018
Glassdoor
Glassdoor: 3.7
WebsiteLinkedInGlassdoor