Senior Production Engineer, Compute

Crusoe
San Francisco, Sunnyvale
Workplace: OnsiteFull timeUSD 170,000 - 205,000Function: Manufacturing & Production OperationsExperience: 5+ yearsSkills: ["Problem-solving","Urgency"]

Build and optimize Crusoe’s compute infrastructure for AI and HPC workloads by supporting virtualization, hypervisors, and kernel-level performance. Develop automation and observability to monitor systems from the kernel to orchestration layers, tune kernel subsystems (scheduler, NUMA, memory, interrupts), and perform root-cause analysis for crashes and performance regressions. Collaborate with Linux kernel and hardware teams to improve reliability, isolation, and hardware offloads using CPU/GPU/DPU-NIC resources.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Crusoe
Crusoe
1 month ago

Senior Production Engineer, Compute

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 16 hours agoStatus: Live

Job Summary

Build and optimize Crusoe’s compute infrastructure for AI and HPC workloads by supporting virtualization, hypervisors, and kernel-level performance. Develop automation and observability to monitor systems from the kernel to orchestration layers, tune kernel subsystems (scheduler, NUMA, memory, interrupts), and perform root-cause analysis for crashes and performance regressions. Collaborate with Linux kernel and hardware teams to improve reliability, isolation, and hardware offloads using CPU/GPU/DPU-NIC resources.
Location: San Francisco, Sunnyvale
Workplace: Onsite
Employment Type: Full time
Job Function: Manufacturing & Production Operations
Seniority: Mid level

Key Responsibilities

  • •Develop automation and observability tools to monitor compute infrastructure from the kernel to orchestration layers.
  • •Support and scale bare-metal and virtualized compute platforms, including virtualization stacks and hypervisors.
  • •Tune kernel subsystems (process scheduler, NUMA configuration, memory management, and interrupt handling) to improve performance.
  • •Perform root-cause analysis for kernel crashes, hardware-software integration problems, and performance regressions.
  • •Integrate hypervisor-level enhancements and validate support for emerging compute hardware (SmartNICs, BlueField devices, and TPUs).

Pay and Benefits

Salary: USD 170,000 - 205,000
Perks:Health InsuranceVisionDentalHsaPaid ParentalLife InsuranceDisability401kPaid LeaveCell PhoneCommuter BenefitsEquity

Key Requirements

  • •5+ years of professional experience in Compute SRE, Linux system engineering, or compute infrastructure roles.
  • •Strong proficiency in Linux kernel internals, including scheduler, memory allocation, and driver subsystems.
  • •Experience with virtualization technologies such as KVM, Xen, QEMU, or VMware.
  • •Familiarity with SmartNICs/DPUs (e.g., NVIDIA CX6/7, BlueField-3) and kernel bypass techniques.
  • •Expert-level skills in at least one programming language: Go, C, or Rust.
Experience:5+ years
Skills:Problem-solvingUrgency
Tech Stack:LinuxLinux kernelKVMQEMUXenVMwareHypervisorsNUMAMemory managementProcess schedulerInterrupt handlingSmartNICsDPUsNVIDIA CX6NVIDIA CX7BlueField-3Kernel bypassTPUsGoC

Company Brief

Crusoe
Builds vertically integrated, energy-first AI infrastructure and purpose-built AI data centers (Crusoe Cloud), leveraging clean/stranded energy to power large-scale GPU compute for AI training and inference.
Industry: Data Centers
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Decacorn (USD 10B+)
Funding: Series E+
Headquarters: Denver, United States
Founded: 2018
Glassdoor
Glassdoor: 3.7
WebsiteLinkedInGlassdoor