Staff Production Engineer, Compute

Crusoe
Sunnyvale
Workplace: OnsiteFull timeUSD 209,000 - 253,000 annuallyFunction: Manufacturing & Production OperationsExperience: 8+ yearsSkills: ["Problem-solving","Root cause analysis","System-level debugging","Collaboration","Urgency"]

Support virtualization, hypervisor, and kernel-level performance for AI-first cloud compute infrastructure. Build automation and observability to monitor the stack from kernel to orchestration, deploy and optimize bare-metal and virtualized platforms, and tune Linux subsystems for CPU/GPU/DPU/NIC performance. Lead root-cause analysis for crashes and performance regressions, and collaborate with platform teams to validate emerging hardware support (SmartNICs, BlueField, TPUs).

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Crusoe
Crusoe
1 month ago

Staff Production Engineer, Compute

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 14 hours agoStatus: Live

Job Summary

Support virtualization, hypervisor, and kernel-level performance for AI-first cloud compute infrastructure. Build automation and observability to monitor the stack from kernel to orchestration, deploy and optimize bare-metal and virtualized platforms, and tune Linux subsystems for CPU/GPU/DPU/NIC performance. Lead root-cause analysis for crashes and performance regressions, and collaborate with platform teams to validate emerging hardware support (SmartNICs, BlueField, TPUs).
Location: Sunnyvale
Workplace: Onsite
Employment Type: Full time
Job Function: Manufacturing & Production Operations

Key Responsibilities

  • •Develop automation and observability tools to monitor compute infrastructure from kernel to orchestration layers.
  • •Support and scale the virtualization stack and deploy/optimize bare-metal and virtualized compute platforms for AI and HPC workloads.
  • •Collaborate with Linux kernel and hardware teams to identify/resolve performance bottlenecks, driver issues, and optimize hardware offloads.
  • •Perform root cause analysis for kernel crashes, hardware-software integration problems, and performance regressions.
  • •Tune kernel subsystems (process scheduler, NUMA, memory management, interrupt handling) and integrate hypervisor enhancements to improve guest VM reliability and isolation.

Pay and Benefits

Salary: USD 209,000 - 253,000 annually
Equity and Bonus:Equity
Perks:Health InsuranceDentalVisionHsaPaid Parental401kRsusPaid LeaveCommuter Benefits

Key Requirements

  • •8+ years in Compute SRE, Linux system engineering, or compute infrastructure roles.
  • •Strong Linux kernel internals proficiency (scheduler, memory allocation, driver subsystems).
  • •Experience with virtualization architectures/technologies such as KVM, Xen, QEMU, or VMware.
  • •Familiarity with SmartNICs/DPUs and kernel bypass techniques.
  • •Expert-level skills in at least one programming language: Go, C, or Rust, plus system-level debugging experience (kdump, kexec, kernel panic analysis).
Experience:8+ yearsAI infrastructureCompute SRELinuxVirtualizationHPCCloud infrastructure
Skills:Problem-solvingRoot cause analysisSystem-level debuggingCollaborationUrgency
Tech Stack:LinuxLinux kernelKVMQEMUXenVMwareSmartNICsDPUsNVIDIA CX6NVIDIA CX7BlueField-3TPUsGoCRustInfrastructure as CodeCI/CDKdumpKexecNUMA

Company Brief

Crusoe
Builds vertically integrated, energy-first AI infrastructure and purpose-built AI data centers (Crusoe Cloud), leveraging clean/stranded energy to power large-scale GPU compute for AI training and inference.
Industry: Data Centers
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Decacorn (USD 10B+)
Funding: Series E+
Headquarters: Denver, United States
Founded: 2018
Glassdoor
Glassdoor: 3.7
WebsiteLinkedInGlassdoor