Senior Hardware Systems Engineer, Performance

Crusoe
Sunnyvale
Workplace: OnsiteFull timeUSD 170,000 - 205,000 annuallyFunction: IT Operations (Systems/Network Admin)Experience: 5+ yearsSkills: ["Analytical problem-solving","Technical communication","Collaboration","Debugging","Automation mindset"]

Own the end-to-end lifecycle of next-generation CPU/GPU compute platforms, from prototype bring-up through production readiness. Drive performance characterization and validation for accelerated computing, run deep workload studies across training and inference models, and translate findings into cluster-level tuning and configuration. Lead complex system debugging across compute, memory, storage, networking, accelerators, and firmware while partnering with hardware, software, infrastructure, and vendor teams to improve reliability and efficiency.

This position is no longer accepting applications.

  • See live roles at Crusoe
  • Search all live jobs
  • Browse companies, collections, and locations hiring now
Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa

This position is no longer accepting applications.

See live roles at CrusoeSearch all live jobsBrowse companies, collections, and locations hiring now

Crusoe
Crusoe
6 months ago

Senior Hardware Systems Engineer, Performance

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 10 hours agoStatus: Closed

Job Summary

Own the end-to-end lifecycle of next-generation CPU/GPU compute platforms, from prototype bring-up through production readiness. Drive performance characterization and validation for accelerated computing, run deep workload studies across training and inference models, and translate findings into cluster-level tuning and configuration. Lead complex system debugging across compute, memory, storage, networking, accelerators, and firmware while partnering with hardware, software, infrastructure, and vendor teams to improve reliability and efficiency.
Location: Sunnyvale
Workplace: Onsite
Employment Type: Full time
Job Function: IT Operations (Systems/Network Admin)
Seniority: Mid level

Key Responsibilities

  • •Drive end-to-end lifecycle of next-generation compute platforms, including evaluation, bring-up, validation, deployment, and production readiness.
  • •Define and execute performance characterization and validation strategies for CPU, GPU, and accelerated computing platforms.
  • •Conduct workload characterization studies across training and inference models to understand compute, memory, communication, and I/O behavior.
  • •Translate workload/platform insights into cluster-level tuning and configuration recommendations to maximize performance and efficiency.
  • •Lead complex system-level debugging across compute, memory, storage, networking, accelerators, and platform firmware, collaborating across engineering disciplines to resolve issues.

Pay and Benefits

Salary: USD 170,000 - 205,000 annually
Equity and Bonus:Equity
Perks:Health InsuranceDentalVision401kHsaPaid ParentalPaid LifeLong-term DisabilityPaid LeaveCell PhoneCommuter Benefits

Key Requirements

  • •5-6+ years of experience in hardware systems engineering, platform engineering, performance engineering, ML systems engineering, infrastructure engineering, or related areas.
  • •Hands-on experience with large-scale GPU or accelerated computing infrastructure for AI/ML or HPC workloads.
  • •Hands-on experience with distributed training and/or inference workloads at scale, including parallelism strategies and performance tuning across the hardware/software stack.
  • •Experience with workload benchmarking, performance profiling, and system performance optimization across hardware and software layers.
  • •Strong understanding of modern server and accelerator architectures (CPU, GPU, memory, storage, networking) and high-speed interconnects like PCIe, InfiniBand, or NVLink.
Experience:5+ yearsAI/MLHPCDistributed trainingInference workloads
Education:
Skills:Analytical problem-solvingTechnical communicationCollaborationDebuggingAutomation mindset
Tech Stack:PythonShellPCIeInfiniBandNVLinkRDMARoCECXLTelemetryTelemetry, benchmarks, profiling toolsObservabilityDiagnosticsLinuxX86ARMMoELong-contextMultimodal models

Company Brief

Crusoe
Builds vertically integrated, energy-first AI infrastructure and purpose-built AI data centers (Crusoe Cloud), leveraging clean/stranded energy to power large-scale GPU compute for AI training and inference.
Industry: Data Centers
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Decacorn (USD 10B+)
Funding: Series E+
Headquarters: Denver, United States
Founded: 2018
Glassdoor
Glassdoor: 3.7
WebsiteLinkedInGlassdoor