Staff Cloud Support Engineer

Crusoe
San Francisco, Sunnyvale
Workplace: OnsiteFull timeFunction: Solutions Engineering & Sales EngineeringExperience: 8+ yearsSkills: ["Communication","Leadership","Mentoring","Problem-solving"]

SeniorCloud Support Engineer leading reliability and incident response for Crusoe Cloud. You’ll design guardrails, influence architecture, mentor engineers, and protect revenue by preventing major incidents, leveraging Linux, Kubernetes, networking, and AI/ML infrastructure expertise to scale high-performance AI workloads globally.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Crusoe
Crusoe
6 months ago

Staff Cloud Support Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 14 hours agoStatus: Live

Job Summary

SeniorCloud Support Engineer leading reliability and incident response for Crusoe Cloud. You’ll design guardrails, influence architecture, mentor engineers, and protect revenue by preventing major incidents, leveraging Linux, Kubernetes, networking, and AI/ML infrastructure expertise to scale high-performance AI workloads globally.
Location: San Francisco, Sunnyvale
Workplace: Onsite
Employment Type: Full time
Job Function: Solutions Engineering & Sales Engineering

Key Responsibilities

  • •Serve as highest-level escalation point for complex P1/P0 incidents.
  • •Lead cross-functional root cause investigations involving compute, networking (IB/RDMA/RoCE), storage, and orchestration layers.
  • •Partner with SRE, Software teams (Storage, Networking, Compute, K8) to design systemic fixes rather than recurring workarounds.
  • •Design and improve node validation, burn-in processes, performance baselining, and release readiness.
  • •Influence Kubernetes architecture, workload orchestration (Slurm, Terraform), and AI/ML cluster stability.

Pay and Benefits

Equity and Bonus:Equity
Perks:Health InsuranceVisionDental401kRsus

Key Requirements

  • •8+ years experience in SRE, DevOps, HPC, or Cloud Infrastructure roles.
  • •Advanced Linux systems expertise.
  • •Deep Kubernetes operational experience (CKA-level or higher).
  • •Strong networking knowledge: Infiniband, RDMA, RoCE, SDN.
  • •Experience supporting AI/ML workloads at scale (GPU clusters).
Experience:8+ yearsCloud computingAI/MLSREHPCCloud infrastructure
Skills:CommunicationLeadershipMentoringProblem-solving
Languages:English
Tech Stack:LinuxKubernetesIBRDMARoCESlurmTerraformNCCLGPU drivers

Company Brief

Crusoe
Builds vertically integrated, energy-first AI infrastructure and purpose-built AI data centers (Crusoe Cloud), leveraging clean/stranded energy to power large-scale GPU compute for AI training and inference.
Industry: Data Centers
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Decacorn (USD 10B+)
Funding: Series E+
Headquarters: Denver, United States
Founded: 2018
Glassdoor
Glassdoor: 3.7
WebsiteLinkedInGlassdoor