Senior Manager, Engineering - AI Inference

Crusoe
San Francisco
Workplace: OnsiteFull timeUSD 250,000 - 300,000 annuallyFunction: Data Science & Machine LearningExperience: 2+ yearsEducation: bachelorsSkills: ["Communication","Leadership","Ownership","Problem-solving","Customer communication"]

Lead an engineering team focused on production inference for large language models—making them faster, cheaper, and more reliable. You’ll stay hands-on with the inference stack end to end: profiling performance, applying modern optimization techniques, and debugging issues down to CUDA kernels. Partner with customer engineering teams to tailor deployments and move workloads from proof of concept to monitored production services, delivering measurable latency, throughput, and cost gains.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Crusoe
Crusoe
1 day ago

Senior Manager, Engineering - AI Inference

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 13 hours agoStatus: Live

Job Summary

Lead an engineering team focused on production inference for large language models—making them faster, cheaper, and more reliable. You’ll stay hands-on with the inference stack end to end: profiling performance, applying modern optimization techniques, and debugging issues down to CUDA kernels. Partner with customer engineering teams to tailor deployments and move workloads from proof of concept to monitored production services, delivering measurable latency, throughput, and cost gains.
Location: San Francisco
Workplace: Onsite
Employment Type: Full time
Job Function: Data Science & Machine Learning
Seniority: Manager level

Key Responsibilities

  • •Lead an engineering team building production inference for large language models, balancing people management with staying deeply technical.
  • •Bring and refine inference techniques in production, including designing and optimizing serving architectures and request routing.
  • •Work down the serving stack—from frameworks (vLLM, SGLang) to CUDA kernels—profiling and fixing performance issues.
  • •Profile and tune deployments against latency, throughput, and cost targets while ensuring reliability under real traffic.
  • •Tailor deployments for each customer’s models and constraints, taking workloads from proof of concept to live, monitored production services.

Pay and Benefits

Salary: USD 250,000 - 300,000 annually
Equity and Bonus:Equity
Perks:Health InsuranceDentalVisionRsusPaid LeavePaid HolidaysParental Leave401kHsaLife InsuranceDisabilityVolunteer Time

Key Requirements

  • •2+ years of experience managing and leading an engineering team in a high-performance or ML-focused environment.
  • •Strong hands-on software engineering and low-level optimization or ML infrastructure experience, staying close to code and architecture.
  • •BS/MS/PhD in Computer Science, Engineering, Mathematics, or a related field.
  • •Hands-on production shipping experience with general-purpose languages, with a strong preference for Python.
  • •Experience optimizing LLM inference for throughput/latency, including comfort with vLLM or SGLang and profiling down to the kernel level.
Experience:2+ yearsAI inferenceLarge language modelsML infrastructureCustomer-facing ML
Education:Bachelor's in Computer Science, Engineering, Mathematics, or a related field
Skills:CommunicationLeadershipOwnershipProblem-solvingCustomer communication
Tech Stack:PythonC++VLLMSGLangCUDACUDA kernelsGPUDockerKubernetesProfilingML pipelines

Company Brief

Crusoe
Builds vertically integrated, energy-first AI infrastructure and purpose-built AI data centers (Crusoe Cloud), leveraging clean/stranded energy to power large-scale GPU compute for AI training and inference.
Industry: Data Centers
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Decacorn (USD 10B+)
Funding: Series E+
Headquarters: Denver, United States
Founded: 2018
Glassdoor
Glassdoor: 3.7
WebsiteLinkedInGlassdoor