Applied AI Inference Engineer

Crusoe
San Francisco, Sunnyvale
Workplace: OnsiteFull timeUSD 250,000 - 300,000 annuallyFunction: Data Science & Machine LearningEducation: bachelorsSkills: ["Communication","Ownership","Problem-solving","Urgency","Tradeoff decision-making"]

Build and optimize the inference stack end to end so large language models run faster, cheaper, and more reliably in production. Profile latency and cost drivers, improve serving architectures and routing, and dive into low-level performance work across frameworks and CUDA kernels. Partner with customer engineering teams to tailor deployments, move workloads from proof of concept to monitored production, and turn performance gains into measurable outcomes.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Crusoe
Crusoe
1 month ago

Applied AI Inference Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 14 hours agoStatus: Live

Job Summary

Build and optimize the inference stack end to end so large language models run faster, cheaper, and more reliably in production. Profile latency and cost drivers, improve serving architectures and routing, and dive into low-level performance work across frameworks and CUDA kernels. Partner with customer engineering teams to tailor deployments, move workloads from proof of concept to monitored production, and turn performance gains into measurable outcomes.
Location: San Francisco, Sunnyvale
Workplace: Onsite
Employment Type: Full time
Job Function: Data Science & Machine Learning

Key Responsibilities

  • •Own the inference stack end to end, profiling where time and cost go and improving real deployments.
  • •Design and optimize serving architectures, including prefill/decode disaggregation and request routing.
  • •Work down into the serving stack from vLLM/SGLang to CUDA kernels to find and fix performance problems.
  • •Adapt and scale optimization methods across ML model types with an emphasis on large language models.
  • •Collaborate with customer engineering teams to tailor deployments and take workloads from proof of concept to monitored production services.

Pay and Benefits

Salary: USD 250,000 - 300,000 annually
Equity and Bonus:Equity
Perks:Health InsuranceDentalVisionEquity401k

Key Requirements

  • •Bachelors', Masters', or Ph.D. in Computer Science, Engineering, Mathematics, or a related field.
  • •Hands-on experience shipping production code using general-purpose languages such as Python or C++, with a strong preference for Python.
  • •Experience optimizing LLMs for high throughput and low latency inference.
  • •Comfort with LLM serving frameworks such as vLLM or SGLang, plus profiling performance down to the kernel level.
  • •A firm grasp of how GPUs work and behave.
Experience:AI infrastructureLLMsAI/ML pipelinesGPU accelerationInference systemsCustomer-facing ML deployments
Education:Bachelor's in Computer Science, Engineering, Mathematics, or a related field
Skills:CommunicationOwnershipProblem-solvingUrgencyTradeoff decision-making
Tech Stack:PythonC++VLLMSGLangCUDAGPUsDockerKubernetesLLMsAI/MLML models

Company Brief

Crusoe
Builds vertically integrated, energy-first AI infrastructure and purpose-built AI data centers (Crusoe Cloud), leveraging clean/stranded energy to power large-scale GPU compute for AI training and inference.
Industry: Data Centers
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Decacorn (USD 10B+)
Funding: Series E+
Headquarters: Denver, United States
Founded: 2018
Glassdoor
Glassdoor: 3.7
WebsiteLinkedInGlassdoor