Staff Applied AI Inference Engineer

Crusoe
San Francisco
Workplace: OnsiteFull timeUSD 215,000 - 260,000 annuallyFunction: Data Science & Machine LearningSkills: ["Communication","Problem-solving","Ownership","Urgency"]

Own the end-to-end AI inference stack to make large language models faster, cheaper, and more reliable in production. Profile latency and cost drivers, apply modern optimization techniques, and dive into the serving stack down to CUDA kernels. Partner with customer engineering teams to tailor deployments, move workloads from proof of concept to monitored services, and deliver measurable performance gains using Python-first development.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Crusoe
Crusoe
3 days ago

Staff Applied AI Inference Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 4 hours agoStatus: Live

Job Summary

Own the end-to-end AI inference stack to make large language models faster, cheaper, and more reliable in production. Profile latency and cost drivers, apply modern optimization techniques, and dive into the serving stack down to CUDA kernels. Partner with customer engineering teams to tailor deployments, move workloads from proof of concept to monitored services, and deliver measurable performance gains using Python-first development.
Location: San Francisco
Workplace: Onsite
Employment Type: Full time
Job Function: Data Science & Machine Learning
Seniority: Mid level

Key Responsibilities

  • •Bring current inference techniques into production and refine them.
  • •Design and optimize serving architectures, including prefill/decode disaggregation and request routing approaches.
  • •Work down into the serving stack, profiling and analyzing performance down to kernel level to find and fix issues.
  • •Profile and tune deployments to targets for latency, throughput, and cost, and keep them dependable under real traffic.
  • •Tailor deployments to customer models and constraints, taking workloads from proof of concept to live, well-monitored production services.

Pay and Benefits

Salary: USD 215,000 - 260,000 annually
Equity and Bonus:Equity
Perks:Health InsuranceDentalVisionRsusPaid LeaveParental Leave401k

Key Requirements

  • •Bachelor’s, Master’s, or Ph.D. in Computer Science, Engineering, Mathematics, or a related field.
  • •Hands-on production experience shipping code in one or more general-purpose languages, with strong preference for Python.
  • •Experience optimizing LLMs for high-throughput/low-latency inference.
  • •Comfort using modern LLM serving frameworks like vLLM or SGLang and profiling performance down to kernel level.
  • •Knowledge of GPUs and hands-on interest/experience with large language models and deploying ML models end to end.
Experience:AI/MLLarge language modelsML pipelinesLLM serving
Skills:CommunicationProblem-solvingOwnershipUrgency
Tech Stack:PythonC++VLLMSGLangCUDAGPUDockerKubernetesLLMsAI/ML pipelines

Company Brief

Crusoe
Builds vertically integrated, energy-first AI infrastructure and purpose-built AI data centers (Crusoe Cloud), leveraging clean/stranded energy to power large-scale GPU compute for AI training and inference.
Industry: Data Centers
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Decacorn (USD 10B+)
Funding: Series E+
Headquarters: Denver, United States
Founded: 2018
Glassdoor
Glassdoor: 3.7
WebsiteLinkedInGlassdoor