Staff Applied AI Inference Engineer

Crusoe
Denver
Workplace: OnsiteFull timeUSD 185,000 - 225,000Function: Data Science & Machine LearningSkills: ["Communication","Problem-solving","Opportunity-finding","Urgency","Ownership"]

Own the end-to-end AI inference stack to make large language models faster, cheaper, and more reliable in production. Bring modern inference optimizations into real deployments by profiling latency and cost, working down into serving code and CUDA kernels, and tuning for GPU performance. Partner with customer engineering teams to take workloads from proof of concept to monitored production services while keeping targets for throughput, latency, and dependability.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Crusoe
Crusoe
2 days ago

Staff Applied AI Inference Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 2 hours agoStatus: Live

Job Summary

Own the end-to-end AI inference stack to make large language models faster, cheaper, and more reliable in production. Bring modern inference optimizations into real deployments by profiling latency and cost, working down into serving code and CUDA kernels, and tuning for GPU performance. Partner with customer engineering teams to take workloads from proof of concept to monitored production services while keeping targets for throughput, latency, and dependability.
Location: Denver
Workplace: Onsite
Employment Type: Full time
Job Function: Data Science & Machine Learning
Seniority: Sr. Director level

Key Responsibilities

  • •Own the inference stack end to end: profile where time and cost go and apply optimization techniques in real deployments.
  • •Design and optimize serving architectures (e.g., prefill/decode disaggregation and request routing) for performance targets.
  • •Work down into the serving stack using frameworks (vLLM, SGLang) and kernel-level profiling to diagnose and fix performance problems.
  • •Adapt and scale optimization methods across many ML models with emphasis on large language models.
  • •Tailor deployments to each customer’s models and constraints, moving workloads from proof of concept to live, well-monitored production services.

Pay and Benefits

Salary: USD 185,000 - 225,000
Perks:Health InsuranceDentalVision401kHsaPaid ParentalPaid LeaveLong-term DisabilityMeal AllowancePaid HolidaysLife InsuranceCell PhoneVolunteer Time

Key Requirements

  • •Bachelor's, Master's, or Ph.D. in Computer Science, Engineering, Mathematics, or a related field.
  • •Hands-on experience shipping production code in one or more general-purpose languages, with a strong preference for Python.
  • •Familiarity with methods for optimizing LLMs for high throughput/low latency inference.
  • •Comfort with modern LLM serving frameworks such as vLLM or SGLang and profiling performance down to the kernel level.
  • •A firm grasp of how GPUs are built and how they behave, plus a working knowledge of AI/ML pipelines for developing and deploying models.
Education:
Skills:CommunicationProblem-solvingOpportunity-findingUrgencyOwnership
Tech Stack:PythonC++VLLMSGLangCUDADockerKubernetesGPULarge language modelsAI/ML pipelines

Company Brief

Crusoe
Builds vertically integrated, energy-first AI infrastructure and purpose-built AI data centers (Crusoe Cloud), leveraging clean/stranded energy to power large-scale GPU compute for AI training and inference.
Industry: Data Centers
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Decacorn (USD 10B+)
Funding: Series E+
Headquarters: Denver, United States
Founded: 2018
Glassdoor
Glassdoor: 3.7
WebsiteLinkedInGlassdoor