Senior Staff Applied AI Inference Engineer

Crusoe
San Francisco
Workplace: OnsiteFull timeUSD 250,000 - 300,000 annuallyFunction: Data Science & Machine LearningEducation: bachelorsSkills: ["Problem-solving","Opportunity-finding","Urgency","Communication","Ownership"]

Build and own the inference stack end to end to make large language model serving faster, cheaper, and more reliable in production. Profile where time and cost go, apply modern optimization techniques, and work deep into serving code—from vLLM and SGLang down to CUDA kernels. Partner with customer engineering teams to tailor deployments to real models, traffic, latency, and cost constraints, taking workloads from proof of concept to monitored production services.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Crusoe
Crusoe
4 days ago

Senior Staff Applied AI Inference Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 7 hours agoStatus: Live

Job Summary

Build and own the inference stack end to end to make large language model serving faster, cheaper, and more reliable in production. Profile where time and cost go, apply modern optimization techniques, and work deep into serving code—from vLLM and SGLang down to CUDA kernels. Partner with customer engineering teams to tailor deployments to real models, traffic, latency, and cost constraints, taking workloads from proof of concept to monitored production services.
Location: San Francisco
Workplace: Onsite
Employment Type: Full time
Job Function: Data Science & Machine Learning
Seniority: Mid level

Key Responsibilities

  • •Own the inference stack end to end by profiling time/cost bottlenecks and applying optimization techniques to production deployments.
  • •Design and optimize serving architectures, including prefill/decode disaggregation and request routing.
  • •Work down into the serving stack from frameworks (vLLM, SGLang) to CUDA kernels to find and fix performance issues.
  • •Adapt and scale optimization methods across different ML models with a focus on large language models.
  • •Partner with customer engineering teams to tailor deployments to model/traffic/latency/cost constraints and move workloads from proof of concept to monitored production services.

Pay and Benefits

Salary: USD 250,000 - 300,000 annually
Equity and Bonus:Equity
Perks:Health InsuranceDentalVisionPaid LeaveParental Leave401kEquityHsa

Key Requirements

  • •Bachelor’s, Master’s, or Ph.D. in Computer Science, Engineering, Mathematics, or a related field.
  • •Hands-on experience shipping production code in one or more general-purpose languages, with strong preference for Python.
  • •Familiarity with optimizing LLMs for high throughput and low latency inference.
  • •Comfort with modern LLM serving frameworks such as vLLM or SGLang, plus profiling down to kernel-level performance.
  • •Strong understanding of GPUs and hands-on interest/experience with large language models.
Experience:AI/MLLarge language modelsLLM inferenceProduction softwareCustomer-facingHigh throughputLow latency
Education:Bachelor's in Computer Science, Engineering, Mathematics, or a related field
Skills:Problem-solvingOpportunity-findingUrgencyCommunicationOwnership
Tech Stack:PythonC++VLLMSGLangCUDAGPUsDockerKubernetes

Company Brief

Crusoe
Builds vertically integrated, energy-first AI infrastructure and purpose-built AI data centers (Crusoe Cloud), leveraging clean/stranded energy to power large-scale GPU compute for AI training and inference.
Industry: Data Centers
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Valuation: Decacorn (USD 10B+)
Funding: Series E+
Headquarters: Denver, United States
Founded: 2018
Glassdoor
Glassdoor: 3.7
WebsiteLinkedInGlassdoor