ML Model Serving Engineer

Sesame
San Francisco, New York, Bellevue
Workplace: OnsiteFull timeUSD 175,000 - 280,000 annuallyFunction: Solutions Engineering & Sales EngineeringSkills: ["Problem-solving","Communication","Collaboration"]

Lead the design and optimization of a high-throughput ML model serving stack for Sesame, spanning LLM, speech, and vision models. Collaborate with ML infrastructure and training engineers to deliver fast, cost-efficient inference, extending frameworks like VLLM and SGLang, using techniques such as in-flight batching, caching, and custom kernels to minimize latency and initialization time.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Sesame
Sesame
1 year ago

ML Model Serving Engineer

âś“ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 1 hour agoStatus: Live

Job Summary

Lead the design and optimization of a high-throughput ML model serving stack for Sesame, spanning LLM, speech, and vision models. Collaborate with ML infrastructure and training engineers to deliver fast, cost-efficient inference, extending frameworks like VLLM and SGLang, using techniques such as in-flight batching, caching, and custom kernels to minimize latency and initialization time.
Location: San Francisco, New York, Bellevue
Workplace: Onsite
Employment Type: Full time
Job Function: Solutions Engineering & Sales Engineering

Key Responsibilities

  • •Turbocharge our serving layer, consisting of a variety of LLM, speech, and vision models
  • •Partner with ML infrastructure and training engineers to build a fast, cost-effective, accurate, and reliable serving layer to power a new consumer product category
  • •Modify and extend LLM serving frameworks like VLLM and SGLang to take advantage of the latest techniques in high-performance model serving
  • •Work with the training team to identify opportunities to produce faster models without sacrificing quality
  • •Use techniques like in-flight batching, caching, and custom kernels to speed up inference

Pay and Benefits

Salary: USD 175,000 - 280,000 annually
Perks:401kHealth InsuranceVisionDentalPaid LeaveMedical Fsa

Key Requirements

  • •Expert in differentiable array computing framework, preferably PyTorch
  • •Expert in optimizing machine learning models for serving reliably at high throughput, with low latency
  • •Significant systems programming experience; ex. Experience working on high-performance server systems—you’d be just as comfortable with the internals of VLLM as you would with a complex PyTorch codebase
  • •Significant performance engineering experience; ex. Bottleneck analysis in high-scale server systems or profiling low-level systems code
  • •Always up to date on the latest techniques for model serving optimization
Experience:Machine learningAI
Skills:Problem-solvingCommunicationCollaboration
Tech Stack:PyTorchVLLMSGlangKubernetesRayAWSGCPAzure

Company Brief

Sesame
Operates an online marketplace connecting patients with affordable direct-pay healthcare services, offering telemedicine and in-person appointments with doctors, dentists, therapists, and specialists without insurance, focusing on transparent pricing and accessible care.
Industry: HealthTech
Company Size: Medium (51 to 250 employees)
Growth: Growth Stage Startup
Headquarters: New York, United States
Founded: 2016
Website