ML Model Serving Engineer
San Francisco, New York, Bellevue
Workplace: OnsiteFull timeUSD 175,000 - 280,000 annuallyFunction: Solutions Engineering & Sales EngineeringSkills: ["Problem-solving","Communication","Collaboration"]Lead the design and optimization of a high-throughput ML model serving stack for Sesame, spanning LLM, speech, and vision models. Collaborate with ML infrastructure and training engineers to deliver fast, cost-efficient inference, extending frameworks like VLLM and SGLang, using techniques such as in-flight batching, caching, and custom kernels to minimize latency and initialization time.
Loading
Loading job details...
Preparing the role view and application actions.

