Member of Technical Staff, Inference & Serving
San Mateo
Workplace: OnsiteFull timeFunction: Solutions Engineering & Sales EngineeringSkills: ["Collaboration","Problem-solving","Communication"]Design, optimize, and scale high-performance model serving systems for diffusion LLMs in production, focusing on low-latency inference, distributed orchestration, and reliable, cost-efficient operation. Collaborate with ML researchers to translate architectural advances into production-ready serving improvements, and implement robust monitoring and deployment workflows to meet SLA requirements.

