Large Language Model Inference System Engineer Graduate (Applied Machine Learning) - 2027 Start
San Jose
Workplace: OnsiteFull timeFunction: Data Science & Machine LearningSkills: ["Coding","Performance analysis","Distributed systems interest","Systems thinking"]Develop and optimize an end-to-end large model inference (MaaS) system for ultra-large-scale, heterogeneous GPU clusters, improving inference performance, stability, and total cost. Use system-level techniques such as disaggregated multi-role inference and distributed KV cache systems, plus heterogeneous and elastic computing and multi-tenant co-located inference. Work within the machine learning platform team supporting training and inference for recommendation, advertising, vision, speech, and NLP scenarios.
Loading
Loading job details...
Preparing the role view and application actions.

