Machine Learning Engineer, Inference & Serving (Speech LLM) - San Francisco
San Francisco
Workplace: HybridFull timeUSD 180,000 - 270,000 annuallyFunction: Data Science & Machine LearningSkills: ["Python","CUDA","GPU","GPUs","TensorRT","Triton","VLLM","SGLang","Kubernetes","WebSockets","WebRTC"]Build and optimize high-throughput, ultra-low-latency inference engines for large language models and speech systems. Collaborate with ML training and backend infra teams to maximize latency/throughput, KV cache performance, and real-time streaming efficiency on multi-GPU clusters, in a hybrid, San Francisco-based setting.
Loading
Loading job details...
Preparing the role view and application actions.

