Senior AI Infra Engineer - Large Model Inference Systems (Multimodal/LLM/VLM)
San Jose
Workplace: OnsiteFull timeFunction: Data Science & Machine LearningEducation: bachelorsSkills: ["System design","Performance optimization","Production system development","Distributed scheduling","Load balancing"]Build and evolve high-performance inference infrastructure for ultra-large multimodal models, enabling scalable distributed serving with low-latency. Own next-generation online serving components like global scheduling across heterogeneous compute, high-concurrency load balancing, and efficient batch formation. Optimize inference for 200B+ models using parallelism strategies and develop CUDA/Triton kernels for frontier architectures such as MoE and multimodal fusion layers. Improve production reliability through AI-driven infrastructure and deployment pipelines.

