AI Infra Engineer - Large Model Inference Systems (Multimodal/LLM/VLM)
San Jose
Workplace: OnsiteFull timeFunction: Data Science & Machine LearningExperience: 2+ yearsEducation: bachelorsSkills: ["System design","Performance optimization","Production reliability","Load balancing","Resource management"]Build and evolve inference infrastructure for ultra-large-scale LLM, VLM, and multimodal models. Work on global scheduling across heterogeneous compute, high-concurrency load balancing, efficient batch formation, and low-latency online serving. Optimize distributed inference for 200B+ models using TP/EP/DP strategies, develop high-performance CUDA/Triton kernels for advanced architectures (e.g., MoE), and improve production reliability through deployment pipelines and intelligent operations.

