Senior AI Infra Engineer - Large Model Training Infrastructure (LLM/VLM /Agent RL)
San Jose
Full timeFunction: Education & TrainingExperience: 4+ yearsEducation: bachelorsSkills: ["Debugging","Performance optimization","Problem-solving"]Build and evolve unified training infrastructure for large language, vision-language, and agentic models, spanning post-training workflows, modalities, and training paradigms. Design and optimize distributed training strategies for 100B–1T parameter models, including parallelism (DP/TP/PP/EP), operator fusion, memory optimization, and cluster-level performance improvements. Develop training and evaluation systems for Reasoning RL and Agent RL, and enable multimodal training across image, text, audio, and video with correctness and convergence validation.

