AI Infra Engineer - Large Model Training Infrastructure (LLM/VLM /Agent RL)
San Jose
Workplace: OnsiteFull timeFunction: Education & TrainingExperience: 2+ yearsEducation: bachelorsSkills: ["Debugging","Problem-solving"]Build and evolve large-model training infrastructure for post-training multimodal learning and reinforcement learning. Design and optimize distributed training strategies for 100B–1T parameter models, covering data/model parallelism, operator fusion, memory optimization, and cluster-level performance. Develop training and evaluation systems for Reasoning RL and Agent RL, including benchmarks and rollout efficiency, and enable multimodal training across image, text, audio, and video with correctness and convergence validation.

