Research Engineer – Multimodal Training Infrastructure (Seed Infra)
ByteDance
San Jose
Workplace: OnsiteFull timeFunction: Education & TrainingSkills: []Build large-scale multimodal training infrastructure for AI foundation models, focusing on efficient distributed training for multimodal LLMs and image/video generation. Design and optimize parallelism strategies, throughput scaling, and cross-layer performance improvements across GPUs and the training stack. Improve reliability with fast checkpointing and fault tolerance, diagnose failures, and analyze exascale bottlenecks to propose data-driven optimizations for production deployment.

