Research Scientist - Seed Multimodal Interaction and World Model
ByteDance
San Jose
Workplace: OnsiteFull timeFunction: Data Science & Machine LearningSkills: []Build and advance large-scale multimodal foundation models focused on visual latent reasoning and human-level multimodal understanding. Develop unified frameworks integrating video, audio, and language, and explore reinforcement learning approaches to improve multimodal visual reasoning and generation. Collaborate with researchers to evaluate world modeling, reasoning, and instruction-conditioned generation tasks, shaping next-generation multimodal assistant capabilities.

