Research Scientist - Seed Multimodal Interaction and World Model

ByteDance
San Jose
Workplace: OnsiteFull timeFunction: Data Science & Machine LearningSkills: []

Build and advance large-scale multimodal foundation models focused on visual latent reasoning and human-level multimodal understanding. Develop unified frameworks integrating video, audio, and language, and explore reinforcement learning approaches to improve multimodal visual reasoning and generation. Collaborate with researchers to evaluate world modeling, reasoning, and instruction-conditioned generation tasks, shaping next-generation multimodal assistant capabilities.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
1 month ago

Research Scientist - Seed Multimodal Interaction and World Model

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 29 days agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Build and advance large-scale multimodal foundation models focused on visual latent reasoning and human-level multimodal understanding. Develop unified frameworks integrating video, audio, and language, and explore reinforcement learning approaches to improve multimodal visual reasoning and generation. Collaborate with researchers to evaluate world modeling, reasoning, and instruction-conditioned generation tasks, shaping next-generation multimodal assistant capabilities.
Location: San Jose
Workplace: Onsite
Employment Type: Full time
Job Function: Data Science & Machine Learning

Key Responsibilities

  • •Research and develop large-scale multimodal foundation models.
  • •Develop unified modeling frameworks integrating video, audio, and language with an emphasis on visual latent reasoning.
  • •Explore reinforcement learning-based methods to connect understanding and generation for multimodal visual reasoning.
  • •Collaborate with researchers to evaluate models on world modeling, reasoning, and instruction-conditioned generation tasks.

Key Requirements

  • •Hold a Master's or PhD in Software Development, Computer Science, Computer Engineering, or a related technical field.
  • •Have publications in accredited venues such as CVPR, ECCV, ICCV, NeurIPS, ICLR, ICML, or similar top conferences.
  • •Demonstrate strong research background in reinforcement learning, multimodal learning, video understanding, or vision-language modeling.
  • •Have experience with reinforcement learning in multimodal or interactive environments (preferred).
  • •Show familiarity with video generation or diffusion-based generative models and experience with large-scale model training (preferred).
Education:
Tech Stack:Reinforcement learningMultimodal learningVideo understandingVision-language modelingVideo generationDiffusion-based generative modelsLarge-scale model trainingTraining pipelinesEvaluation pipelinesWorld modelingInstruction-conditioned generation

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn