Student Researcher (Seed – Multimodal Interaction & World Model - RL Focused) – 2026 Start (PhD)

ByteDance
San Jose
Workplace: OnsiteFull timeFunction: Research & Scientific (R&D)Education: phdSkills: []

Join the Seed Multimodal Interaction and World Model team to build and evaluate large-scale multimodal foundation models. You’ll design reinforcement learning (RL) training systems, develop unified frameworks that combine video, audio, and language for visual latent reasoning, and explore RL methods that improve multimodal understanding and generation. Work with researchers to assess world modeling, reasoning, and instruction-conditioned generation performance.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
2 hours ago

Student Researcher (Seed – Multimodal Interaction & World Model - RL Focused) – 2026 Start (PhD)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 2 hours agoStatus: Live

Job Summary

Join the Seed Multimodal Interaction and World Model team to build and evaluate large-scale multimodal foundation models. You’ll design reinforcement learning (RL) training systems, develop unified frameworks that combine video, audio, and language for visual latent reasoning, and explore RL methods that improve multimodal understanding and generation. Work with researchers to assess world modeling, reasoning, and instruction-conditioned generation performance.
Location: San Jose
Workplace: Onsite
Employment Type: Full time
Job Function: Research & Scientific (R&D)
Seniority: Intern level

Key Responsibilities

  • •Design and implement reinforcement learning (RL) training systems for large-scale multimodal foundation models.
  • •Develop unified modeling frameworks integrating video, audio, and language with a focus on visual latent reasoning.
  • •Explore RL-based approaches to improve understanding and generation for multimodal visual reasoning.
  • •Collaborate with researchers to evaluate models on world modeling, reasoning, and instruction-conditioned generation tasks.

Key Requirements

  • •Currently pursuing a PhD in Software Development, Computer Science, Computer Engineering, or a related technical discipline.
  • •Publications in accredited venues such as CVPR, ECCV, ICCV, NeurIPS, ICLR, or ICML.
  • •Strong research background in reinforcement learning, multimodal learning, video understanding, or vision-language modeling.
  • •Experience with reinforcement learning in multimodal or interactive environments (preferred).
  • •Solid programming skills building ML training or evaluation pipelines (preferred).
Experience:AI and MLMultimodal learningReinforcement learningVision-language modelingComputer vision
Education:PhD / Doctorate in Computer Science
Tech Stack:Reinforcement learningMultimodal learningVideo understandingVision-language modelingDiffusion modelsTransformersDistributed trainingCurriculum learningMemory-augmented transformers

Eligibility

Work Authorization:Authorization required. Sponsorship not provided.

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn