Research Scientist Project Intern (Multimodal Interaction & World Model) - 2026 Start (PhD)

ByteDance
Singapore
Workplace: OnsiteInternshipFunction: Data Science & Machine LearningEducation: phdSkills: ["Analytical thinking","Problem-solving","Communication","Collaboration","Independent exploration"]

Join the Multimodal Interaction & World Model team to research and develop multimodal intelligence for virtual/real-world interaction. You’ll work on large-scale multimodal understanding and generation, data construction and fine-tuning, evaluation systems, and advanced world-model capabilities like multimodal RAG, visual CoT, and agents for GUI/games and virtual worlds. This is a short-term project internship requiring at least 3 months commitment.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
1 month ago

Research Scientist Project Intern (Multimodal Interaction & World Model) - 2026 Start (PhD)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 29 days agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Join the Multimodal Interaction & World Model team to research and develop multimodal intelligence for virtual/real-world interaction. You’ll work on large-scale multimodal understanding and generation, data construction and fine-tuning, evaluation systems, and advanced world-model capabilities like multimodal RAG, visual CoT, and agents for GUI/games and virtual worlds. This is a short-term project internship requiring at least 3 months commitment.
Location: Singapore
Workplace: Onsite
Employment Type: Internship
Job Function: Data Science & Machine Learning
Seniority: Intern level

Key Responsibilities

  • •Research multimodal understanding, generative modeling, machine learning, reinforcement learning, AIGC, computer vision, and related AI technologies.
  • •Work on basic multimodal understanding/generation models, including extreme system optimization, data construction, instruction fine-tuning, preference alignment, and model optimization.
  • •Improve data synthesis, scalable oversight, model reasoning and planning, and build evaluation systems to assess large-model capabilities.
  • •Explore advanced multimodal and world-model capabilities such as multimodal RAG, visual CoT, and agent-based systems for GUI/games and virtual worlds.
  • •Use pre-training, simulation, and other techniques to model environments in virtual/real worlds and enable multimodal interactive exploration and new AI-driven product development.

Key Requirements

  • •PhD degree in computer, electronics, mathematics, or related majors.
  • •In-depth research in areas such as computer vision, multimodal, AIGC, machine learning, or rendering generation.
  • •Excellent analytical and problem-solving skills, including ability to tackle large-model training and application problems independently.
  • •Good communication and collaboration skills; proactive and able to work harmoniously with the team.
Education:PhD / Doctorate in computer, electronics, mathematics (or related majors)
Skills:Analytical thinkingProblem-solvingCommunicationCollaborationIndependent exploration
Tech Stack:Multimodal understandingGenerationMachine learningReinforcement learningAIGCComputer visionLarge modelsMultimodal RAGVisual CoTAgentsPre-trainingSimulationData constructionInstruction fine-tuningPreference alignmentModel optimizationPythonC/C++

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn