Research Scientist Graduate (Multimodal Interaction & World Model) - 2026 Start (PhD)

ByteDance
Singapore
Workplace: OnsiteFull timeFunction: Data Science & Machine LearningEducation: phdSkills: ["Analytical problem-solving","Communication","Collaboration","Independent exploration","Proactive work"]

Join the Multimodal Interaction & World Model team to advance multimodal intelligence for virtual/real-world interaction. You’ll explore large-scale multimodal understanding and generation, work on model/data construction, fine-tuning, alignment, and evaluation, and push capabilities like multimodal RAG, visual COT, and agentic systems. Using pre-training and simulation, you’ll develop technologies and AI-powered products that improve interactive exploration and user experiences.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
1 month ago

Research Scientist Graduate (Multimodal Interaction & World Model) - 2026 Start (PhD)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 29 days agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Join the Multimodal Interaction & World Model team to advance multimodal intelligence for virtual/real-world interaction. You’ll explore large-scale multimodal understanding and generation, work on model/data construction, fine-tuning, alignment, and evaluation, and push capabilities like multimodal RAG, visual COT, and agentic systems. Using pre-training and simulation, you’ll develop technologies and AI-powered products that improve interactive exploration and user experiences.
Location: Singapore
Workplace: Onsite
Employment Type: Full time
Job Function: Data Science & Machine Learning
Seniority: Graduate level

Key Responsibilities

  • •Research multimodal understanding, generative machine learning, reinforcement learning, AIGC, computer vision, and related AI technologies.
  • •Develop and optimize large/ultra-large multimodal understanding and generation models, including data construction, instruction fine-tuning, preference alignment, and model optimization.
  • •Advance multimodal model and world-model capabilities such as multimodal RAG, visual COT, and agent systems.
  • •Build evaluation systems and improve large-model reasoning and planning abilities.
  • •Use pre-training and simulation to model virtual/real environments and enable multimodal interactive exploration and application development.

Key Requirements

  • •Final year Ph.D or recent Ph.D graduate in Computer Science, Software Engineering, Electronics, Mathematics, or related majors.
  • •In-depth research in areas such as computer vision, multimodal, AIGC, machine learning, or rendering generation.
  • •Excellent analytical and problem-solving skills, with the ability to solve large-model training and application problems independently.
  • •Good communication and collaboration skills, with a proactive, cooperative approach to exploring new technologies.
  • •Ability to explore and improve large-model reasoning/planning and build comprehensive evaluation systems.
Experience:ResearchComputer visionMultimodalAIGCMachine learningReinforcement learningLarge modelsWorld models
Education:PhD / Doctorate
Skills:Analytical problem-solvingCommunicationCollaborationIndependent explorationProactive work
Tech Stack:PythonC/C++

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn