Student Researcher (Multimodal Interaction and World Model - Seed) – 2027 Start (PhD)

ByteDance
San Jose
Workplace: OnsiteFull timeFunction: Research & Scientific (R&D)Education: phdSkills: ["Problem-solving","Independent research","Collaboration"]

Work with the Seed Multimodal Interaction and World Model team to advance multimodal foundation model research for human-level understanding and interaction. You’ll conduct research on multimodal learning systems, explore vision-language and world modeling approaches, and design/prototype algorithms and model components. Collaborate to push research directions, leveraging a strong programming foundation and an active PhD research track with publications in relevant ML/AI areas.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
1 month ago

Student Researcher (Multimodal Interaction and World Model - Seed) – 2027 Start (PhD)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 30 days agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Work with the Seed Multimodal Interaction and World Model team to advance multimodal foundation model research for human-level understanding and interaction. You’ll conduct research on multimodal learning systems, explore vision-language and world modeling approaches, and design/prototype algorithms and model components. Collaborate to push research directions, leveraging a strong programming foundation and an active PhD research track with publications in relevant ML/AI areas.
Location: San Jose
Workplace: Onsite
Employment Type: Full time
Job Function: Research & Scientific (R&D)
Seniority: Graduate level

Key Responsibilities

  • •Conduct research on multimodal foundation models and related systems.
  • •Explore methods to improve model capabilities across modalities, including vision-language modeling, world modeling, and representation learning.
  • •Design and prototype algorithms, models, or system components.
  • •Collaborate with the team to advance research directions.

Key Requirements

  • •Currently pursuing a PhD in computer science, mathematics, engineering, or a related field.
  • •Strong programming skills with solid algorithms and data structures; proficient in Python or C/C++.
  • •Demonstrated research track record with publications in conferences related to multimodal learning, machine learning, or AI.
  • •Experience in multimodal modeling, including vision-language modeling, world modeling, simulation, or 3D representations.
  • •Strong problem-solving skills and experience conducting independent research while collaborating effectively in a research environment.
Experience:Multimodal modelingVision-language modelingWorld modelingMachine learningArtificial intelligenceSimulation3D representations
Education:PhD / Doctorate
Skills:Problem-solvingIndependent researchCollaboration
Tech Stack:PythonC/C++

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn