Research Scientist Graduate (Multimodal Interaction and World Model) - 2026 Start (PhD)

ByteDance
San Jose
Workplace: OnsiteFull timeFunction: Data Science & Machine LearningEducation: phdSkills: ["Research","Problem-solving","Systematic experimentation"]

Join the ByteDance Seed Multimodal Interaction and World Model team to drive research and engineering that improves multimodal understanding and strengthens reasoning. You’ll explore ideas that raise both model performance and efficiency, and develop scaling laws supported by systematic ablations. Work on research areas spanning reinforcement learning, multimodal learning, video understanding, and vision-language modeling, applying techniques for training multimodal models.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
6 hours ago

Research Scientist Graduate (Multimodal Interaction and World Model) - 2026 Start (PhD)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 6 hours agoStatus: Live

Job Summary

Join the ByteDance Seed Multimodal Interaction and World Model team to drive research and engineering that improves multimodal understanding and strengthens reasoning. You’ll explore ideas that raise both model performance and efficiency, and develop scaling laws supported by systematic ablations. Work on research areas spanning reinforcement learning, multimodal learning, video understanding, and vision-language modeling, applying techniques for training multimodal models.
Location: San Jose
Workplace: Onsite
Employment Type: Full time
Job Function: Data Science & Machine Learning
Seniority: Graduate level

Key Responsibilities

  • •Drive research and engineering to advance models that enhance understanding of multimodal data and expand reasoning capabilities.
  • •Explore research ideas that optimize both model performance and efficiency.
  • •Establish scaling laws and design and conduct systematic ablations to produce transferrable conclusions.

Key Requirements

  • •Completing or recently completed a PhD in Software Development, Computer Science, Computer Engineering, or a related technical discipline.
  • •Publications in accredited AI/ML venues such as CVPR, ECCV, ICCV, NeurIPS, ICLR, or ICML.
  • •Strong research background in at least one of reinforcement learning, multimodal learning, video understanding, or vision-language modeling.
  • •Expertise with Transformers (Dense and MoE) and experience scaling Transformers on GPUs or TPUs.
  • •Hands-on experience with PyTorch or JAX and distributed training frameworks, plus familiarity with state-of-the-art multimodal training data preparation.
Education:PhD / Doctorate in Software Development, Computer Science, Computer Engineering, or a related technical discipline
Skills:ResearchProblem-solvingSystematic experimentation
Tech Stack:TransformersMoEGPUTPUPyTorchJAXDistributed trainingMultimodal training data

Eligibility

Work Authorization:Authorization required. Sponsorship not provided.

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn