Research Scientist Graduate (Seed-Vision Foundation Model) - 2027 Start (PhD)

ByteDance
San Jose
Workplace: OnsiteFull timeFunction: Data Science & Machine LearningEducation: phdSkills: ["Problem-solving","Communication"]

Build and scale vision foundation models for image and video, advancing multimodal generative modeling in GenAI. You’ll design end-to-end data pipelines, pre- and post-training strategies, and improve perception, reasoning, and multimodal understanding. Partner with the team to optimize model architectures, training efficiency, and evaluation frameworks, while exploring real-world multimodal applications.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
1 month ago

Research Scientist Graduate (Seed-Vision Foundation Model) - 2027 Start (PhD)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 30 days agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Build and scale vision foundation models for image and video, advancing multimodal generative modeling in GenAI. You’ll design end-to-end data pipelines, pre- and post-training strategies, and improve perception, reasoning, and multimodal understanding. Partner with the team to optimize model architectures, training efficiency, and evaluation frameworks, while exploring real-world multimodal applications.
Location: San Jose
Workplace: Onsite
Employment Type: Full time
Job Function: Data Science & Machine Learning
Seniority: Graduate level

Key Responsibilities

  • •Develop and scale vision foundation models across image and video modalities.
  • •Design data pipelines, pre-training strategies, and post-training methods for vision tasks.
  • •Improve core capabilities such as perception, reasoning, and multimodal understanding.
  • •Optimize model architectures, training efficiency, and evaluation frameworks.
  • •Explore real-world applications of vision models in multimodal systems.

Key Requirements

  • •Currently pursuing a PhD in computer science, mathematics, engineering, or a related field, with an expected graduation date in 2027 and the ability to commit to an onboarding date by the end of 2027.
  • •Strong coding ability with data structures and fundamental algorithms; proficient in C/C++ or Python.
  • •Experience with computer vision, multimodal learning, or large-scale model training.
  • •Strong understanding of deep learning architectures and training methodologies.
  • •Demonstrated research impact through impactful papers or projects.
Experience:Computer visionMultimodal learningLarge-scale model trainingGenAI
Education:PhD / Doctorate
Skills:Problem-solvingCommunication
Tech Stack:CC++Python

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn