Research Scientist Graduate (Seed Vision Foundation Model) - 2027 Start

ByteDance
San Jose
Workplace: OnsiteInternshipFunction: Data Science & Machine LearningSkills: ["Problem-solving","Communication","Coding"]

Join the Seed Vision team working on vision foundation models for visual generation and multimodal generative AI. You’ll develop and scale models across image and video, design data pipelines and pre-/post-training strategies, and improve core capabilities like perception, reasoning, and multimodal understanding. The role also focuses on optimizing architectures and training efficiency and exploring real-world applications of vision models within multimodal systems.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
1 month ago

Research Scientist Graduate (Seed Vision Foundation Model) - 2027 Start

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 30 days agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Join the Seed Vision team working on vision foundation models for visual generation and multimodal generative AI. You’ll develop and scale models across image and video, design data pipelines and pre-/post-training strategies, and improve core capabilities like perception, reasoning, and multimodal understanding. The role also focuses on optimizing architectures and training efficiency and exploring real-world applications of vision models within multimodal systems.
Location: San Jose
Workplace: Onsite
Employment Type: Internship
Job Function: Data Science & Machine Learning
Seniority: Intern level

Key Responsibilities

  • •Develop and scale vision foundation models across image and video modalities.
  • •Design data pipelines, pre-training strategies, and post-training methods for vision tasks.
  • •Improve core capabilities such as perception, reasoning, and multimodal understanding.
  • •Optimize model architectures, training efficiency, and evaluation frameworks.
  • •Explore real-world applications of vision models in multimodal systems.

Key Requirements

  • •Currently pursuing a Bachelor's or Master's degree in computer science, mathematics, engineering, or a related field, with an expected graduation date in 2027, and able to commit to an onboarding date by the end of 2027.
  • •Excellent coding ability with strong data structures and fundamental algorithm skills.
  • •Proficient in C/C++ or Python.
  • •Demonstrated interest or project experience in relevant areas.
  • •Experience with computer vision, multimodal learning, or machine learning through internships (preferred).
Experience:Computer visionMultimodal learningMachine learningGenAI
Education:
Skills:Problem-solvingCommunicationCoding
Tech Stack:C/C++Python

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn