Research Scientist Graduate (Seed-Speech Foundation Model) - 2027 Start

ByteDance
San Jose
Workplace: OnsiteInternshipFunction: Data Science & Machine LearningSkills: ["Problem-solving","Collaboration"]

Join the Seed Speech team to enrich interactive experiences with multimodal speech technologies. You’ll develop and scale speech foundation models for understanding and generation, design training pipelines (data construction, instruction tuning, and model alignment), and improve capabilities like speech recognition, synthesis, reasoning, and robustness. Work on optimizing model architectures, training efficiency, and system performance while exploring natural, interactive interfaces for speech-based systems.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
1 month ago

Research Scientist Graduate (Seed-Speech Foundation Model) - 2027 Start

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 30 days agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Join the Seed Speech team to enrich interactive experiences with multimodal speech technologies. You’ll develop and scale speech foundation models for understanding and generation, design training pipelines (data construction, instruction tuning, and model alignment), and improve capabilities like speech recognition, synthesis, reasoning, and robustness. Work on optimizing model architectures, training efficiency, and system performance while exploring natural, interactive interfaces for speech-based systems.
Location: San Jose
Workplace: Onsite
Employment Type: Internship
Job Function: Data Science & Machine Learning
Seniority: Graduate level

Key Responsibilities

  • •Develop and scale speech foundation models for understanding and generation tasks.
  • •Design training pipelines including data construction, instruction tuning, and model alignment.
  • •Improve core capabilities such as speech recognition, synthesis, reasoning, and robustness.
  • •Optimize model architectures, training efficiency, and system performance.
  • •Explore natural and interactive interfaces for speech-based systems.

Key Requirements

  • •Currently pursuing a Bachelor's or Master's degree in computer science, mathematics, engineering, or a related field, expected to graduate in 2027, and able to commit to onboarding by end of 2027.
  • •Excellent coding ability, strong data structures, and fundamental algorithm skills; proficient in C/C++ or Python.
  • •Demonstrated interest or project experience in relevant areas.
  • •Experience in speech processing or audio modeling through internships is preferred.
  • •Strong problem-solving and collaboration skills.
Education:
Skills:Problem-solvingCollaboration
Tech Stack:C/C++Python

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn