Research Scientist Graduate (Seed-Speech Foundation Model) - 2027 Start (PhD)

ByteDance
San Jose
Workplace: OnsiteFull timeFunction: Data Science & Machine LearningEducation: phdSkills: ["Problem-solving","Collaboration"]

Join the Seed Speech team to develop and scale speech foundation models for understanding and generation. You’ll design training pipelines (data construction, instruction tuning, and model alignment), improve capabilities like speech recognition, synthesis, reasoning, and robustness, and optimize architectures for training efficiency and system performance. The role also includes exploring natural, interactive interfaces for speech-based systems as part of cutting-edge speech, audio, and multimodal deep learning research.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
1 month ago

Research Scientist Graduate (Seed-Speech Foundation Model) - 2027 Start (PhD)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 30 days agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Join the Seed Speech team to develop and scale speech foundation models for understanding and generation. You’ll design training pipelines (data construction, instruction tuning, and model alignment), improve capabilities like speech recognition, synthesis, reasoning, and robustness, and optimize architectures for training efficiency and system performance. The role also includes exploring natural, interactive interfaces for speech-based systems as part of cutting-edge speech, audio, and multimodal deep learning research.
Location: San Jose
Workplace: Onsite
Employment Type: Full time
Job Function: Data Science & Machine Learning
Seniority: Graduate level

Key Responsibilities

  • •Develop and scale speech foundation models for understanding and generation tasks.
  • •Design training pipelines including data construction, instruction tuning, and model alignment.
  • •Improve core capabilities such as speech recognition, synthesis, reasoning, and robustness.
  • •Optimize model architectures, training efficiency, and system performance.
  • •Explore natural and interactive interfaces for speech-based systems.

Key Requirements

  • •Currently pursuing a PhD in computer science, mathematics, engineering, or a related field, expected to graduate in 2027, and able to commit to an onboarding date by the end of 2027.
  • •Excellent coding ability with data structures and fundamental algorithm skills, proficient in C/C++ or Python.
  • •Experience in speech processing, audio modeling, or related areas.
  • •Familiarity with deep learning approaches for speech understanding and generation.
  • •Strong research track record in relevant areas, with strong problem-solving and collaboration skills.
Education:PhD / Doctorate
Skills:Problem-solvingCollaboration
Tech Stack:CC++Python

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn