Research Scientist Graduate (Foundation Model-Speech-Interaction & Learning) - 2026 Start (PhD)

ByteDance
San Jose
Workplace: OnsiteFull timeFunction: Data Science & Machine LearningEducation: phdSkills: []

Join the ByteDance Seed Speech team to contribute cutting-edge research that advances foundation model capabilities for multimodal, interactive audio and speech experiences. You’ll develop and evaluate novel machine learning models and algorithms for areas such as speech synthesis, voice conversion, audio language modeling, and audio-video systems. Work with globally distributed researchers and engineering teams to translate research into product evolution impacting billions of users.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
12 hours ago

Research Scientist Graduate (Foundation Model-Speech-Interaction & Learning) - 2026 Start (PhD)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 12 hours agoStatus: Live
Reposted: similar role first listed 3 weeks ago

Job Summary

Join the ByteDance Seed Speech team to contribute cutting-edge research that advances foundation model capabilities for multimodal, interactive audio and speech experiences. You’ll develop and evaluate novel machine learning models and algorithms for areas such as speech synthesis, voice conversion, audio language modeling, and audio-video systems. Work with globally distributed researchers and engineering teams to translate research into product evolution impacting billions of users.
Location: San Jose
Workplace: Onsite
Employment Type: Full time
Job Function: Data Science & Machine Learning
Seniority: Graduate level

Key Responsibilities

  • •Contribute cutting-edge research that advances ByteDance product evolution (e.g., Douyin, CapCut) at global scale.
  • •Work on advanced audio processing and generation, including dialogue systems, audio-video models, speech synthesis, voice conversion, audio codec learning, and audio language modeling.
  • •Research, design, develop, and evaluate novel machine learning models and algorithms.
  • •Collaborate with globally based researchers and engineering teams to develop machine learning models and algorithms.

Key Requirements

  • •Completing or recently completed a PhD in Computer Science, Electrical Engineering, Electrical and Computer Engineering, Physics, Mathematics, or a related field.
  • •Good knowledge of theoretical and empirical research methods for addressing research problems.
  • •Solid knowledge and experience using at least one popular deep learning framework (e.g., PyTorch or TensorFlow) and familiarity with deep neural network architectures.
  • •Experience with both neural and non-neural, classical machine learning models and algorithms.
  • •Ability to establish authorization to work in the United States (no sponsorship or immigration-related benefits for this position).
Experience:Deep learningMachine learningSpeech synthesisAudio generation
Education:PhD / Doctorate
Tech Stack:PyTorchTensorFlowCC++PythonShell

Eligibility

Work Authorization:Authorization required. Sponsorship not provided.

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn