Student Researcher (Speech Foundation Model - Seed) – 2027 Start (PhD)

ByteDance
San Jose
Workplace: OnsiteInternshipFunction: Research & Scientific (R&D)Education: phdSkills: ["Problem-solving","Independent research","Collaboration"]

Join the Seed Speech team to conduct research on speech foundation models and related systems. You’ll explore techniques to improve speech generation, speech understanding, and multimodal modeling across speech, language, and vision, and design/prototype algorithms, models, and system components. Collaborate with the team to advance research directions while building on a strong programming background and an active PhD in a relevant technical field.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
1 month ago

Student Researcher (Speech Foundation Model - Seed) – 2027 Start (PhD)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 30 days agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Join the Seed Speech team to conduct research on speech foundation models and related systems. You’ll explore techniques to improve speech generation, speech understanding, and multimodal modeling across speech, language, and vision, and design/prototype algorithms, models, and system components. Collaborate with the team to advance research directions while building on a strong programming background and an active PhD in a relevant technical field.
Location: San Jose
Workplace: Onsite
Employment Type: Internship
Job Function: Research & Scientific (R&D)
Seniority: Intern level

Key Responsibilities

  • •Conduct research on speech foundation models and related systems.
  • •Explore methods to advance model capabilities for speech generation, speech understanding, and multimodal modeling involving speech, language, or vision.
  • •Design and prototype algorithms, models, or system components.
  • •Collaborate with the team to advance research directions.

Key Requirements

  • •Currently pursuing a PhD in computer science, electrical engineering, mathematics, or a related field.
  • •Strong programming skills and solid foundation in algorithms and data structures; proficient in Python or C/C++.
  • •Demonstrated research track record with publications in conferences related to speech, audio, multimodal learning, or machine learning.
  • •Experience related to speech generation, speech understanding, multimodal modeling, or representation learning.
  • •Strong problem-solving ability with experience conducting independent research and collaborating effectively in a research environment.
Experience:ResearchSpeechAudioMultimodal learningMachine learning
Education:PhD / Doctorate
Skills:Problem-solvingIndependent researchCollaboration
Tech Stack:PythonC/C++Speech generationSpeech understandingMultimodal modelingNatural language understandingMultimodal deep learningAlgorithmsData structures

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn