Research Scientist - Foundation Model, Speech Understanding

ByteDance
San Jose
Workplace: OnsiteFull timeFunction: Data Science & Machine LearningSkills: ["Collaboration","Cross-functional teamwork","Ability to work in fast-paced environments","Coding/programming proficiency"]

Develop speech/audio foundation models and advance research in multimodal speech technologies. Collaborate with cross-functional teams to define key research directions, then work with product engineering to translate research results into practical applications for ByteDance and other platforms. Contribute to team-driven projects addressing complex challenges, leveraging expertise in machine learning, deep learning, and large-scale model training to improve research effectiveness.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
1 month ago

Research Scientist - Foundation Model, Speech Understanding

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 30 days agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Develop speech/audio foundation models and advance research in multimodal speech technologies. Collaborate with cross-functional teams to define key research directions, then work with product engineering to translate research results into practical applications for ByteDance and other platforms. Contribute to team-driven projects addressing complex challenges, leveraging expertise in machine learning, deep learning, and large-scale model training to improve research effectiveness.
Location: San Jose
Workplace: Onsite
Employment Type: Full time
Job Function: Data Science & Machine Learning
Seniority: Mid level

Key Responsibilities

  • •Conduct research and development for speech/audio foundation models.
  • •Collaborate with cross-functional teams to identify key research areas and develop innovative speech/audio models.
  • •Partner with product development teams to integrate research findings into practical applications for ByteDance and other platforms.
  • •Contribute to team-driven projects to solve complex challenges and improve overall research effectiveness.

Key Requirements

  • •Master’s or PhD in computer science, mathematics, engineering, or a related field.
  • •3+ years of experience in machine learning and deep learning, including areas such as automatic speech recognition/translation, speech/audio self-supervised learning, foundation models, or multimodal foundation models.
  • •Experience with large language model pre-training and fine-tuning.
  • •Familiarity with distributed computing and large-scale model training.
  • •Proficiency with deep learning frameworks such as Tensorflow and Pytorch, plus strong coding skills in C/C++ and Python.
Experience:Machine learningDeep learningSpeech recognitionSpeech translationMultimodalFoundation modelsLarge language modelsSelf-supervised learningDistributed training
Skills:CollaborationCross-functional teamworkAbility to work in fast-paced environmentsCoding/programming proficiency
Tech Stack:TensorflowPytorchC/C++PythonDistributed computingLarge-scale model training

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn