Senior Research Scientist, Intelligent Editing (Multimodality)

ByteDance
Seattle
Workplace: OnsiteFull timeFunction: Data Science & Machine LearningExperience: 3+ yearsSkills: ["Algorithms","Coding","Research","Technology transfer"]

Join the Intelligent Creation Team, building AI and multimodal technologies for content understanding and intelligent editing. You’ll conduct cutting-edge research in computer vision and machine learning—covering video and audio understanding, vision-and-language, language model downstream tasks, and large-scale training with RLHF. You’ll then transfer promising methods into ByteDance products and help explore new AI-first product experiences.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
1 month ago

Senior Research Scientist, Intelligent Editing (Multimodality)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 30 days agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Join the Intelligent Creation Team, building AI and multimodal technologies for content understanding and intelligent editing. You’ll conduct cutting-edge research in computer vision and machine learning—covering video and audio understanding, vision-and-language, language model downstream tasks, and large-scale training with RLHF. You’ll then transfer promising methods into ByteDance products and help explore new AI-first product experiences.
Location: Seattle
Workplace: Onsite
Employment Type: Full time
Job Function: Data Science & Machine Learning
Seniority: Mid level

Key Responsibilities

  • •Conduct research and development in computer vision and machine learning, including multimodal understanding, vision and language, and large-scale training.
  • •Transfer advanced technologies to ByteDance products.
  • •Explore new AI-first products at its core.

Key Requirements

  • •3+ years of research and practical experience in computer vision.
  • •Experience with multimodal understanding (e.g., video highlight detection and slicing; audio/music understanding).
  • •Experience in vision-and-language tasks such as captioning, retrieval, and VQA.
  • •Experience applying language models to downstream tasks, especially for intelligent editing.
  • •Experience with large-scale training and RLHF; strong coding skills in C/C++ and Python; publications preferred.
Experience:3+ yearsComputer visionMachine learningMultimodalLanguage modelsDeep learningLarge-scale trainingRLHF
Skills:AlgorithmsCodingResearchTechnology transfer
Tech Stack:C/C++PythonDeep learningComputer visionLanguage modelsRLHF

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn