Research Scientist, Intelligent Editing (Multimodality)

ByteDance
Seattle
Workplace: OnsiteFull timeFunction: Data Science & Machine LearningExperience: 1+ yearsSkills: ["Collaboration","Independent work"]

Work on ByteDance’s Intelligent Creation Team to conduct cutting-edge R&D in computer vision and machine learning, focused on multimodal understanding, vision-and-language modeling, and large-scale training. Apply language models to intelligent editing workflows, transfer advances into ByteDance products, and explore new AI-native product ideas. Collaborate on algorithms and implementation, including RLHF-driven training where relevant.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
1 month ago

Research Scientist, Intelligent Editing (Multimodality)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 30 days agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Work on ByteDance’s Intelligent Creation Team to conduct cutting-edge R&D in computer vision and machine learning, focused on multimodal understanding, vision-and-language modeling, and large-scale training. Apply language models to intelligent editing workflows, transfer advances into ByteDance products, and explore new AI-native product ideas. Collaborate on algorithms and implementation, including RLHF-driven training where relevant.
Location: Seattle
Workplace: Onsite
Employment Type: Full time
Job Function: Data Science & Machine Learning
Seniority: Entry level

Key Responsibilities

  • •Conduct cutting-edge research and development in computer vision and machine learning, especially multimodal understanding, vision and language, and large-scale training.
  • •Transfer advanced technologies to ByteDance products.
  • •Explore new AI-core products with artificial intelligence technology at the core.

Key Requirements

  • •At least 1 year of research and practical experience in computer vision, including multimodal understanding (e.g., video highlight detection/slicing, audio/music understanding).
  • •Experience in vision-and-language tasks such as image/video captioning, retrieval, and VQA.
  • •Experience with language models and applying them to downstream tasks, especially intelligent editing.
  • •Experience with large-scale training and RLHF.
  • •Strong algorithms and programming capability, including coding skills in C/C++ and Python.
Experience:1+ yearsComputer visionMultimodalVision and languageLanguage modelsLarge-scale trainingRLHF
Skills:CollaborationIndependent work
Tech Stack:Computer visionMachine learningDeep learningMultimodal understandingVision and languageImage/video captioningRetrievalVQALanguage modelsIntelligent editingLarge-scale trainingRLHFC/C++Python

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn