Senior Research Scientist (Multimodal Large Language Model) - PICO

ByteDance
San Jose
Workplace: OnsiteFull timeFunction: Data Science & Machine LearningEducation: high_schoolSkills: ["Problem-solving","Independent research","Collaboration","Communication"]

Lead R&D for multimodal large language models tailored to MR scenarios, integrating vision, point clouds, and text with model optimization, cross-modal alignment, data construction, and accelerated training/inference. Develop and implement tool-use capabilities for single- and multi-turn MR interactions, tackling long-horizon, memory, tool selection, and error correction challenges. Collaborate across software, product design, and hardware teams to deploy research into PICO MR products.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
1 month ago

Senior Research Scientist (Multimodal Large Language Model) - PICO

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 31 days agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Lead R&D for multimodal large language models tailored to MR scenarios, integrating vision, point clouds, and text with model optimization, cross-modal alignment, data construction, and accelerated training/inference. Develop and implement tool-use capabilities for single- and multi-turn MR interactions, tackling long-horizon, memory, tool selection, and error correction challenges. Collaborate across software, product design, and hardware teams to deploy research into PICO MR products.
Location: San Jose
Workplace: Onsite
Employment Type: Full time
Job Function: Data Science & Machine Learning

Key Responsibilities

  • •Lead R&D of multimodal large language models for MR, including architecture optimization, cross-modal alignment, data construction, evaluation system enhancement, and training/inference acceleration.
  • •Research and implement MLLM tool-use capabilities in MR, enabling tool calls for single-turn and multi-turn conversations and task completion via interaction.
  • •Solve long-horizon, multi-turn tool-augmented MR challenges including context memory management, tool selection strategy, and error correction mechanisms.
  • •Stay current with cutting-edge MLLM, multimodal intelligence, and tool-use research, and drive application and deployment in PICO MR products.
  • •Collaborate with cross-functional teams (software engineering, product design, hardware development) to translate research into user-experience features.

Key Requirements

  • •Master's or Ph.D. in Computer Science, Electrical Engineering, Machine Learning, Artificial Intelligence, or a related quantitative field.
  • •Expertise in multimodal large model pre-training, post-training, fine-tuning, or cross-modal fusion, including model optimization and training workflow design.
  • •Research experience in LLM tool use, reinforcement learning, LLM agents, or interactive learning with understanding of single- and multi-turn interaction.
  • •Proficiency in core 2D/3D computer vision tasks including detection, segmentation, depth estimation, image matching, and 3D scene perception.
  • •Skilled in Python and C++, and experience building large-scale models with mainstream deep learning frameworks (PyTorch/TensorFlow).
Education:High School
Skills:Problem-solvingIndependent researchCollaborationCommunication
Tech Stack:PythonC++PyTorchTensorFlow

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn