Visual Generation & Multimodal Evaluation Researcher Graduate (AML-Ark-US) - 2027 Start (PhD)

ByteDance
San Jose
Workplace: OnsiteFull timeFunction: Research & Scientific (R&D)Education: phdSkills: ["Research ability","Engineering ability","Problem-solving","Data-driven decision making"]

Build evaluation systems for multimodal (text + image/video) models and AI agents, focusing on generation quality, instruction following, multimodal understanding, and safety. Develop automated metrics and reproducible human evaluation protocols, and design/debug video generation agents for multi-step creative workflows. Create large-scale image/video data pipelines and turn evaluation results into model and product improvements through continuous, data-driven feedback loops.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
19 hours ago

Visual Generation & Multimodal Evaluation Researcher Graduate (AML-Ark-US) - 2027 Start (PhD)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 19 hours agoStatus: Live

Job Summary

Build evaluation systems for multimodal (text + image/video) models and AI agents, focusing on generation quality, instruction following, multimodal understanding, and safety. Develop automated metrics and reproducible human evaluation protocols, and design/debug video generation agents for multi-step creative workflows. Create large-scale image/video data pipelines and turn evaluation results into model and product improvements through continuous, data-driven feedback loops.
Location: San Jose
Workplace: Onsite
Employment Type: Full time
Job Function: Research & Scientific (R&D)
Seniority: Graduate level

Key Responsibilities

  • •Build evaluation systems for image/video models and agents, covering generation quality, instruction following, multimodal understanding, and safety.
  • •Develop automated metrics and model-based evaluators, and design reproducible human evaluation protocols.
  • •Design and develop video generation/debugging agents that orchestrate multi-step creative workflows.
  • •Build large-scale image and video data pipelines.
  • •Turn evaluation findings into model and product improvements.

Key Requirements

  • •Completing or recently completed a PhD in Computer Science, Artificial Intelligence, Machine Learning, Computer Vision, or a related field.
  • •Strong foundation in deep learning and computer vision, including generative modeling fundamentals.
  • •Practical experience in visual generation, multimodal LLMs, video understanding, or visual quality assessment.
  • •Strong Python skills with proficiency in PyTorch (or an equivalent framework), or experience with multimodal evaluation frameworks.
  • •Demonstrated research or engineering ability via publications, substantial projects, internships, or open-source work.
Experience:Deep learningComputer visionMultimodalLLMsOpen sourceVideo understandingVisual generation
Education:PhD / Doctorate in Computer Science, Artificial Intelligence, Machine Learning, Computer Vision, or a related field
Skills:Research abilityEngineering abilityProblem-solvingData-driven decision making
Tech Stack:PythonPyTorchLLMMultimodal LLMsComputer visionDeep learningDiffusion models

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn