Visual Generation & Multimodal Evaluation Researcher Graduate (AML-Ark-US) - 2027 Start (PhD)

ByteDance
Seattle
Workplace: OnsiteFull timeFunction: Research & Scientific (R&D)Education: phdSkills: []

Work on multimodal evaluation for image and video models and AI agents within the Applied Machine Learning Ark team. Build evaluation systems and automated metrics, including reproducible human evaluation protocols, and develop video generation/debugging agents for multi-step creative workflows. Create large-scale image/video data pipelines and translate evaluation findings into model and product improvements for MaaS services used across markets.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
2 days ago

Visual Generation & Multimodal Evaluation Researcher Graduate (AML-Ark-US) - 2027 Start (PhD)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 14 hours agoStatus: Live

Job Summary

Work on multimodal evaluation for image and video models and AI agents within the Applied Machine Learning Ark team. Build evaluation systems and automated metrics, including reproducible human evaluation protocols, and develop video generation/debugging agents for multi-step creative workflows. Create large-scale image/video data pipelines and translate evaluation findings into model and product improvements for MaaS services used across markets.
Location: Seattle
Workplace: Onsite
Employment Type: Full time
Job Function: Research & Scientific (R&D)
Seniority: Graduate level

Key Responsibilities

  • •Build evaluation systems for image and video models/agents, covering generation quality, instruction following, multimodal understanding, and safety.
  • •Develop automated metrics and model-based evaluators, and design reproducible human evaluation protocols.
  • •Design and develop video generation/debugging agents that orchestrate multi-step creative workflows.
  • •Build large-scale image and video data pipelines and extract actionable insights from evaluation findings.
  • •Turn evaluation findings into model and product improvements.

Key Requirements

  • •Completing or recently completed a PhD in Computer Science, Artificial Intelligence, Machine Learning, Computer Vision, or a related field.
  • •Solid foundation in deep learning and computer vision, including generative modeling fundamentals.
  • •Practical experience in visual generation, multimodal LLMs, video understanding, or visual quality assessment.
  • •Strong Python skills and proficiency with PyTorch (or an equivalent framework), or a multimodal evaluation framework.
  • •Demonstrated research or engineering ability through publications, substantial projects, internships, or open-source work.
Education:PhD / Doctorate in Computer Science, Artificial Intelligence, Machine Learning, Computer Vision, or a related field
Tech Stack:PythonPyTorchDiffusion-based models

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn