Visual Generation & Multimodal Evaluation Machine Learning Engineer Intern (AML-Ark-US) - 2027 Summer

ByteDance
San Jose
Workplace: OnsiteInternshipFunction: Data Science & Machine LearningSkills: []

Build evaluation systems for image/video generation and multimodal agents, measuring quality, instruction following, understanding, and safety. Develop automated metrics and reproducible human evaluation protocols, and design video generation/debugging agents for multi-step creative workflows. Work on large-scale image/video data pipelines, then translate evaluation findings into model and product improvements within the Applied Machine Learning Ark team.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
17 hours ago

Visual Generation & Multimodal Evaluation Machine Learning Engineer Intern (AML-Ark-US) - 2027 Summer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 17 hours agoStatus: Live

Job Summary

Build evaluation systems for image/video generation and multimodal agents, measuring quality, instruction following, understanding, and safety. Develop automated metrics and reproducible human evaluation protocols, and design video generation/debugging agents for multi-step creative workflows. Work on large-scale image/video data pipelines, then translate evaluation findings into model and product improvements within the Applied Machine Learning Ark team.
Location: San Jose
Workplace: Onsite
Employment Type: Internship
Job Function: Data Science & Machine Learning
Seniority: Intern level

Key Responsibilities

  • •Build evaluation systems for image and video models/agents, covering generation quality, instruction following, multimodal understanding, and safety.
  • •Develop automated metrics and model-based evaluators, and design reproducible human evaluation protocols.
  • •Design and develop video generation/debugging agents that orchestrate multi-step creative workflows.
  • •Build large-scale image and video data pipelines, and turn evaluation findings into model and product improvements.

Key Requirements

  • •Currently pursuing a Bachelor's or Master's degree in Computer Science, Artificial Intelligence, Machine Learning, Computer Vision, or a related field.
  • •Strong foundation in deep learning and computer vision, including generative modeling fundamentals.
  • •Practical experience in visual generation, multimodal LLMs, video understanding, or visual quality assessment.
  • •Strong Python skills and proficiency with PyTorch (or an equivalent framework), or multimodal evaluation frameworks.
  • •Demonstrated research or engineering ability via publications, substantial projects, internships, or open-source work.
Education:
Tech Stack:PythonPyTorch

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn