Visual Generation & Multimodal Evaluation Machine Learning Engineer Graduate (AML-Ark-US) - 2027 Start

ByteDance
San Jose
Workplace: OnsiteFull timeFunction: Data Science & Machine LearningSkills: ["Research capability","Engineering ability","Problem-solving"]

Build and evaluate multimodal LLM and video/image generation systems for MaaS solutions. You’ll develop automated metrics and model- and human-evaluation protocols, create video generation and debugging agents for multi-step creative workflows, and build large-scale image/video data pipelines. Work with large invocation log analytics to extract quality signals and usage patterns, feeding continuous improvements to models, systems, and product decisions.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
18 hours ago

Visual Generation & Multimodal Evaluation Machine Learning Engineer Graduate (AML-Ark-US) - 2027 Start

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 18 hours agoStatus: Live

Job Summary

Build and evaluate multimodal LLM and video/image generation systems for MaaS solutions. You’ll develop automated metrics and model- and human-evaluation protocols, create video generation and debugging agents for multi-step creative workflows, and build large-scale image/video data pipelines. Work with large invocation log analytics to extract quality signals and usage patterns, feeding continuous improvements to models, systems, and product decisions.
Location: San Jose
Workplace: Onsite
Employment Type: Full time
Job Function: Data Science & Machine Learning
Seniority: Graduate level

Key Responsibilities

  • •Build evaluation systems for image/video models and agents, covering generation quality, instruction following, multimodal understanding, and safety.
  • •Develop automated metrics and model-based evaluators, and design reproducible human evaluation protocols.
  • •Design and develop video generation/debugging agents that orchestrate multi-step creative workflows.
  • •Build large-scale image/video data pipelines and use evaluation findings to improve models and products.

Key Requirements

  • •Completing or recently completed a Bachelor's/ Master's degree in Computer Science, Artificial Intelligence, Machine Learning, Computer Vision, or a related field.
  • •Solid foundation in deep learning and computer vision, including generative modeling fundamentals.
  • •Practical experience in at least one of visual generation, multimodal LLMs, video understanding, or visual quality assessment.
  • •Strong Python skills and proficiency with PyTorch (or an equivalent framework) or a multimodal evaluation framework.
  • •Demonstrated research/engineering ability via publications, major projects, internships, or open-source work.
Experience:Large language modelsMultimodalComputer visionGenerative modelingVideo understandingAI agentsOpen source
Education:
Skills:Research capabilityEngineering abilityProblem-solving
Tech Stack:PythonPyTorchDiffusion modelsLLMsMultimodal evaluation

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn