Agent Evaluation & Evolution Researcher Graduate (AML-Ark-US) - 2027 Start (PhD)

ByteDance
San Jose
Full timeFunction: Research & Scientific (R&D)Education: phdSkills: ["Research ability","Engineering ability","Analytical skills","Problem-solving","Communication"]

Build and improve evaluation systems for LLM-based agents, including benchmark and automated judging pipelines that combine rule checks, model-based judging, and human review. Analyze agent execution traces and user feedback to identify failure patterns and translate them into production-ready system improvements. Work across research, platform, and product teams on closed-loop methods for capability growth, supporting MaaS platforms and large-scale log analytics.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
17 hours ago

Agent Evaluation & Evolution Researcher Graduate (AML-Ark-US) - 2027 Start (PhD)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 17 hours agoStatus: Live

Job Summary

Build and improve evaluation systems for LLM-based agents, including benchmark and automated judging pipelines that combine rule checks, model-based judging, and human review. Analyze agent execution traces and user feedback to identify failure patterns and translate them into production-ready system improvements. Work across research, platform, and product teams on closed-loop methods for capability growth, supporting MaaS platforms and large-scale log analytics.
Location: San Jose
Employment Type: Full time
Job Function: Research & Scientific (R&D)
Seniority: Graduate level

Key Responsibilities

  • •Design evaluation systems for LLM-based agents covering task success, tool use, reasoning quality, and reliability.
  • •Build benchmarks and automated judging pipelines combining rule-based checks, model-based judging, and human review.
  • •Analyze agent execution traces and user feedback to identify failure patterns and drive concrete system improvements.
  • •Support closed-loop learning from experience to capability, partnering with research, platform, and product teams to bring methods into production.

Key Requirements

  • •Completing or recently completed a PhD in Computer Science, Artificial Intelligence, Machine Learning, Data Science, or a related field.
  • •Solid foundation in machine learning and deep learning.
  • •Hands-on experience with LLM-based systems (e.g., agents, tool calling, retrieval, multi-agent systems) through research, internships, or projects.
  • •Strong Python skills and experience with a mainstream ML or agent evaluation framework.
  • •Demonstrated research or engineering ability via publications, substantial projects, internships, or open-source work.
Experience:Machine learningDeep learningLarge language modelsAI agentsNLPRetrievalMulti-agent systemsLog analysis
Education:PhD / Doctorate in Computer Science, Artificial Intelligence, Machine Learning, or Data Science (or related field)
Skills:Research abilityEngineering abilityAnalytical skillsProblem-solvingCommunication
Tech Stack:PythonMachine LearningDeep LearningLarge Language Models (LLMs)Tool callingRetrievalMulti-agent systemsLog analytics

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn