Student Researcher (LLM Post Training – Agent & Reinforcement Learning) - 2026 Start (PhD)

ByteDance
San Jose
Workplace: OnsiteFull timeFunction: Research & Scientific (R&D)Education: phdSkills: ["Research","Experimentation","Programming","Model optimization","In-depth research"]

Join the Seed LLM Post Training team to research and improve post-training methods for unified multimodal large models. You’ll explore and optimize post-training technologies such as SFT, RM, and RL, and work on data construction, instruction tuning, preference alignment, and model optimization. Focus areas include advancing reasoning, coding, math, agent capabilities, and investigating future use cases through large-scale model research and systems optimization.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
2 hours ago

Student Researcher (LLM Post Training – Agent & Reinforcement Learning) - 2026 Start (PhD)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 2 hours agoStatus: Live

Job Summary

Join the Seed LLM Post Training team to research and improve post-training methods for unified multimodal large models. You’ll explore and optimize post-training technologies such as SFT, RM, and RL, and work on data construction, instruction tuning, preference alignment, and model optimization. Focus areas include advancing reasoning, coding, math, agent capabilities, and investigating future use cases through large-scale model research and systems optimization.
Location: San Jose
Workplace: Onsite
Employment Type: Full time
Job Function: Research & Scientific (R&D)
Seniority: Intern level

Key Responsibilities

  • •Explore large-scale models and optimize systems for post-training.
  • •Build datasets and perform instruction tuning, preference alignment, and model optimization.
  • •Improve model capabilities such as reasoning, coding, and math.
  • •Research and explore advanced agent and self-learning use cases in the post-training phase.
  • •Advance relevant technologies like SFT, RM, and RL for unified multimodal large models.

Key Requirements

  • •Currently pursuing a PhD in Computer Science, AI, or a related field.
  • •Research experience in reinforcement learning, sequential decision-making, or agent behavior.
  • •First-author publications in accredited ML/AI conferences (e.g., NeurIPS, ICLR, ICML).
  • •Solid programming and experimentation skills, including with RL or LLM frameworks.
  • •Experience with LLM agents, tool use, or prompt-based control (preferred).
Experience:Machine learningReinforcement learningLLM agentsAgent behavior
Education:PhD / Doctorate in Computer Science, AI, or a related field
Skills:ResearchExperimentationProgrammingModel optimizationIn-depth research
Tech Stack:LLM frameworks

Eligibility

Work Authorization:Authorization required. Sponsorship not provided.

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn