Research Engineer Graduate (AI Training Systems & RL Infrastructure - Seed Infra) - 2026 Start (PhD)

ByteDance
San Jose
Full timeFunction: Education & TrainingEducation: phdSkills: []

Join the Seed Infrastructures team to build large-scale AI training and post-training infrastructure for foundation models, multimodal LLMs, and image/video generation. You’ll design and optimize distributed training strategies, prototype end-to-end reinforcement learning (RL) training systems, and develop fault-tolerant infrastructure under dynamic, heterogeneous workloads. Work with researchers on system–algorithm co-design and create observability tooling to diagnose and improve throughput, efficiency, and stability.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
2 days ago

Research Engineer Graduate (AI Training Systems & RL Infrastructure - Seed Infra) - 2026 Start (PhD)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 14 hours agoStatus: Live
Reposted: similar role first listed 3 weeks ago

Job Summary

Join the Seed Infrastructures team to build large-scale AI training and post-training infrastructure for foundation models, multimodal LLMs, and image/video generation. You’ll design and optimize distributed training strategies, prototype end-to-end reinforcement learning (RL) training systems, and develop fault-tolerant infrastructure under dynamic, heterogeneous workloads. Work with researchers on system–algorithm co-design and create observability tooling to diagnose and improve throughput, efficiency, and stability.
Location: San Jose
Employment Type: Full time
Job Function: Education & Training
Seniority: Graduate level

Key Responsibilities

  • •Conduct research and development on large-scale AI infrastructure for efficient training and post-training of foundation models and multimodal generation models.
  • •Design and optimize distributed training strategies including data/model/tensor/pipeline/expert parallelism, computation–communication overlap, and GPU cluster scaling.
  • •Prototype and improve end-to-end reinforcement learning training systems covering rollout generation, policy optimization, evaluation, and iterative deployment workflows.
  • •Build scalable, fault-tolerant infrastructure that operates reliably under dynamic workloads and heterogeneous compute environments.
  • •Collaborate on system–algorithm co-design and translate research prototypes into scalable, production-ready infrastructure; develop observability tooling for reliability.

Key Requirements

  • •Complete or recently completed a PhD in Computer Science, Electrical Engineering, Electrical and Computer Engineering, Physics, Mathematics, or related fields.
  • •Strong background in distributed systems and large-scale machine learning/deep learning infrastructure.
  • •Experience training or optimizing large-scale models (e.g., LLMs, multimodal models, RL systems).
  • •Understanding of parallelism strategies (data, model/tensor, pipeline, expert parallelism) and distributed training concepts.
  • •Experience programming with Python and/or C++ and modern ML frameworks such as PyTorch with distributed training tools.
Experience:Large-scale machine learningDeep learning infrastructureDistributed systemsReinforcement learningLLMsMultimodal models
Education:PhD / Doctorate
Tech Stack:PythonC++PyTorch

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn