Research Engineer – Reinforcement Learning (RL) Systems & Infrastructure (Seed Infra)

ByteDance
San Jose
Workplace: OnsiteFull timeFunction: Education & TrainingSkills: ["Cross-team collaboration","System observability","Debugging","Performance optimization"]

Build end-to-end reinforcement learning systems for large-scale models, including rollout, training, evaluation, and deployment pipelines. Develop scalable, fault-tolerant RL infrastructure that handles dynamic workloads and heterogeneous compute. Optimize distributed training across GPU clusters to improve throughput, utilization, and stability. Collaborate with researchers on system–algorithm co-design and create tooling for monitoring, debugging, and observability of large-scale RL training.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
1 month ago

Research Engineer – Reinforcement Learning (RL) Systems & Infrastructure (Seed Infra)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 30 days agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Build end-to-end reinforcement learning systems for large-scale models, including rollout, training, evaluation, and deployment pipelines. Develop scalable, fault-tolerant RL infrastructure that handles dynamic workloads and heterogeneous compute. Optimize distributed training across GPU clusters to improve throughput, utilization, and stability. Collaborate with researchers on system–algorithm co-design and create tooling for monitoring, debugging, and observability of large-scale RL training.
Location: San Jose
Workplace: Onsite
Employment Type: Full time
Job Function: Education & Training
Seniority: Mid level

Key Responsibilities

  • •Design and build end-to-end RL systems for large-scale models, including rollout, training, evaluation, and deployment pipelines.
  • •Develop scalable, fault-tolerant RL infrastructure for dynamic workloads and heterogeneous compute environments.
  • •Optimize distributed training performance across GPU clusters to improve throughput, resource utilization, and system stability.
  • •Collaborate with cross-team researchers on system–algorithm co-design for production-grade implementations.
  • •Build tooling, monitoring, and debugging frameworks for reliability and observability of large-scale RL training systems.

Key Requirements

  • •Strong background in distributed systems, large-scale ML systems, or deep learning infrastructure.
  • •Experience building or optimizing large-scale training systems (RL, LLM, multimodal models).
  • •Solid engineering skills in Python/C++ and familiarity with modern ML stacks (e.g., PyTorch, distributed training frameworks).
  • •Experience with GPU optimization, parallelism strategies, and system-level performance tuning.
  • •Understanding of reinforcement learning workflows (rollout, policy update, evaluation loops).
Experience:Large-scale MLDeep learningReinforcement learningLLM trainingDistributed systemsGPU optimization
Skills:Cross-team collaborationSystem observabilityDebuggingPerformance optimization
Tech Stack:PythonC++PyTorchGPUDistributed training frameworks

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn