Research Engineer - LLM Training Infrastructure - Seed Infra

ByteDance
San Jose
Workplace: OnsiteFull timeFunction: Education & TrainingSkills: ["Cross-functional collaboration","Performance optimization","Research and development","Systems thinking","Problem-solving"]

Build and optimize large-scale LLM training infrastructure for distributed reinforcement learning and AI foundation models. Design efficient distributed training strategies (parallelism, compute/communication optimization, throughput scaling) and improve system reliability via fast checkpointing, fault tolerance, and failure diagnosis. Analyze bottlenecks in exascale training systems, optimize networking and GPU memory management, and translate research ideas into production-grade, real-world infrastructure.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
1 month ago

Research Engineer - LLM Training Infrastructure - Seed Infra

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 30 days agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Build and optimize large-scale LLM training infrastructure for distributed reinforcement learning and AI foundation models. Design efficient distributed training strategies (parallelism, compute/communication optimization, throughput scaling) and improve system reliability via fast checkpointing, fault tolerance, and failure diagnosis. Analyze bottlenecks in exascale training systems, optimize networking and GPU memory management, and translate research ideas into production-grade, real-world infrastructure.
Location: San Jose
Workplace: Onsite
Employment Type: Full time
Job Function: Education & Training

Key Responsibilities

  • •Conduct research and development on large-scale LLM training infrastructure and efficiency.
  • •Design and optimize distributed training strategies, including parallelism schemes and compute/communication optimization for throughput scaling on GPU clusters.
  • •Improve reliability and resilience with techniques like fast checkpointing, fault tolerance, and failure diagnosis for long-running training workloads.
  • •Optimize network, scheduling, and GPU memory management across the training stack to drive cross-layer performance improvements.
  • •Analyze exascale training bottlenecks and translate data-driven optimization ideas into scalable, production deployment solutions.

Key Requirements

  • •Experience with large-scale distributed training for LLMs.
  • •Strong programming skills in Python and/or C++.
  • •Strong background in ML systems and training infrastructure development.
  • •Proficiency in parallelism strategies (DDP, FSDP, model/pipeline/expert parallelism).
  • •Solid understanding of training stack internals (PyTorch, CUDA, NCCL) and performance optimization (memory, communication, throughput).
Experience:LLMDistributed trainingML systemsTraining infrastructureHPC
Skills:Cross-functional collaborationPerformance optimizationResearch and developmentSystems thinkingProblem-solving
Tech Stack:PythonC++PyTorchCUDANCCLDDPFSDPModel parallelismPipeline parallelismExpert parallelism

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn