Research Scientist - LLM Training System as a Service - Global Frontier Tech Recruitment Program - 2027 Start (PhD)

ByteDance
San Jose
Workplace: OnsiteFull timeFunction: Data Science & Machine LearningEducation: phdSkills: []

Develop and optimize distributed LLM training, inference, and reinforcement learning frameworks in collaboration with model researchers. Work on GPU and CUDA performance optimization to deliver high-performance, reliable, scalable training systems. Contribute to next-generation agent training workflows by designing architectures that separate logical control from compute execution. This PhD recruitment program focuses on accelerating LLM model optimization and supporting frontier AI workloads worldwide.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
1 month ago

Research Scientist - LLM Training System as a Service - Global Frontier Tech Recruitment Program - 2027 Start (PhD)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 30 days agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Develop and optimize distributed LLM training, inference, and reinforcement learning frameworks in collaboration with model researchers. Work on GPU and CUDA performance optimization to deliver high-performance, reliable, scalable training systems. Contribute to next-generation agent training workflows by designing architectures that separate logical control from compute execution. This PhD recruitment program focuses on accelerating LLM model optimization and supporting frontier AI workloads worldwide.
Location: San Jose
Workplace: Onsite
Employment Type: Full time
Job Function: Data Science & Machine Learning
Seniority: Graduate level

Key Responsibilities

  • •Develop and optimize LLM training and inference and reinforcement learning frameworks.
  • •Collaborate with model researchers to scale LLM training and reinforcement learning.
  • •Optimize GPU and CUDA performance to build a high-performance LLM training, inference, and RL engine.
  • •Help design scalable training workflows by separating logical control from compute execution.

Key Requirements

  • •Currently pursuing a PhD in computer science, automation, electronics engineering, or a related technical discipline.
  • •Proficient in algorithms and data structures, and familiar with Python.
  • •Understand basic principles of deep learning algorithms, neural network architectures, and deep learning training frameworks such as PyTorch.
Experience:Deep learningLLM trainingDistributed computing
Education:PhD / Doctorate in computer science, automation, electronics engineering or a related technical discipline
Tech Stack:PythonPyTorchGPUCUDADeep learningNeural networksFSDPDeepspeedJAXMegatron-LMVerlTensorRT-LLMORCAVLLMSGLang

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn