Applied Scientist - LLM Training System as a Service - Global Frontier Tech Recruitment Program - 2027 Start (PhD)

ByteDance
San Jose
Workplace: OnsiteFull timeFunction: Research & Scientific (R&D)Education: phdSkills: ["Collaboration"]

Develop and optimize LLM training, inference, and reinforcement learning frameworks as part of a massively distributed ML training/inference systems team. Collaborate with model researchers to scale training and RL, and drive GPU and CUDA performance optimization for high-performance, reliable LLM training/inference engines. Contribute to a scalable architecture that supports emerging AI agent training workloads with more dynamic patterns.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
1 month ago

Applied Scientist - LLM Training System as a Service - Global Frontier Tech Recruitment Program - 2027 Start (PhD)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 30 days agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Develop and optimize LLM training, inference, and reinforcement learning frameworks as part of a massively distributed ML training/inference systems team. Collaborate with model researchers to scale training and RL, and drive GPU and CUDA performance optimization for high-performance, reliable LLM training/inference engines. Contribute to a scalable architecture that supports emerging AI agent training workloads with more dynamic patterns.
Location: San Jose
Workplace: Onsite
Employment Type: Full time
Job Function: Research & Scientific (R&D)
Seniority: Graduate level

Key Responsibilities

  • •Develop and optimize LLM training, inference, and reinforcement learning frameworks.
  • •Work closely with model researchers to scale LLM training and reinforcement learning to the next level.
  • •Optimize GPU and CUDA performance to build an industry-leading high-performance LLM training and inference and RL engine.
  • •Help design scalable architectures that separate logical control from compute execution for more adaptable training workflows.

Key Requirements

  • •Pursuing a PhD in computer science, automation, electronics engineering, or a related technical discipline.
  • •Proficient in algorithms and data structures, and familiar with Python.
  • •Understand basic deep learning principles, neural network architectures, and deep learning training frameworks such as PyTorch.
  • •Proficient in GPU high-performance computing optimization with CUDA, including parallel computing and memory access optimization.
  • •Familiar with distributed training/inference frameworks such as FSDP, DeepSpeed, JAX SPMD, Megatron-LM, and related tooling (e.g., TensorRT-LLM, vLLM, SGLang).
Experience:LLM trainingAIGCAI AgentsDistributed MLReinforcement learningGPU computing
Education:PhD / Doctorate in computer science, automation, electronics engineering, or related technical discipline
Skills:Collaboration
Tech Stack:PythonPyTorchDeep learningNeural networksReinforcement learningRLGPUCUDAFSDPDeepSpeedJAXMegatron-LMVerlTensorRT-LLMORCAVLLMSGLangLow-bit computing

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn