Research Scientist Graduate (Applied Machine Learning - ML System) - 2026 Start (PhD)

ByteDance
San Jose
Workplace: OnsiteFull timeFunction: Data Science & Machine LearningEducation: phdSkills: []

Build and optimize large-scale LLM training, inference, and RL frameworks as part of a distributed ML systems team. Collaborate with model researchers to scale training and RL to the next level, and drive GPU/CUDA performance optimization for a high-performance training/inference engine. Work on heterogeneous infrastructure integrating GPU/NPU/RDMA/Storage while contributing to stable, reliable system operations across a global team.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
1 month ago

Research Scientist Graduate (Applied Machine Learning - ML System) - 2026 Start (PhD)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 31 days agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Build and optimize large-scale LLM training, inference, and RL frameworks as part of a distributed ML systems team. Collaborate with model researchers to scale training and RL to the next level, and drive GPU/CUDA performance optimization for a high-performance training/inference engine. Work on heterogeneous infrastructure integrating GPU/NPU/RDMA/Storage while contributing to stable, reliable system operations across a global team.
Location: San Jose
Workplace: Onsite
Employment Type: Full time
Job Function: Data Science & Machine Learning
Seniority: Graduate level

Key Responsibilities

  • •Develop and optimize LLM training, inference, and RL frameworks.
  • •Partner with model researchers to scale LLM training and RL to the next level.
  • •Optimize GPU and CUDA performance to build a high-performance LLM training, inference, and RL engine.
  • •Build and maintain massively distributed ML training and inference systems and services.
  • •Integrate and operate heterogeneous infrastructure combining GPU/NPU/RDMA/Storage for stable, reliable scalability.

Key Requirements

  • •Be completing or have recently completed a PhD in Software Development, Computer Science, Computer Engineering, or a related technical discipline.
  • •Proficient in algorithms and data structures and familiar with Python.
  • •Understand deep learning fundamentals, neural network architecture, and deep learning training frameworks such as PyTorch.
  • •Have GPU high-performance computing optimization experience with CUDA, including computer architecture and parallel/memory optimization (including low-bit computing).
  • •Be familiar with distributed training/acceleration frameworks such as FSDP, DeepSpeed, JAX SPMD, Megatron-LM, Verl, TensorRT-LLM, ORCA, vLLM, or SGLang, and know LLM model optimization concepts.
Education:PhD / Doctorate in Software Development, Computer Science, Computer Engineering, or related technical discipline
Tech Stack:PythonPyTorchGPUNPUCUDARDMAStorageFSDPDeepSpeedJAXMegatron-LMVerlTensorRT-LLMORCAVLLMSGLang

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn