Research Engineer Graduate (AI Training Systems Reliability & Performance - Seed Infra) - 2026 Start (PhD)

ByteDance
Seattle
Workplace: OnsiteFull timeFunction: Education & TrainingEducation: phdSkills: ["Problem-solving","Collaboration","Analytical skills"]

Build reliability and performance improvements for large-scale AI training systems spanning pre-training, fine-tuning, evaluation, and inference. Develop observability, profiling, and debugging tools for distributed ML workloads, and identify bottlenecks across GPU, networking, and storage. Work on multi-GPU and multi-node distributed training frameworks, collaborate with model and infrastructure teams, and support incident analysis to improve operational stability.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
12 hours ago

Research Engineer Graduate (AI Training Systems Reliability & Performance - Seed Infra) - 2026 Start (PhD)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 12 hours agoStatus: Live
Reposted: similar role first listed 3 weeks ago

Job Summary

Build reliability and performance improvements for large-scale AI training systems spanning pre-training, fine-tuning, evaluation, and inference. Develop observability, profiling, and debugging tools for distributed ML workloads, and identify bottlenecks across GPU, networking, and storage. Work on multi-GPU and multi-node distributed training frameworks, collaborate with model and infrastructure teams, and support incident analysis to improve operational stability.
Location: Seattle
Workplace: Onsite
Employment Type: Full time
Job Function: Education & Training
Seniority: Graduate level

Key Responsibilities

  • •Improve the reliability and performance of large-scale training systems across pre-training, fine-tuning, evaluation, and inference.
  • •Build observability, profiling, and debugging tools for distributed ML workloads.
  • •Identify and optimize performance bottlenecks across GPU, networking, and storage layers.
  • •Contribute to distributed training frameworks in multi-GPU and multi-node environments.
  • •Support incident analysis and operational stability.

Key Requirements

  • •Completing or recently completed a PhD in Computer Science, Electrical Engineering, Electrical and Computer Engineering, Physics, Mathematics, or a related discipline.
  • •Strong programming skills in C++ and Python.
  • •Solid understanding of PyTorch training workflows and distributed runtime behavior.
  • •Familiarity with CUDA execution and NCCL communication, plus GPU systems fundamentals.
  • •Ability to analyze and optimize performance in complex ML training systems.
Education:PhD / Doctorate
Skills:Problem-solvingCollaborationAnalytical skills
Tech Stack:C++PythonPyTorchCUDANCCLGPUTorch.profilerNsightFSDPMegatron-LM

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn