Research Engineer – Multimodal Training Infrastructure (Seed Infra)

ByteDance
San Jose
Workplace: OnsiteFull timeFunction: Education & TrainingSkills: []

Build large-scale multimodal training infrastructure for AI foundation models, focusing on efficient distributed training for multimodal LLMs and image/video generation. Design and optimize parallelism strategies, throughput scaling, and cross-layer performance improvements across GPUs and the training stack. Improve reliability with fast checkpointing and fault tolerance, diagnose failures, and analyze exascale bottlenecks to propose data-driven optimizations for production deployment.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
1 month ago

Research Engineer – Multimodal Training Infrastructure (Seed Infra)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 31 days agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Build large-scale multimodal training infrastructure for AI foundation models, focusing on efficient distributed training for multimodal LLMs and image/video generation. Design and optimize parallelism strategies, throughput scaling, and cross-layer performance improvements across GPUs and the training stack. Improve reliability with fast checkpointing and fault tolerance, diagnose failures, and analyze exascale bottlenecks to propose data-driven optimizations for production deployment.
Location: San Jose
Workplace: Onsite
Employment Type: Full time
Job Function: Education & Training

Key Responsibilities

  • •Conduct R&D on large-scale infrastructure for efficient training of foundation models, multimodal LLMs, and image/video generation models.
  • •Design and optimize distributed training strategies, including parallelism schemes and computation/communication optimization for throughput scaling on large GPU clusters.
  • •Investigate reliability and resilience techniques such as fast checkpointing, fault tolerance, and failure diagnosis for long-running training workloads.
  • •Research and optimize network, scheduling, and GPU memory management across the training stack to drive cross-layer performance improvements.
  • •Analyze performance bottlenecks in exascale training systems and translate research ideas into scalable production infrastructure solutions.

Key Requirements

  • •Deep expertise in large-scale distributed training of LLMs and multimodal models.
  • •Strong systems research background with ability to design, build, and optimize large-scale ML systems.
  • •Experience implementing parallelism strategies (data, model, pipeline, expert) and performance optimization on large GPU clusters.
  • •Strong programming skills with hands-on experience building production-grade ML systems or infrastructure.
  • •Solid understanding of algorithm–system co-design and cross-layer optimization for training efficiency, scalability, and reliability.
Experience:AIMultimodalDeep learningDistributed trainingReinforcement learning
Tech Stack:Distributed trainingMultimodal LLMsImage/video generationReinforcement learningHigh-performance inferenceGPU clustersParallelismCheckpointingFault toleranceGPU memory managementNetwork optimizationSchedulingHeterogeneous hardware compilationExascale systems

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn