Member of Technical Staff - Distributed Training Engineer

Liquid AI
San Francisco, United States
Workplace: HybridFull timeFunction: Education & TrainingSkills: ["Distributed systems","Debugging","Profiling","Performance tuning","Problem-solving"]

We are seeking a Distributed Training Engineer to design and optimize high-ownership training systems for large-scale models. You’ll build core distributed infrastructure, optimize GPU cluster training, improve data loading and checkpointing, and create tools to monitor performance and stability, working in a small team with fast feedback loops across evolving model architectures.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Liquid AI
Liquid AI
1 year ago

Member of Technical Staff - Distributed Training Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 7 hours agoStatus: Live

Job Summary

We are seeking a Distributed Training Engineer to design and optimize high-ownership training systems for large-scale models. You’ll build core distributed infrastructure, optimize GPU cluster training, improve data loading and checkpointing, and create tools to monitor performance and stability, working in a small team with fast feedback loops across evolving model architectures.
Location: San Francisco, United States
Workplace: Hybrid
Employment Type: Full time
Job Function: Education & Training

Key Responsibilities

  • •Design and build core systems that make large training runs fast and reliable
  • •Build scalable distributed training infrastructure for GPU clusters
  • •Implement and tune parallelism/sharding strategies for evolving architectures
  • •Optimize distributed efficiency (topology-aware collectives, comm/compute overlap, straggler mitigation)
  • •Build data loading systems that eliminate I/O bottlenecks for multimodal datasets

Pay and Benefits

Perks:Health InsuranceDentalVision401kPaid LeaveTime OffEquity

Key Requirements

  • •Hands-on experience building distributed training infrastructure (PyTorch Distributed DDP/FSDP, DeepSpeed ZeRO, Megatron-LM TP/PP)
  • •Experience diagnosing performance bottlenecks and failure modes (profiling, NCCL/collectives issues, hangs, OOMs, stragglers)
  • •Understanding of hardware accelerators and networking topologies
  • •Experience optimizing data pipelines for ML workloads
  • •Nice-to-have: MoE (Mixture of Experts) training experience
Experience:AIMachine learningDistributed systems
Skills:Distributed systemsDebuggingProfilingPerformance tuningProblem-solving
Languages:English
Tech Stack:PyTorchDDPFSDPDeepSpeedMegatron-LMTP/PPNCCLGPUsCUDA

Company Brief

Liquid AI
Builds AI infrastructure and tooling to enable real-time, distributed machine learning and orchestration across edge and cloud environments, simplifying deployment and management of intelligent applications.
Industry: AI & Machine Learning
Website