Member of Technical Staff - Pre-Training Infra

Reflection AI
San Francisco, New York, London
Workplace: OnsiteFull timeFunction: Education & TrainingSkills: ["Communication","Problem-solving","Collaboration","Teamwork","Diagnostic reasoning"]

Build and scale distributed training systems for frontier model pre-training, collaborating with researchers to run large-scale foundation model training across thousands of GPUs. Develop and optimize infrastructure and pipelines to improve throughput, stability, and experiment iteration in a production-ready environment.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Reflection AI
Reflection AI
4 months ago

Member of Technical Staff - Pre-Training Infra

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 14 hours agoStatus: Live

Job Summary

Build and scale distributed training systems for frontier model pre-training, collaborating with researchers to run large-scale foundation model training across thousands of GPUs. Develop and optimize infrastructure and pipelines to improve throughput, stability, and experiment iteration in a production-ready environment.
Location: San Francisco, New York, London
Workplace: Onsite
Employment Type: Full time
Job Function: Education & Training

Key Responsibilities

  • •Design, build, and scale distributed training systems and pipelines for large-model pre-training.
  • •Collaborate with ML researchers to translate experiments into scalable, production-ready training systems.
  • •Tune and optimize infrastructure to maximize training throughput, stability, and GPU utilization.
  • •Debug performance bottlenecks across distributed stacks, including model parallelism and GPU communication.
  • •Develop and maintain training pipelines for large datasets, checkpointing, and experiment iteration.

Pay and Benefits

Perks:Health InsuranceDentalVisionLife InsuranceDisability InsuranceRemote WorkLearning BudgetGym Membership

Key Requirements

  • •Experience building or operating distributed training systems for large machine learning models.
  • •Strong experience with distributed training frameworks such as Megatron, DeepSpeed, or similar systems.
  • •Familiarity with large-scale model parallelism strategies (data, tensor, pipeline, or expert parallelism).
  • •Experience optimizing training throughput and GPU utilization in large distributed environments.
  • •Knowledge of GPU communication libraries such as NCCL and performance tuning for distributed workloads.
Experience:Machine learningAIDistributed trainingFoundation models
Skills:CommunicationProblem-solvingCollaborationTeamworkDiagnostic reasoning
Tech Stack:MegatronDeepSpeedNCCLGPUCUDADistributed trainingModel parallelism

Company Brief

Reflection AI
Builds frontier autonomous AI systems focused on autonomous coding agents (product: Asimov) to create organizational superintelligence, founded by former DeepMind/Google researchers and hiring across SF, NYC, London.
Industry: AI & Machine Learning
Company Size: Small (11 to 50 employees)
Growth: Growth Stage Startup
Funding: Series A
Headquarters: San Francisco, United States
Founded: 2024
WebsiteLinkedInGlassdoor