Member of Technical Staff - Mid-Training Infra

Reflection AI
San Francisco, New York, London
Workplace: OnsiteFull timeFunction: Education & TrainingSkills: ["Problem-solving","Collaboration","Communication"]

Design, build, and operate large-scale GPU infrastructure for high-throughput model inference and mid-training workloads; develop systems for synthetic data generation and reinforcement learning pipelines; optimize throughput, latency, and GPU utilization across thousands of GPUs; collaborate with research teams on distributed RL workloads and large-scale evaluation infrastructure; diagnose performance bottlenecks in inference runtimes, GPU kernels, networking, and distributed compute systems.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Reflection AI
Reflection AI
4 months ago

Member of Technical Staff - Mid-Training Infra

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 14 hours agoStatus: Live

Job Summary

Design, build, and operate large-scale GPU infrastructure for high-throughput model inference and mid-training workloads; develop systems for synthetic data generation and reinforcement learning pipelines; optimize throughput, latency, and GPU utilization across thousands of GPUs; collaborate with research teams on distributed RL workloads and large-scale evaluation infrastructure; diagnose performance bottlenecks in inference runtimes, GPU kernels, networking, and distributed compute systems.
Location: San Francisco, New York, London
Workplace: Onsite
Employment Type: Full time
Job Function: Education & Training
Seniority: Mid level

Key Responsibilities

  • •Design, build, and operate large-scale GPU infrastructure for high-throughput model inference and mid-training workloads.
  • •Develop systems that power synthetic data generation and reinforcement learning pipelines at scale.
  • •Build high-performance inference platforms capable of serving and evaluating models across thousands of GPUs.
  • •Optimize throughput, latency, and GPU utilization for large language model inference and rollout workloads.
  • •Build infrastructure that supports reinforcement learning pipelines, including large-scale rollout generation, evaluation, and policy improvement loops.

Pay and Benefits

Perks:Health InsuranceDentalVisionParental LeaveRelocationEquity

Key Requirements

  • •Experience deploying and operating large-scale GPU systems for inference or model serving.
  • •Several years of hands-on experience building and running production infrastructure.
  • •Strong understanding of GPU performance characteristics and optimization techniques.
  • •Experience working with modern inference frameworks such as SGLang, Megatron, or similar high-performance LLM runtimes.
  • •Familiarity with distributed reinforcement learning infrastructure or rollout generation systems.
Experience:AIInfrastructureMachine learning
Skills:Problem-solvingCollaborationCommunication
Languages:English
Tech Stack:SGLangMegatronGPULLM runtimesKernel optimizationDistributed systemsReinforcement learning

Company Brief

Reflection AI
Builds frontier autonomous AI systems focused on autonomous coding agents (product: Asimov) to create organizational superintelligence, founded by former DeepMind/Google researchers and hiring across SF, NYC, London.
Industry: AI & Machine Learning
Company Size: Small (11 to 50 employees)
Growth: Growth Stage Startup
Funding: Series A
Headquarters: San Francisco, United States
Founded: 2024
WebsiteLinkedInGlassdoor