Member of Technical Staff - Research Software Engineer

Reflection AI
New York, London, San Francisco
Workplace: OnsiteFull timeFunction: Software EngineeringSkills: ["Problem-solving","Communication","Collaboration"]

Bridge the gap between AI research and production by designing and optimizing core training infrastructure for frontier models. You’ll build scalable RL training loops, distributed GPU systems, and large-scale data pipelines, collaborating with researchers to turn cutting-edge ideas into reliable, production-grade training systems used across thousands of GPUs and petabyte-scale datasets.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Reflection AI
Reflection AI
4 months ago

Member of Technical Staff - Research Software Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 8 hours agoStatus: Live

Job Summary

Bridge the gap between AI research and production by designing and optimizing core training infrastructure for frontier models. You’ll build scalable RL training loops, distributed GPU systems, and large-scale data pipelines, collaborating with researchers to turn cutting-edge ideas into reliable, production-grade training systems used across thousands of GPUs and petabyte-scale datasets.
Location: New York, London, San Francisco
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering

Key Responsibilities

  • •Designing and optimizing large-scale training loops and data pipelines.
  • •Implementing state-of-the-art techniques and ensuring numerical stability and computational efficiency.
  • •Building internal tooling for launching, monitoring, and reproducing complex experiments.
  • •Diagnosing bottlenecks across the training stack (GPU memory, communication overhead, dataloader stalls).
  • •Translating research prototypes into reusable, production-grade infrastructure.

Pay and Benefits

Equity and Bonus:Equity
Perks:Health InsuranceDentalVisionLife InsuranceDisability InsuranceRelocationPaid Leave

Key Requirements

  • •Strong software engineer with experience in machine learning or ML infrastructure
  • •Deep experience in Reinforcement Learning Systems, Distributed Training & Inference, or Data Infrastructure
  • •Ability to translate research prototypes into production-grade infrastructure
  • •Experience optimizing large-scale training loops, data pipelines, and GPU memory usage
  • •Comfort building internal tooling for launching, monitoring, and reproducing experiments
Experience:Machine learningOpen foundational modelsAi infrastructure
Skills:Problem-solvingCommunicationCollaboration
Tech Stack:PyTorchJAXMegatronTritonRayKubernetesSlurmNCCLRDMAFSDPZeROGPUMegatron-style training stacksDataloaderDistributed training

Company Brief

Reflection AI
Builds frontier autonomous AI systems focused on autonomous coding agents (product: Asimov) to create organizational superintelligence, founded by former DeepMind/Google researchers and hiring across SF, NYC, London.
Industry: AI & Machine Learning
Company Size: Small (11 to 50 employees)
Growth: Growth Stage Startup
Funding: Series A
Headquarters: San Francisco, United States
Founded: 2024
WebsiteLinkedInGlassdoor