Research Scientist / Engineer – Reinforcement Learning Infrastructure

Luma
Redwood City
Workplace: HybridFull timeFunction: Data Science & Machine LearningSkills: ["System design","Debugging","Building scalable infrastructure","Attention to correctness","Collaboration"]

Build and scale distributed reinforcement learning post-training infrastructure at frontier scale. Own the full RL loop—trainer orchestration, high-throughput rollout generation, environment execution, and reward computation—across thousands of GPUs. Design scalable RL environments for agentic, multi-step tasks, develop robust reward and evaluation/monitoring tooling, and improve training efficiency and stability to turn new ideas into production runs.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Luma
Luma
2 weeks ago

Research Scientist / Engineer – Reinforcement Learning Infrastructure

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 6 hours agoStatus: Live

Job Summary

Build and scale distributed reinforcement learning post-training infrastructure at frontier scale. Own the full RL loop—trainer orchestration, high-throughput rollout generation, environment execution, and reward computation—across thousands of GPUs. Design scalable RL environments for agentic, multi-step tasks, develop robust reward and evaluation/monitoring tooling, and improve training efficiency and stability to turn new ideas into production runs.
Location: Redwood City
Workplace: Hybrid
Employment Type: Full time
Job Function: Data Science & Machine Learning

Key Responsibilities

  • •Design, build, and scale distributed RL post-training systems, orchestrating trainer, rollout, environment, and reward workloads across thousands of GPUs.
  • •Build high-throughput rollout generation, integrating inference engines (vLLM, SGLang), weight synchronization, and asynchronous/off-policy schemes.
  • •Design scalable RL environments for agentic, multi-step tasks (sandboxed code execution, tool use, computer use, multimodal interaction) for reproducible large-scale training.
  • •Build reward infrastructure including verifiable/programmatic rewards, reward-model serving, LLM-as-judge pipelines, and defenses against reward hacking.
  • •Develop evaluation, monitoring, and debugging tooling to keep large RL runs stable while advancing training efficiency and stability.

Key Requirements

  • •Hands-on experience with post-training LLMs with RL (PPO/GRPO-family, RLHF, RLVR) at meaningful scale.
  • •Extensive distributed PyTorch training and parallelism (FSDP, Tensor/Pipeline/Expert Parallel) for foundation models.
  • •Experience building RL environments, reward functions, verifiers, or evaluation harnesses for LLM agents, including sandboxed execution and multi-turn tool use.
  • •Deep familiarity with RL post-training frameworks (veRL, OpenRLHF, TRL, Ray orchestration) and rollout inference engines (vLLM, SGLang).
  • •Strong understanding of GPU clusters, networking, and communication libraries (NCCL, MPI) under mixed training and inference workloads.
Experience:Reinforcement learningLLM post-trainingDistributed trainingAgentic environments
Skills:System designDebuggingBuilding scalable infrastructureAttention to correctnessCollaboration
Tech Stack:PyTorchFSDPTensor ParallelPipeline ParallelExpert ParallelVLLMSGLangVeRLOpenRLHFTRLRayKubernetesRay orchestrationNCCLMPIPPOGRPORLHFRLVR

Company Brief

Luma
Develops AI-powered tools for capturing, editing, and rendering high-quality 3D scenes from photos and videos, enabling creators to generate photorealistic 3D assets and spatial experiences.
Industry: AR/VR & Spatial Computing
Website