Research Scientist / Engineer – Reinforcement Learning Infrastructure
Luma
Redwood City, London
Workplace: HybridFull timeUSD 187,500 - 395,000 annuallyFunction: Data Science & Machine LearningSkills: ["Debugging","Collaboration"]Build and scale distributed reinforcement learning post-training infrastructure for large multimodal foundation models. You’ll orchestrate trainer/rollout/environment/reward workloads across thousands of GPUs, optimize high-throughput rollout generation with inference engines (vLLM, SGLang), and design hermetic RL environments. Own reward, evaluation, monitoring, and debugging tooling to keep runs stable and diagnose regressions, while collaborating with researchers to productionize new RL approaches.

