Research Engineer - RL Infrastructure

Prime Intellect
San Francisco
Workplace: HybridFull timeFunction: Research & Scientific (R&D)Skills: ["Ownership","Problem-solving","Collaboration","Proactive"]

A Research Engineer to design and optimize the systems layer behind large-scale RL training, focusing on kernels, memory and communication efficiency, and distributed training workloads. You’ll work with researchers and infra engineers to push throughput and reliability toward hardware limits, shaping the RL training stack and contributing to open-source infrastructure.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Prime Intellect
Prime Intellect
5 months ago

Research Engineer - RL Infrastructure

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 11 hours agoStatus: Live

Job Summary

A Research Engineer to design and optimize the systems layer behind large-scale RL training, focusing on kernels, memory and communication efficiency, and distributed training workloads. You’ll work with researchers and infra engineers to push throughput and reliability toward hardware limits, shaping the RL training stack and contributing to open-source infrastructure.
Location: San Francisco
Workplace: Hybrid
Employment Type: Full time
Job Function: Research & Scientific (R&D)

Key Responsibilities

  • •Build and optimize the systems infrastructure behind large-scale RL and distributed training workloads.
  • •Improve end-to-end training efficiency across compute, memory, networking, and scheduling layers.
  • •Design and implement low-level performance optimizations, including kernels, communication paths, and runtime improvements.
  • •Work on distributed training systems spanning data, tensor, and pipeline parallel workloads.
  • •Collaborate closely with researchers and infrastructure engineers to translate bottlenecks into concrete systems improvements.

Pay and Benefits

Equity and Bonus:Equity
Perks:Visa SponsorshipRemote WorkEquity

Key Requirements

  • •Strong systems engineering experience in AI/ML infrastructure, especially around large-scale model training or inference.
  • •Deep familiarity with PyTorch and distributed training frameworks such as PyTorch Distributed, DeepSpeed, FSDP, Megatron, vLLM, Ray, or related tooling.
  • •Experience optimizing training performance across kernels, memory movement, communication overhead, or parallelization strategy.
  • •Hands-on experience with large-scale training techniques including data parallelism, tensor parallelism, and pipeline parallelism.
  • •Strong understanding of GPU architecture, profiling, and performance debugging.
Experience:Ai/ml infrastructureDistributed trainingReinforcement learning
Skills:OwnershipProblem-solvingCollaborationProactive
Languages:English
Tech Stack:PyTorchPyTorch DistributedDeepSpeedFSDPMegatronVLLMRayCUDATritonKernelsGPUProfilingDistributed training

Company Brief

Prime Intellect
Builds a decentralized, open compute and training platform that enables distributed training and collective ownership of AI models, aggregating global GPU resources and offering tools for agentic RL and model evaluation.
Industry: AI & Machine Learning
Company Size: Small (11 to 50 employees)
Growth: Early Stage Startup
Funding: Seed
Headquarters: San Francisco, United States
Founded: 2024
WebsiteLinkedIn