Research Engineer - Post-Training

Pluralis Research
United States, Australia
Workplace: RemoteFull timeFunction: Education & TrainingSkills: ["Mission alignment","Ownership","Ability to defend work","Cross-time-zone collaboration"]

Build and operate end-to-end RL post-training for decentralized, high-latency training over consumer GPUs and Macs. Own the rollout ingestion, reward computation, policy updates, and returning updated weights across a geo-distributed network. Adapt RL algorithms to asynchronous, partially trusted generation (staleness tolerance, off-policy corrections, efficient updates), create evaluation tooling, and ship the first decentralized post-trained model release to public artifacts.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Pluralis Research
Pluralis Research
5 days ago

Research Engineer - Post-Training

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 10 hours agoStatus: Live

Job Summary

Build and operate end-to-end RL post-training for decentralized, high-latency training over consumer GPUs and Macs. Own the rollout ingestion, reward computation, policy updates, and returning updated weights across a geo-distributed network. Adapt RL algorithms to asynchronous, partially trusted generation (staleness tolerance, off-policy corrections, efficient updates), create evaluation tooling, and ship the first decentralized post-trained model release to public artifacts.
Location: United States, Australia
Workplace: Remote
Employment Type: Full time
Job Function: Education & Training

Key Responsibilities

  • •Build the RL post-training stack end-to-end: rollout ingestion, reward computation, policy updates, and distributing updated weights back to the network.
  • •Adapt RL algorithms for asynchronous, high-latency, partially trusted generation (e.g., staleness tolerance, off-policy corrections, communication-efficient policy updates).
  • •Create evaluation tooling to demonstrate post-trained model improvements and support the first decentralized release as a public artifact.
  • •Establish technical direction for the post-training system and drive execution across algorithms and infrastructure.
  • •Develop the training workflow to run on consumer GPUs and Macs connected over the public internet with high latencies.

Pay and Benefits

Perks:EquityRemote WorkRelocation

Key Requirements

  • •Hands-on RL post-training experience on large language models (e.g., RLHF, RLVR, or reasoning-focused RL), including touching the systems layer for rollout generation, async training, and weight synchronization.
  • •Production-quality engineering with Python and PyTorch, including concurrency, failure handling, and profiling before optimizing.
  • •Research ability in RL post-training, asynchronous/distributed RL, or adjacent areas, with publications or unpublished work you can defend.
  • •Mission alignment with Protocol Learning and decentralized, trustless, sovereign AI.
  • •Comfort working across time zones with a remote-first, distributed team.
Experience:Reinforcement learningLLMRLHFReasoning-focused RLDistributed RLDecentralized training
Skills:Mission alignmentOwnershipAbility to defend workCross-time-zone collaboration
Languages:English
Tech Stack:PythonPyTorchRLHFRLVRVLLMSGLangP2P networkingNAT traversalGeo-distributed inference

Company Brief

Pluralis Research
Develops Protocol Learning for decentralized, multi‑participant training of foundation models so models remain unmaterialized and community‑owned, enabling open-source large‑scale AI without single‑party control.
Industry: AI & Machine Learning
Company Size: Micro (1 to 10 employees)
Growth: Early Stage Startup
Funding: Seed
Founded: 2024
Glassdoor
Glassdoor: 4.0
WebsiteLinkedIn