Research Engineer - Post-Training
United States, Australia
Workplace: RemoteFull timeFunction: Education & TrainingSkills: ["Mission alignment","Ownership","Ability to defend work","Cross-time-zone collaboration"]Build and operate end-to-end RL post-training for decentralized, high-latency training over consumer GPUs and Macs. Own the rollout ingestion, reward computation, policy updates, and returning updated weights across a geo-distributed network. Adapt RL algorithms to asynchronous, partially trusted generation (staleness tolerance, off-policy corrections, efficient updates), create evaluation tooling, and ship the first decentralized post-trained model release to public artifacts.
Loading
Loading job details...
Preparing the role view and application actions.

