Research Engineer - Pre-training

Pluralis Research
United States, Australia
Workplace: RemoteFull timeFunction: Education & TrainingSkills: []

Build the distributed training system behind Protocol Learning, scaling from an 8B decentralized run to frontier-scale large models. You’ll implement and optimize parallelism across heterogeneous GPUs under low-bandwidth, high-latency conditions, add communication-efficient performance improvements, and make runs resilient to node churn. You’ll also develop run instrumentation to surface throughput, bottlenecks, and model quality across hundreds of connected devices.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Pluralis Research
Pluralis Research
5 days ago

Research Engineer - Pre-training

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 10 hours agoStatus: Live

Job Summary

Build the distributed training system behind Protocol Learning, scaling from an 8B decentralized run to frontier-scale large models. You’ll implement and optimize parallelism across heterogeneous GPUs under low-bandwidth, high-latency conditions, add communication-efficient performance improvements, and make runs resilient to node churn. You’ll also develop run instrumentation to surface throughput, bottlenecks, and model quality across hundreds of connected devices.
Location: United States, Australia
Workplace: Remote
Employment Type: Full time
Job Function: Education & Training

Key Responsibilities

  • •Implement and optimize distributed model-parallel training, including data, pipeline, and tensor parallelism on heterogeneous GPUs under constrained network conditions.
  • •Reduce communication overhead while maintaining model convergence in challenging environments.
  • •Engineer elasticity and fault tolerance so training survives node churn via checkpointing, state synchronization, and recovery.
  • •Build monitoring and instrumentation to show throughput, bottlenecks, and model quality across hundreds of devices.

Pay and Benefits

Equity and Bonus:Equity
Perks:EquityRemote WorkVisa SponsorshipRelocation Support

Key Requirements

  • •Hands-on distributed training across many devices in PyTorch with FSDP, DeepSpeed, Megatron (or equivalent), including data, tensor, and pipeline parallelism.
  • •Production-quality Python with concurrency, failure handling, and profiling before optimizing.
  • •Evidence of execution through shipped systems, research code, open-source work, or serious personal projects.
  • •Deep alignment with Protocol Learning as the viable third path for collective, trustless, and sovereign AI.
  • •Experience with training or serving large language models (e.g., Nemotron, Qwen, OLMo).
Experience:Machine learningDistributed trainingOpen-sourceLarge language modelsDecentralized training
Languages:English
Tech Stack:PythonPyTorchFSDPDeepSpeedMegatronModel-parallel trainingData parallelismTensor parallelismPipeline parallelismCheckpointingRLNAT traversalP2P networking

Eligibility

Work Authorization:Sponsorship available.

Company Brief

Pluralis Research
Develops Protocol Learning for decentralized, multi‑participant training of foundation models so models remain unmaterialized and community‑owned, enabling open-source large‑scale AI without single‑party control.
Industry: AI & Machine Learning
Company Size: Micro (1 to 10 employees)
Growth: Early Stage Startup
Funding: Seed
Founded: 2024
Glassdoor
Glassdoor: 4.0
WebsiteLinkedIn