Machine Learning Engineer - Distributed ML Systems

Pluralis Research
San Francisco
Workplace: RemoteFull timeFunction: Data Science & Machine LearningExperience: 5+ yearsSkills: ["Python","DeepSpeed","Megatron","FSDP","GPU","GPUs","NAT traversal","GRPC","Distributed training","Model parallelism","Data parallelism","Tensor parallelism","Pipeline parallelism"]

Senior/Staff engineers will design and implement large-scale distributed training systems for foundation-model style protocols, focusing on training across heterogeneous hardware over consumer-grade networks, with emphasis on model-parallelism, low-bandwidth optimization, and resilient, decentralized coordination.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Pluralis Research
Pluralis Research
4 months ago

Machine Learning Engineer - Distributed ML Systems

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 11 hours agoStatus: Live

Job Summary

Senior/Staff engineers will design and implement large-scale distributed training systems for foundation-model style protocols, focusing on training across heterogeneous hardware over consumer-grade networks, with emphasis on model-parallelism, low-bandwidth optimization, and resilient, decentralized coordination.
Location: San Francisco
Workplace: Remote
Employment Type: Full time
Job Function: Data Science & Machine Learning

Key Responsibilities

  • •Design and implement large-scale distributed training systems optimized for heterogeneous hardware operating under low-bandwidth, high-latency conditions.
  • •Develop and optimize model-parallel training strategies (data, tensor, pipeline parallelism) with custom sharding techniques that minimize communication overhead.
  • •Optimize GPU utilization, memory efficiency, and compute performance across distributed nodes.
  • •Implement robust checkpointing, state synchronization, and recovery mechanisms for long-running, fault-prone training jobs.
  • •Build monitoring and metrics systems to track training progress, model quality, and system bottlenecks.

Pay and Benefits

Equity and Bonus:Equity
Perks:Remote FirstVisa SponsorshipEquity

Key Requirements

  • •Strong experience building and operating distributed systems in production.
  • •Hands-on expertise with distributed training frameworks (FSDP, DeepSpeed, Megatron, or similar).
  • •Deep understanding of model parallelism (data, tensor, pipeline parallelism).
  • •Expert-level Python with production experience (concurrency, error handling, retry logic, clean architecture).
  • •Strong networking fundamentals: P2P systems, gRPC, routing, NAT traversal, distributed coordination.
Experience:5+ yearsAI/MLDistributed systems
Skills:PythonDeepSpeedMegatronFSDPGPUGPUsNAT traversalGRPCDistributed trainingModel parallelismData parallelismTensor parallelismPipeline parallelism
Languages:English
Tech Stack:PythonDeepSpeedMegatronFSDPGPUGRPCNAT traversalDistributed systems

Eligibility

Work Authorization:Sponsorship available.

Company Brief

Pluralis Research
Develops Protocol Learning for decentralized, multi‑participant training of foundation models so models remain unmaterialized and community‑owned, enabling open-source large‑scale AI without single‑party control.
Industry: AI & Machine Learning
Company Size: Micro (1 to 10 employees)
Growth: Early Stage Startup
Funding: Seed
Founded: 2024
Glassdoor
Glassdoor: 4.0
WebsiteLinkedIn