Machine Learning Engineer - ML Training Platform

Pluralis Research
United States, Australia
Workplace: RemoteFull timeFunction: Data Science & Machine LearningSkills: ["Mission alignment","Systems ownership","Incident response","Observability"]

Architect and scale a fully decentralized ML training and inference platform that runs across consumer nodes and multi-cloud instances. Own the resource orchestration layer for dynamic, fault-tolerant distributed GPU workloads across AWS, GCP, and Azure, including checkpointing, streaming datasets, health monitoring, and resilient retries. Build networking systems that simulate real-world conditions (bandwidth, latency, packet loss) and handle node churn. Collaborate on continuous experimentation at frontier scale.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Pluralis Research
Pluralis Research
4 days ago

Machine Learning Engineer - ML Training Platform

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 12 hours agoStatus: Live

Job Summary

Architect and scale a fully decentralized ML training and inference platform that runs across consumer nodes and multi-cloud instances. Own the resource orchestration layer for dynamic, fault-tolerant distributed GPU workloads across AWS, GCP, and Azure, including checkpointing, streaming datasets, health monitoring, and resilient retries. Build networking systems that simulate real-world conditions (bandwidth, latency, packet loss) and handle node churn. Collaborate on continuous experimentation at frontier scale.
Location: United States, Australia
Workplace: Remote
Employment Type: Full time
Job Function: Data Science & Machine Learning
Seniority: Mid level

Key Responsibilities

  • •Design resource management systems to provision and orchestrate compute across AWS, GCP, and Azure using infrastructure-as-code.
  • •Architect fault-tolerant distributed ML infrastructure for training and inference, including GPU clusters, checkpointing, and resilient retry strategies.
  • •Build infrastructure for real-world networking conditions (bandwidth shaping, latency injection, packet loss) and handle node churn.
  • •Manage state synchronization and concurrent operations across hundreds of heterogeneous nodes.
  • •Develop services that keep continuous experimentation and large-scale training running across decentralized environments.

Pay and Benefits

Equity and Bonus:Equity
Perks:EquityRemote WorkVisa SponsorshipRelocation

Key Requirements

  • •Production experience with infrastructure-as-code (Pulumi/Terraform/CloudFormation) managing multi-cloud deployments at scale.
  • •Hands-on Docker/Kubernetes (EKS) experience running GPU workloads and heterogeneous clusters.
  • •Strong understanding of distributed training workflows including checkpointing, data sharding, model versioning, and long-running job orchestration.
  • •Experience with decentralized networking concepts such as P2P, NAT traversal, and traffic shaping under real bandwidth constraints.
  • •Strong Python engineering (asyncio/concurrency/retry logic) with observability and SRE practices (Prometheus/Grafana, profiling, incident response).
Experience:Machine learningDistributed systemsMulti-cloudDistributed trainingSRE
Skills:Mission alignmentSystems ownershipIncident responseObservability
Languages:English
Tech Stack:AWSGCPAzurePulumiTerraformCloudFormationDockerKubernetesEKSNVIDIAPythonAsyncioS3PrometheusGrafanaNAT traversalP2PTraffic shapingCloud SDKsCLI tooling

Eligibility

Work Authorization:Sponsorship available.

Company Brief

Pluralis Research
Develops Protocol Learning for decentralized, multi‑participant training of foundation models so models remain unmaterialized and community‑owned, enabling open-source large‑scale AI without single‑party control.
Industry: AI & Machine Learning
Company Size: Micro (1 to 10 employees)
Growth: Early Stage Startup
Funding: Seed
Founded: 2024
Glassdoor
Glassdoor: 4.0
WebsiteLinkedIn