Research Engineer - Distributed Training

Prime Intellect
San Francisco, United States
Workplace: OnsiteFull timeFunction: Education & TrainingSkills: ["Communication","Collaboration","Problem-solving"]

Research Engineer focused on distributed training in Prime Intellect’s open superintelligence stack. Develop, optimize, and publish on decentralized training orchestration, AI/ML pipelines, and open-source frameworks, with a focus on scalable, secure infrastructure and advancing AI capabilities for researchers, startups, and enterprises.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Prime Intellect
Prime Intellect
2 years ago

Research Engineer - Distributed Training

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 5 hours agoStatus: Live

Job Summary

Research Engineer focused on distributed training in Prime Intellect’s open superintelligence stack. Develop, optimize, and publish on decentralized training orchestration, AI/ML pipelines, and open-source frameworks, with a focus on scalable, secure infrastructure and advancing AI capabilities for researchers, startups, and enterprises.
Location: San Francisco, United States
Workplace: Onsite
Employment Type: Full time
Job Function: Education & Training

Key Responsibilities

  • •Lead and participate in novel research to build a massive scale, highly reliable and secure decentralized training orchestration solution.
  • •Optimize the performance, cost, and resource utilization of AI workloads by leveraging the most recent advances for compute & memory optimization techniques.
  • •Contribute to the development of our open-source libraries and frameworks for distributed model training.
  • •Publish research in top-tier AI conferences such as ICML & NeurIPS.
  • •Distill highly technical project outcomes in layman approachable technical blogs to our customers and developers.

Pay and Benefits

Equity and Bonus:Equity
Perks:Equity IncentivesVisa SponsorshipRemote Work

Key Requirements

  • •Strong background in AI/ML engineering with end-to-end pipelines for training and deploying large-scale AI models.
  • •Deep expertise in distributed training techniques and frameworks (e.g., PyTorch Distributed, DeepSpeed, MosaicML’s LLM Foundry) and tools (e.g. Ray).
  • •Experience in large-scale model training with data, tensor and pipeline parallelism.
  • •Solid understanding of MLOps, including model versioning, experiment tracking, and CI/CD pipelines.
  • •Passion for decentralized AI model training and democratizing access to AI capabilities.
Experience:AI/MLDistributed trainingMLOps
Skills:CommunicationCollaborationProblem-solving
Languages:English
Tech Stack:PyTorch DistributedDeepSpeedMosaicML LLM FoundryRay

Company Brief

Prime Intellect
Builds a decentralized, open compute and training platform that enables distributed training and collective ownership of AI models, aggregating global GPU resources and offering tools for agentic RL and model evaluation.
Industry: AI & Machine Learning
Company Size: Small (11 to 50 employees)
Growth: Early Stage Startup
Funding: Seed
Headquarters: San Francisco, United States
Founded: 2024
WebsiteLinkedIn