Member of Technical Staff - Training Platform

Prime Intellect
San Francisco
Workplace: OnsiteFull timeUSD 150,000 - 300,000 annuallyFunction: Education & TrainingSkills: ["Communication","Problem-solving","Teamwork"]

Join Prime Intellect to build a hosted training platform for open superintelligence. You’ll architect Kubernetes-based training/inference orchestration, develop Python control-plane agents, and ship a Frontend for job submission, monitoring, and logs. This role spans platform, infrastructure, and developer tooling, enabling researchers and enterprises to train and fine-tune models at frontier scale.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Prime Intellect
Prime Intellect
3 months ago

Member of Technical Staff - Training Platform

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 3 hours agoStatus: Live

Job Summary

Join Prime Intellect to build a hosted training platform for open superintelligence. You’ll architect Kubernetes-based training/inference orchestration, develop Python control-plane agents, and ship a Frontend for job submission, monitoring, and logs. This role spans platform, infrastructure, and developer tooling, enabling researchers and enterprises to train and fine-tune models at frontier scale.
Location: San Francisco
Workplace: Onsite
Employment Type: Full time
Job Function: Education & Training

Key Responsibilities

  • •Design and operate Kubernetes-based training and inference orchestration across multi-cluster, multi-cloud GPU fleets
  • •Build and maintain Helm charts that compose trainers, inference servers, environment servers, and supporting services into reproducible Training stacks
  • •Develop the Python control-plane agents that watch pods, report run state to the platform, and keep clusters in sync
  • •Implement scheduling and autoscaling for heterogeneous hardware (H100/H200/B200) using KEDA, LeaderWorkerSet, taints/tolerations, and gang scheduling
  • •Run a tight GitOps workflow - every change ships through PRs, Helm values, and CI

Pay and Benefits

Salary: USD 150,000 - 300,000 annually
Equity and Bonus:Equity
Perks:Remote WorkVisa SponsorshipRelocationLearning Budget

Key Requirements

  • •Proficient in Python backend development with FastAPI and SQLAlchemy.
  • •Strong Kubernetes operations experience (Helm, CRDs, operators, KEDA) and multi-cluster orchestration.
  • •Cloud platform experience (GCP preferred) and infrastructure automation (Terraform, Ansible) with a GitOps mindset.
  • •Experience building developer tools, dashboards, and live-monitoring UIs (React/Next.js, TypeScript).
  • •Understanding of distributed training fundamentals and RL/LLM workflows (LoRA, QLoRA, fine-tuning).
Experience:Ai infrastructureMachine learningDistributed systemsCloud
Skills:CommunicationProblem-solvingTeamwork
Languages:English
Tech Stack:KubernetesHelmKEDATerraformAnsibleGCPGKEFastAPISQLAlchemyPythonTypeScriptReactNext.jsTailwindShadcnTRPCTanStack QueryPrometheusGrafanaLoki

Eligibility

Work Authorization:Sponsorship available.

Company Brief

Prime Intellect
Builds a decentralized, open compute and training platform that enables distributed training and collective ownership of AI models, aggregating global GPU resources and offering tools for agentic RL and model evaluation.
Industry: AI & Machine Learning
Company Size: Small (11 to 50 employees)
Growth: Early Stage Startup
Funding: Seed
Headquarters: San Francisco, United States
Founded: 2024
WebsiteLinkedIn