Member of Technical Staff - Compute Platform

Prime Intellect
San Francisco
Workplace: HybridFull timeFunction: Software EngineeringSkills: ["Collaboration","Reliability focus","Open development"]

Build and evolve Prime Intellect’s full-stack Compute Platform, spanning AI platform software and infrastructure. You’ll develop web interfaces, REST APIs, real-time monitoring, and job control features in Python, while also designing distributed training infrastructure and reliability improvements using Rust and automation. Work with container orchestration, cloud resources, scheduling for CPU/GPU/TPU, and observability tooling to support post-training at frontier scale.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Prime Intellect
Prime Intellect
2 months ago

Member of Technical Staff - Compute Platform

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 5 hours agoStatus: Live

Job Summary

Build and evolve Prime Intellect’s full-stack Compute Platform, spanning AI platform software and infrastructure. You’ll develop web interfaces, REST APIs, real-time monitoring, and job control features in Python, while also designing distributed training infrastructure and reliability improvements using Rust and automation. Work with container orchestration, cloud resources, scheduling for CPU/GPU/TPU, and observability tooling to support post-training at frontier scale.
Location: San Francisco
Workplace: Hybrid
Employment Type: Full time
Job Function: Software Engineering
Seniority: Sr. Manager level

Key Responsibilities

  • •Build intuitive web interfaces for AI workload management and monitoring.
  • •Develop REST APIs and backend services in Python, and create real-time monitoring and debugging tools.
  • •Implement user-facing resource management and job control features.
  • •Design and implement distributed training infrastructure in Rust and build high-performance networking and coordination components.
  • •Create infrastructure automation pipelines and manage cloud resources, container orchestration, and scheduling for heterogeneous hardware.

Key Requirements

  • •Strong Python backend development (FastAPI, async) with RESTful API design and implementation
  • •Modern frontend development using TypeScript with React/Next.js and Tailwind, including developer tools and dashboards
  • •Systems programming experience with Rust, including building distributed training infrastructure
  • •Infrastructure automation and orchestration experience using Ansible and Terraform, plus Kubernetes and cloud platform expertise (GCP preferred)
  • •Observability experience with Prometheus and Grafana, and ability to maintain reliability and security when integrating new features
Experience:AI infrastructureDeveloper toolsOpen-source
Skills:CollaborationReliability focusOpen development
Tech Stack:PythonFastAPIAsyncREST APIsTypeScriptReactNext.jsTailwindRustAnsibleTerraformKubernetesGCPPrometheusGrafanaWebSocketCPUGPUTPU

Company Brief

Prime Intellect
Builds a decentralized, open compute and training platform that enables distributed training and collective ownership of AI models, aggregating global GPU resources and offering tools for agentic RL and model evaluation.
Industry: AI & Machine Learning
Company Size: Small (11 to 50 employees)
Growth: Early Stage Startup
Funding: Seed
Headquarters: San Francisco, United States
Founded: 2024
WebsiteLinkedIn