Member of Technical Staff - Storage Infrastructure

Prime Intellect
San Francisco, United States
Workplace: RemoteFull timeUSD 150,000 - 300,000Function: DevOps, Cloud & InfrastructureSkills: []

Build and operate storage systems that power frontier AI workloads. Own high-throughput, reliable access to datasets, checkpoints, and model artifacts as GPU clusters scale. Design and tune parallel filesystems, object storage, and NVMe caching; benchmark performance and concurrency; implement replication, recovery, backup, and durability targets; and run reliable storage operations through monitoring, automation, and access controls.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Prime Intellect
Prime Intellect
1 day ago

Member of Technical Staff - Storage Infrastructure

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 17 minutes agoStatus: Live

Job Summary

Build and operate storage systems that power frontier AI workloads. Own high-throughput, reliable access to datasets, checkpoints, and model artifacts as GPU clusters scale. Design and tune parallel filesystems, object storage, and NVMe caching; benchmark performance and concurrency; implement replication, recovery, backup, and durability targets; and run reliable storage operations through monitoring, automation, and access controls.
Location: San Francisco, United States
Workplace: Remote
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Design and operate storage architectures for training datasets, checkpointing, inference artifacts, and shared research workflows
  • •Deploy and tune parallel filesystems, object storage, and local NVMe caching for AI workloads
  • •Benchmark throughput, latency, metadata performance, and concurrent access for training and checkpoint workloads
  • •Build storage provisioning, capacity planning, lifecycle management, and operational automation
  • •Design and test replication, recovery, backup, and failure-handling procedures with durability/availability targets; troubleshoot performance and reliability issues across systems

Pay and Benefits

Salary: USD 150,000 - 300,000

Key Requirements

  • •3+ years building or operating production distributed storage systems
  • •Hands-on experience with at least one parallel/distributed filesystem or object storage (Lustre, BeeGFS, Ceph, GPFS)
  • •Strong Linux administration and performance troubleshooting skills
  • •Experience automating infrastructure operations in Python, Go, Bash (or similar)
  • •Understanding storage failure modes, data integrity, consistency, replication, and recovery
Experience:Distributed systemsAI infrastructureCloudOpen source
Tech Stack:LustreBeeGFSCephGPFSLinuxPythonGoBashNVMeSSDS3-compatibleRDMAGPUDirect StorageKubernetesSLURM

Company Brief

Prime Intellect
Builds a decentralized, open compute and training platform that enables distributed training and collective ownership of AI models, aggregating global GPU resources and offering tools for agentic RL and model evaluation.
Industry: AI & Machine Learning
Company Size: Small (11 to 50 employees)
Growth: Early Stage Startup
Funding: Seed
Headquarters: San Francisco, United States
Founded: 2024
WebsiteLinkedIn