Member of Technical Staff - Datacenter Operations

Prime Intellect
San Francisco
Workplace: HybridFull timeUSD 150,000 - 300,000 annuallyFunction: DevOps, Cloud & InfrastructureExperience: 3+ yearsSkills: ["Operational judgment","Documentation","Ownership","Customer obsession"]

Own operational readiness for the physical infrastructure behind a GPU cloud—coordinating rack deployments, maintenance, and incident response with datacenter partners. You’ll manage capacity readiness, hardware health documentation, and RMA workflows, while partnering with facilities on power, cooling, and monitoring. Build runbooks and escalation procedures, track failure trends and repair times, and automate operational reporting to reduce time to repair.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Prime Intellect
Prime Intellect
1 day ago

Member of Technical Staff - Datacenter Operations

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 16 minutes agoStatus: Live

Job Summary

Own operational readiness for the physical infrastructure behind a GPU cloud—coordinating rack deployments, maintenance, and incident response with datacenter partners. You’ll manage capacity readiness, hardware health documentation, and RMA workflows, while partnering with facilities on power, cooling, and monitoring. Build runbooks and escalation procedures, track failure trends and repair times, and automate operational reporting to reduce time to repair.
Location: San Francisco
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Mid level

Key Responsibilities

  • •Coordinate rack deployment, cabling, inventory, and acceptance testing for new GPU capacity with datacenter partners and engineering teams.
  • •Maintain accurate asset records, rack layouts, power allocations, cabling documentation, and spare-parts inventories.
  • •Lead hardware fault triage and coordinate remote hands, vendor escalations, component replacement, and RMA workflows.
  • •Establish maintenance plans and change procedures that minimize disruption and protect equipment and data.
  • •Track capacity readiness and hardware failure trends; automate repetitive reporting and workflows, and create runbooks for incident response.

Pay and Benefits

Salary: USD 150,000 - 300,000 annually
Equity and Bonus:Equity

Key Requirements

  • •3+ years in datacenter operations, hardware infrastructure, or production systems operations.
  • •Hands-on experience deploying and troubleshooting rack-mounted servers, networking equipment, and structured cabling.
  • •Experience coordinating datacenter providers, remote hands, and hardware vendors through deployments and incidents.
  • •Working knowledge of Linux diagnostics, BMC consoles, and server hardware health tools.
  • •Strong operational judgment, documentation habits, and ownership of issues through resolution.
Experience:3+ yearsDatacenter operationsProduction systems operations
Skills:Operational judgmentDocumentationOwnershipCustomer obsession
Tech Stack:LinuxBMC consolesServer hardware diagnosticsGPU serversPCIeRack power budgetingRedundant power pathsAirflowFiber cablingCopper cablingOpticsInventory automationHealth checksRMA workflowsNVIDIA DGX/HGX

Company Brief

Prime Intellect
Builds a decentralized, open compute and training platform that enables distributed training and collective ownership of AI models, aggregating global GPU resources and offering tools for agentic RL and model evaluation.
Industry: AI & Machine Learning
Company Size: Small (11 to 50 employees)
Growth: Early Stage Startup
Funding: Seed
Headquarters: San Francisco, United States
Founded: 2024
WebsiteLinkedIn