Member of Technical Staff - Bare Metal & Fleet Provisioning

Prime Intellect
San Francisco
Workplace: OnsiteFull timeUSD 150,000 - 300,000 annuallyFunction: Transportation & Fleet OperationsExperience: 3+ yearsSkills: ["Debugging","Incident ownership","Collaboration","Reliability-focused automation","Problem-solving"]

Build the systems that turn bare-metal GPU servers into reliable, production-ready compute. Own the machine lifecycle from discovery and provisioning through validation, upgrades, repair, and secure reuse. Design automated discovery, PXE/iPXE boot, imaging, and configuration workflows, including safe rollouts for BIOS/BMC/NIC/GPU drivers/firmware. Integrate with SLURM, Kubernetes, and compute allocation systems, and deliver observability, quarantine, and repair workflows using Python/Go/Bash with Linux and infrastructure automation.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Prime Intellect
Prime Intellect
1 day ago

Member of Technical Staff - Bare Metal & Fleet Provisioning

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 15 minutes agoStatus: Live

Job Summary

Build the systems that turn bare-metal GPU servers into reliable, production-ready compute. Own the machine lifecycle from discovery and provisioning through validation, upgrades, repair, and secure reuse. Design automated discovery, PXE/iPXE boot, imaging, and configuration workflows, including safe rollouts for BIOS/BMC/NIC/GPU drivers/firmware. Integrate with SLURM, Kubernetes, and compute allocation systems, and deliver observability, quarantine, and repair workflows using Python/Go/Bash with Linux and infrastructure automation.
Location: San Francisco
Workplace: Onsite
Employment Type: Full time
Job Function: Transportation & Fleet Operations

Key Responsibilities

  • •Build automated discovery, network boot, OS imaging, and configuration workflows for GPU servers
  • •Automate BIOS, BMC, NIC, GPU driver, and firmware configuration with staged rollouts and safe recovery paths
  • •Develop hardware inventory and lifecycle services tracking identity, configuration, health, and readiness
  • •Create acceptance tests and burn-in workflows for GPUs, memory, storage, and interconnects before production
  • •Integrate provisioning and health checks with SLURM, Kubernetes, and compute allocation systems

Pay and Benefits

Salary: USD 150,000 - 300,000 annually

Key Requirements

  • •3+ years operating Linux servers or building bare-metal infrastructure automation in production
  • •Hands-on experience with PXE/iPXE, DHCP, image provisioning, and out-of-band management (Redfish or IPMI)
  • •Strong software engineering and debugging skills in Python, Go, or a comparable language, plus Bash
  • •Design reliable automation for partial failures, retries, and configuration drift
  • •Own operational incidents and collaborate across hardware, networking, and platform teams
Experience:3+ yearsInfrastructure automationBare-metalGPU infrastructure
Skills:DebuggingIncident ownershipCollaborationReliability-focused automationProblem-solving
Tech Stack:PythonGoBashLinuxSystemdPXEIPXEDHCPRedfishIPMIAnsibleTerraformSLURMKubernetesBMCNICBIOSKernelDriversOS image management

Company Brief

Prime Intellect
Builds a decentralized, open compute and training platform that enables distributed training and collective ownership of AI models, aggregating global GPU resources and offering tools for agentic RL and model evaluation.
Industry: AI & Machine Learning
Company Size: Small (11 to 50 employees)
Growth: Early Stage Startup
Funding: Seed
Headquarters: San Francisco, United States
Founded: 2024
WebsiteLinkedIn