Member of Technical Staff - GPU Infrastructure

Prime Intellect
San Francisco, United States
Workplace: OnsiteFull timeFunction: Solutions Engineering & Sales EngineeringExperience: 3+ yearsSkills: ["Communication","Leadership","Problem-solving"]

Technical, customer-facing role designing and deploying GPU-accelerated infrastructure for large-scale AI/ML workloads. Partners with clients to define GPU cluster architectures, develops deployment strategies for LLM training and HPC workloads, and leads production operations with SLURM/Kubernetes, InfiniBand, and advanced GPU tooling. Focused on delivering production-ready systems and ongoing optimization for high-density GPU deployments.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Prime Intellect
Prime Intellect
1 year ago

Member of Technical Staff - GPU Infrastructure

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 5 hours agoStatus: Live

Job Summary

Technical, customer-facing role designing and deploying GPU-accelerated infrastructure for large-scale AI/ML workloads. Partners with clients to define GPU cluster architectures, develops deployment strategies for LLM training and HPC workloads, and leads production operations with SLURM/Kubernetes, InfiniBand, and advanced GPU tooling. Focused on delivering production-ready systems and ongoing optimization for high-density GPU deployments.
Location: San Francisco, United States
Workplace: Onsite
Employment Type: Full time
Job Function: Solutions Engineering & Sales Engineering

Key Responsibilities

  • •Serve as primary technical escalation point for customer infrastructure issues across hardware, drivers, networking, and software.
  • •Diagnose and resolve complex problems across the full stack in GPU HPC environments.
  • •Implement monitoring, alerting, and automated remediation systems for GPU deployments.
  • •Provide 24/7 on-call support for critical customer deployments and create runbooks/documentation for operations teams.
  • •Design and present architectural recommendations to technical and executive stakeholders for GPU clusters and associated workflows.

Key Requirements

  • •3+ years hands-on experience with GPU clusters and HPC environments.
  • •Deep expertise with SLURM and Kubernetes in production GPU settings.
  • •Proven experience with InfiniBand configuration and troubleshooting.
  • •Strong understanding of NVIDIA GPU architecture, CUDA ecosystem, and driver stack.
  • •Experience with infrastructure automation tools (Ansible, Terraform) and proficiency in Python, Bash, and systems programming
Experience:3+ yearsGPU infrastructureHPCCUDASLURMKubernetes
Skills:CommunicationLeadershipProblem-solving
Languages:English
Tech Stack:PythonBashCUDANVIDIASLURMKubernetesInfiniBandRoCENVLinkLustreBeeGFSGPFSDockerContainerdAnsibleTerraformLinux

Company Brief

Prime Intellect
Builds a decentralized, open compute and training platform that enables distributed training and collective ownership of AI models, aggregating global GPU resources and offering tools for agentic RL and model evaluation.
Industry: AI & Machine Learning
Company Size: Small (11 to 50 employees)
Growth: Early Stage Startup
Funding: Seed
Headquarters: San Francisco, United States
Founded: 2024
WebsiteLinkedIn