Director of AI Infrastructure

AI2
Seattle
Workplace: OnsiteFull time176,400 - 264,600 annuallyFunction: Data Science & Machine LearningExperience: 12+ yearsEducation: bachelorsSkills: ["Leadership","Communication","Problem-solving","Strategic thinking","Collaboration","Infrastructure","Hpc","Gpu","Clustering","Orchestration","Cloud"]

Lead the design and operation of AI2's HPC infrastructure powering frontier AI research. You will oversee on-prem GPU clusters and a hybrid cloud orchestration layer across AWS/GCP, shape storage strategy for petascale data, and manage the GPU compute budget while ensuring researchers have reliable, high-velocity access.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
AI2
AI2
4 months ago

Director of AI Infrastructure

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 22 hours agoStatus: Live

Job Summary

Lead the design and operation of AI2's HPC infrastructure powering frontier AI research. You will oversee on-prem GPU clusters and a hybrid cloud orchestration layer across AWS/GCP, shape storage strategy for petascale data, and manage the GPU compute budget while ensuring researchers have reliable, high-velocity access.
Location: Seattle
Workplace: Onsite
Employment Type: Full time
Job Function: Data Science & Machine Learning
Seniority: Director level

Key Responsibilities

  • •Oversee the availability and performance of dense on-prem GPU clusters and partner with hardware vendors/internal teams to meet frontier model training demands
  • •Direct the strategy for Beaker, the internal orchestration platform, to optimize job scheduling across on-prem and elastic cloud resources (AWS/GCP)
  • •Develop and execute a long-term storage roadmap balancing high-throughput training with durable petascale data storage
  • •Act as the primary steward of the GPU compute budget and make data-driven decisions on cloud bursting vs. on-prem capacity
  • •Serve as the technical bridge to research teams to ensure infrastructure accelerates research objectives

Pay and Benefits

Salary: 176,400 - 264,600 annually
Perks:Health InsuranceDentalVision401kPaid LeavePersonal DaysCommuting AllowanceGym MembershipHolidaysAnnual Bonus

Key Requirements

  • •12+ years in infrastructure, systems engineering, or HPC, with at least 5 years in a leadership role managing multi-disciplinary engineering teams
  • •Bachelor’s degree in related field; relevant advanced degree may substitute for equivalent years of technical work experience
  • •Direct experience managing large-scale NVIDIA GPU clusters and high-performance networking (InfiniBand/RoCE)
  • •Strong background in Kubernetes, Slurm, or similar orchestration frameworks, particularly in hybrid-cloud configurations
  • •Experience with distributed filesystems (e.g., WEKA, Ceph, Lustre) and cloud storage integration at scale
Experience:12+ yearsAI researchHigh-performance computingOpen source
Education:Bachelor's
Skills:LeadershipCommunicationProblem-solvingStrategic thinkingCollaborationInfrastructureHpcGpuClusteringOrchestrationCloud
Languages:English
Tech Stack:LinuxInfiniBandNCCLKubernetesSlurmWEKACephLustrePythonGoGPUNVIDIAAWSGCP

Company Brief

AI2
The Allen Institute for AI (AI2) is a nonprofit research institute that builds AI systems and conducts fundamental research to contribute to scientific discovery, natural language understanding, and open-source tools for the AI community.
Industry: AI & Machine Learning
Company Size: Medium (51 to 250 employees)
Growth: Nonprofit & NGO
Funding: Grant Funded
Headquarters: Seattle, United States
Founded: 2014
WebsiteLinkedIn