Director, Site Operations

X AI
Memphis
Workplace: OnsiteFull timeFunction: Executive & General ManagementExperience: 7+ yearsEducation: bachelorsSkills: ["Communication","Prioritization","Analytical skills","Leadership","Accountability"]

Own node and rack uptime for SpaceXAI’s AI supercompute cluster across 5+ 24/7 sites, holding accountability for customer SLAs. Lead a 250+ person multi-shift operations organization and partner with facilities, network engineering, and tenants to drive remediation and planned/unplanned downtime reductions. Oversee SRE monitoring, incident response, root-cause analysis, and reliability procedures, while standardizing best practices and using operational data to continuously improve uptime and repair performance as the cluster scales.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
X AI
X AI
11 hours ago

Director, Site Operations

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 1 hour agoStatus: Live

Job Summary

Own node and rack uptime for SpaceXAI’s AI supercompute cluster across 5+ 24/7 sites, holding accountability for customer SLAs. Lead a 250+ person multi-shift operations organization and partner with facilities, network engineering, and tenants to drive remediation and planned/unplanned downtime reductions. Oversee SRE monitoring, incident response, root-cause analysis, and reliability procedures, while standardizing best practices and using operational data to continuously improve uptime and repair performance as the cluster scales.
Location: Memphis
Workplace: Onsite
Employment Type: Full time
Job Function: Executive & General Management
Seniority: Director level

Key Responsibilities

  • •Own node, rack, and cluster health across 5+ sites, accountable for customer SLAs and exceptional uptime 24/7.
  • •Lead a 250+ person organization of site managers, shift supervisors, and technicians across four 24/7 shifts, building excellence and accountability.
  • •Drive systematic remediation of failed nodes and racks, including command-line and physical intervention to reduce repair time.
  • •Partner across functions (facilities, network engineering, and tenants) to coordinate power/cooling fault mitigation, cluster upgrades, and downtime planning.
  • •Own site reliability engineering: proactive monitoring, reactive fault mitigation, root cause analyses, and site-wide reliability procedures and documentation.
Travel: High travel

Key Requirements

  • •Bachelor’s degree and 7+ years in large-scale operations, including 5+ years leading people leaders of technical teams.
  • •10+ years in large-scale operations with 5+ years leading people leaders of technical teams.
  • •Proven ability to lead large, multi-site, 24/7 operations organizations in fast-paced environments.
  • •Deep expertise in server hardware, cluster reliability, and data center technologies from deployment through lifecycle management.
  • •Experience owning uptime/SLA or reliability metrics for large compute clusters, including site reliability engineering experience and root cause analysis/procedure ownership.
Experience:7+ yearsData centerAIMachine learningHigh-performance computingSupercomputing
Education:Bachelor's
Skills:CommunicationPrioritizationAnalytical skillsLeadershipAccountability
Languages:English
Tech Stack:JiraPythonBash

Company Brief

X AI
Develops advanced artificial intelligence models and research aimed at building safe, general AI and understanding the fundamental nature of the universe. Focuses on large-scale AI systems, research publications, and building foundational AI capabilities.
Industry: AI & Machine Learning
Company Size: Medium (51 to 250 employees)
Growth: Early Stage Startup
Headquarters: San Francisco, United States
Founded: 2023
Website