Hardware Operations Engineer

OpenAI
United States
Workplace: RemoteFull timeUSD 86,400 - 228,000 annuallyFunction: Field Service, Maintenance & Skilled TradesExperience: 8+ yearsSkills: ["Communication","Leadership","Troubleshooting","Problem-solving","Cross-functional collaboration"]

Senior on-site hardware operations lead for OpenAI’s flagship AI campus, overseeing server, GPU, storage, and rack-level hardware reliability. Drives triage of complex hardware failures, partners with Fleet Health Engineering, conducts root cause analysis, and coordinates with Oracle and OEMs on repairs and lifecycle activities. Establishes maintenance standards and operational runbooks to scale across Stargate campuses.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
OpenAI
OpenAI
11 months ago

Hardware Operations Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 1 hour agoStatus: Live

Job Summary

Senior on-site hardware operations lead for OpenAI’s flagship AI campus, overseeing server, GPU, storage, and rack-level hardware reliability. Drives triage of complex hardware failures, partners with Fleet Health Engineering, conducts root cause analysis, and coordinates with Oracle and OEMs on repairs and lifecycle activities. Establishes maintenance standards and operational runbooks to scale across Stargate campuses.
Location: United States
Workplace: Remote
Employment Type: Full time
Job Function: Field Service, Maintenance & Skilled Trades

Key Responsibilities

  • •Serve as OpenAI’s senior on-site hardware operations lead for server, GPU, storage, and rack-level infrastructure.
  • •Drive technical triage and resolution of complex hardware failures impacting production systems.
  • •Partner with Fleet Health Engineering to investigate recurring hardware issues, identify failure patterns, and improve fleet reliability.
  • •Lead root cause analysis (RCA) efforts for critical hardware incidents and develop corrective and preventive action plans.
  • •Collaborate with Oracle operations teams and OEM vendors to coordinate repairs, replacements, upgrades, and hardware lifecycle activities.
Travel: Low travel

Pay and Benefits

Salary: USD 86,400 - 228,000 annually

Key Requirements

  • •8+ years of experience supporting large-scale datacenter hardware infrastructure in a senior technician, sustaining engineering, or hardware operations leadership role.
  • •Deep expertise with server platforms, GPU systems, storage infrastructure, rack integration, and datacenter hardware architecture.
  • •Strong experience diagnosing complex hardware failures and leading repair efforts in production environments.
  • •Experience conducting root cause analysis and driving long-term corrective actions.
  • •Ability to travel occasionally to support new campus deployments and operational readiness activities.
Experience:8+ yearsDatacenterGPUHardware operationsInfrastructureAI
Skills:CommunicationLeadershipTroubleshootingProblem-solvingCross-functional collaboration
Tech Stack:LinuxHardware monitoringTelemetryRCCAFRACASFMEA

Company Brief

OpenAI
Develops and deploys advanced generative AI models (including ChatGPT and DALL·E) and AI infrastructure, providing APIs and consumer products to accelerate safe AGI for broad benefit.
Industry: AI & Machine Learning
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Scaleup
Valuation: Hectocorn (USD 100B+)
Funding: Series E+
Headquarters: San Francisco, United States
Founded: 2015
Glassdoor
Glassdoor: 4.4
WebsiteLinkedInGlassdoor