Software Engineer, Compute Foundations Systems

OpenAI
San Francisco
Workplace: OnsiteFull timeUSD 230,000 - 490,000 annuallyFunction: Software EngineeringSkills: ["Problem-solving","Communication","Collaboration","Adaptability","Time management"]

Hardware-software hybrid role operating advanced frontier compute clusters. You’ll scale massive Kubernetes clusters, automate bare-metal bring-up, and create software layers that unify multiple clusters for training workloads. The role blends distributed systems engineering with hands-on data-center infrastructure, focusing on uptime, automation, and end-to-end reliability across servers and network gear in large-scale OpenAI datacenters.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
OpenAI
OpenAI
1 year ago

Software Engineer, Compute Foundations Systems

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 16 minutes agoStatus: Live

Job Summary

Hardware-software hybrid role operating advanced frontier compute clusters. You’ll scale massive Kubernetes clusters, automate bare-metal bring-up, and create software layers that unify multiple clusters for training workloads. The role blends distributed systems engineering with hands-on data-center infrastructure, focusing on uptime, automation, and end-to-end reliability across servers and network gear in large-scale OpenAI datacenters.
Location: San Francisco
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering

Key Responsibilities

  • •Spin up and scale large Kubernetes clusters, including automation for provisioning, bootstrapping, and cluster lifecycle management
  • •Build software abstractions that unify multiple clusters and present a seamless interface to training workloads
  • •Own node bring-up from bare metal through firmware upgrades, ensuring fast, repeatable deployment at massive scale
  • •Improve operational metrics such as reducing cluster restart times and accelerating firmware or OS upgrade cycles
  • •Integrate networking and hardware health systems to deliver end-to-end reliability across servers, switches, and data center infrastructure

Pay and Benefits

Salary: USD 230,000 - 490,000 annually
Equity and Bonus:Equity
Perks:Equity

Key Requirements

  • •Kubernetes experience in large-scale or hyperscale environments
  • •Strong programming/scripting skills (Python, Go, or similar)
  • •Familiarity with Infrastructure-as-Code tools (Terraform or CloudFormation)
  • •Experience with bare-metal Linux environments and GPU hardware
  • •Ability to diagnose and fix issues quickly and build automation to reduce manual work
Experience:High-scale computingCloud computing
Skills:Problem-solvingCommunicationCollaborationAdaptabilityTime management
Tech Stack:KubernetesPythonGoTerraformCloudFormationLinuxGPUNetworking

Company Brief

OpenAI
Develops and deploys advanced generative AI models (including ChatGPT and DALL·E) and AI infrastructure, providing APIs and consumer products to accelerate safe AGI for broad benefit.
Industry: AI & Machine Learning
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Scaleup
Valuation: Hectocorn (USD 100B+)
Funding: Series E+
Headquarters: San Francisco, United States
Founded: 2015
Glassdoor
Glassdoor: 4.4
WebsiteLinkedInGlassdoor