Software Engineer, Compute Foundations

OpenAI
San Francisco
Workplace: OnsiteFull timeUSD 255,000 - 490,000 annuallyFunction: Software EngineeringSkills: ["Communication","Reliability focus","Problem diagnosis","Cross-team collaboration","Technical tradeoff reasoning"]

Build distributed systems that provision, configure, and manage GPU compute infrastructure across sites. Create Kubernetes-based control planes, controllers, and APIs to coordinate machine and cluster lifecycles, including network boot, firmware/OS images, drivers, and host configuration. Design reliable reconciliation and recovery under concurrency and partial failures, improve control-plane throughput and latency, and integrate new sites and GPU hardware generations into a consistent platform.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
OpenAI
OpenAI
1 day ago

Software Engineer, Compute Foundations

āœ“ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 8 hours agoStatus: Live

Job Summary

Build distributed systems that provision, configure, and manage GPU compute infrastructure across sites. Create Kubernetes-based control planes, controllers, and APIs to coordinate machine and cluster lifecycles, including network boot, firmware/OS images, drivers, and host configuration. Design reliable reconciliation and recovery under concurrency and partial failures, improve control-plane throughput and latency, and integrate new sites and GPU hardware generations into a consistent platform.
Location: San Francisco
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering

Key Responsibilities

  • •Design, build, and operate Kubernetes-based controllers and distributed services to coordinate infrastructure across sites, isolate failures, and scale with GPU capacity.
  • •Define APIs and resource models so clients can request and track lifecycle operations through consistent interfaces across hardware platforms and providers.
  • •Build provisioning and configuration services that coordinate network boot, hardware management interfaces, and deployment of firmware, OS images, drivers, and host configuration.
  • •Develop lifecycle management for discovery, allocation, provisioning, upgrades, maintenance, recovery, and decommissioning, integrating with health and validation systems.
  • •Design reliable reconciliation and recovery for concurrent changes, interrupted operations, and partial failures, including staged rollouts to limit disruption across nodes, racks, and clusters.

Pay and Benefits

Salary: USD 255,000 - 490,000 annually
Equity and Bonus:Equity

Key Requirements

  • •Design, implement, and own production distributed systems or infrastructure services.
  • •Develop infrastructure systems that use Kubernetes APIs and reconciliation to manage resources.
  • •Understand bare-metal node lifecycle from power-on to workload-ready, with depth in one or more areas such as PXE, DHCP/DNS, BMCs, firmware, Linux, drivers, images, or configuration management.
  • •Design reliable APIs and asynchronous workflows, reasoning about concurrency, consistency, idempotency, and failures across service and provider boundaries.
  • •Diagnose reliability and performance problems across service, operating-system, and machine boundaries, turning production evidence into lasting improvements.
Experience:Distributed systemsInfrastructure servicesKubernetesBare-metalGPU or HPC infrastructureData centers
Skills:CommunicationReliability focusProblem diagnosisCross-team collaborationTechnical tradeoff reasoning
Tech Stack:KubernetesKubernetes APIsControllersAPIsDistributed servicesPXEDHCPDNSBaseboard management controllers (BMCs)FirmwareLinuxDriversOperating-system imagesConfiguration managementNetwork boot

Company Brief

OpenAI
Develops and deploys advanced generative AI models (including ChatGPT and DALLĀ·E) and AI infrastructure, providing APIs and consumer products to accelerate safe AGI for broad benefit.
Industry: AI & Machine Learning
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Scaleup
Valuation: Hectocorn (USD 100B+)
Funding: Series E+
Headquarters: San Francisco, United States
Founded: 2015
Glassdoor
Glassdoor: 4.4
WebsiteLinkedInGlassdoor