Software Engineer, GPU Infrastructure - HPC

OpenAI
San Francisco, New York
Workplace: OnsiteFull timeUSD 230,000 - 490,000 annuallyFunction: Software EngineeringSkills: ["Problem-solving","Collaboration"]

OpenAI’s Fleet HPC software engineer role focuses on ensuring reliability and uptime of the compute fleet across data centers, GPUs, and networking. You’ll build automation for provisioning, monitoring, and lifecycle events; diagnose performance bottlenecks; collaborate with clusters and infrastructure teams; and push automation to scale, enabling safe, high-performance AI research and products.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
OpenAI
OpenAI
1 year ago

Software Engineer, GPU Infrastructure - HPC

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 13 minutes agoStatus: Live

Job Summary

OpenAI’s Fleet HPC software engineer role focuses on ensuring reliability and uptime of the compute fleet across data centers, GPUs, and networking. You’ll build automation for provisioning, monitoring, and lifecycle events; diagnose performance bottlenecks; collaborate with clusters and infrastructure teams; and push automation to scale, enabling safe, high-performance AI research and products.
Location: San Francisco, New York
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering

Key Responsibilities

  • •Build and maintain automation systems for provisioning and managing server fleets.
  • •Develop tools to monitor server health, performance, and lifecycle events.
  • •Collaborate with clusters, networking, and infrastructure teams.
  • •Partner with external operators to ensure a high level of quality.
  • •Identify and fix performance bottlenecks and inefficiencies.

Pay and Benefits

Salary: USD 230,000 - 490,000 annually
Equity and Bonus:Equity
Perks:Equity

Key Requirements

  • •Experience managing large-scale server environments.
  • •Proficiency in Python, Go, or similar languages.
  • •Strong Linux, networking, and server hardware knowledge.
  • •Comfort digging into noisy data with SQL, PromQL, and Pandas or equivalent tools.
  • •Experience building and operationalizing reliable automation for large compute fleets.
Experience:HPCDistributed systemsGPU computing
Skills:Problem-solvingCollaboration
Tech Stack:PythonGoLinuxSQLPrometheusGrafanaPandasPromQLPCIeInfiniband

Company Brief

OpenAI
Develops and deploys advanced generative AI models (including ChatGPT and DALL·E) and AI infrastructure, providing APIs and consumer products to accelerate safe AGI for broad benefit.
Industry: AI & Machine Learning
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Scaleup
Valuation: Hectocorn (USD 100B+)
Funding: Series E+
Headquarters: San Francisco, United States
Founded: 2015
Glassdoor
Glassdoor: 4.4
WebsiteLinkedInGlassdoor