Software Engineer, Platform Systems

OpenAI
San Francisco
Workplace: OnsiteFull timeUSD 310,000 - 460,000 annuallyFunction: Software EngineeringSkills: ["Observability","Debugging","Tracing","Profiling","Distributed systems","Fault tolerance"]

Design and build distributed failure-detection, tracing, and observability systems for large-scale AI training workloads. Build tooling to identify slow or faulty nodes, surface performance bottlenecks, and improve observability and reliability of OpenAI’s training platform. Collaborate with systems, infrastructure, and research teams to evolve platform capabilities and support new training paradigms.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
OpenAI
OpenAI
7 months ago

Software Engineer, Platform Systems

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 1 hour agoStatus: Live

Job Summary

Design and build distributed failure-detection, tracing, and observability systems for large-scale AI training workloads. Build tooling to identify slow or faulty nodes, surface performance bottlenecks, and improve observability and reliability of OpenAI’s training platform. Collaborate with systems, infrastructure, and research teams to evolve platform capabilities and support new training paradigms.
Location: San Francisco
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering

Key Responsibilities

  • •Design and build distributed failure-detection, tracing, and profiling systems for large-scale AI training jobs.
  • •Develop tooling to identify slow, faulty, or misbehaving nodes and provide actionable visibility into system behavior.
  • •Improve observability, reliability, and performance across OpenAI’s training platform.
  • •Debug and resolve issues in complex, high-throughput distributed systems.
  • •Collaborate with systems, infrastructure, and research teams to evolve platform capabilities and extend failure-detection or tracing to support new training paradigms.

Pay and Benefits

Salary: USD 310,000 - 460,000 annually
Equity and Bonus:Equity

Key Requirements

  • •Design and implement distributed failure-detection, tracing, and profiling systems for large-scale AI training jobs; experience with low-level software where system details matter preferred.
  • •Develop tooling to identify slow, faulty, or misbehaving nodes and provide actionable visibility into system behavior.
  • •Strong focus on observability, reliability, and performance across a distributed training platform.
  • •Ability to debug and resolve issues in complex, high-throughput distributed systems; familiarity with performance analysis and debugging at scale.
  • •Experience collaborating across systems, infrastructure, and research teams to evolve platform capabilities and support new training paradigms.
Experience:Distributed systemsHigh-performance computingPlatform infrastructure
Skills:ObservabilityDebuggingTracingProfilingDistributed systemsFault tolerance
Tech Stack:ObservabilityProfilingTracingFault tolerance

Company Brief

OpenAI
Develops and deploys advanced generative AI models (including ChatGPT and DALL·E) and AI infrastructure, providing APIs and consumer products to accelerate safe AGI for broad benefit.
Industry: AI & Machine Learning
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Scaleup
Valuation: Hectocorn (USD 100B+)
Funding: Series E+
Headquarters: San Francisco, United States
Founded: 2015
Glassdoor
Glassdoor: 4.4
WebsiteLinkedInGlassdoor