AI Infrastructure Engineer, pAGI

OpenAI
San Francisco
Workplace: HybridFull timeUSD 266,000 - 500,000 annuallyFunction: DevOps, Cloud & InfrastructureSkills: ["Ownership","Debugging","Collaboration","Self-motivation","Measurement-driven improvement"]

Build and operate the infrastructure that powers large-scale model training and evaluation, improving reliability, throughput, and GPU efficiency. Develop shared inference and grading platforms with health monitoring, capacity management, and performance visibility. Enhance compute scheduling and resource allocation to reduce idle GPU time and speed recovery from failures. Diagnose bottlenecks across the training, inference, and orchestration stack, and deliver self-service, observability, and automated validation for researchers.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
OpenAI
OpenAI
1 day ago

AI Infrastructure Engineer, pAGI

āœ“ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 15 hours agoStatus: Live

Job Summary

Build and operate the infrastructure that powers large-scale model training and evaluation, improving reliability, throughput, and GPU efficiency. Develop shared inference and grading platforms with health monitoring, capacity management, and performance visibility. Enhance compute scheduling and resource allocation to reduce idle GPU time and speed recovery from failures. Diagnose bottlenecks across the training, inference, and orchestration stack, and deliver self-service, observability, and automated validation for researchers.
Location: San Francisco
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure

Key Responsibilities

  • •Build and operate infrastructure for large-scale training and evaluation to improve reliability, throughput, and resource efficiency.
  • •Develop shared inference and grading platforms with automated capacity management, health monitoring, and performance visibility.
  • •Improve compute scheduling and resource allocation to reduce idle GPU time and speed workload recovery from failures.
  • •Diagnose bottlenecks across training, inference, and orchestration and coordinate with teams to improve end-to-end performance.
  • •Build self-service tools, automated validation, and observability to help researchers launch experiments and troubleshoot with less manual effort.

Pay and Benefits

Salary: USD 266,000 - 500,000 annually
Equity and Bonus:Equity
Perks:Equity

Key Requirements

  • •Strong software engineering fundamentals and experience building or operating large-scale distributed systems.
  • •Experience in ML infrastructure, inference systems, GPU performance, or infrastructure tooling.
  • •Ability to diagnose bottlenecks across training, inference, and orchestration and improve end-to-end performance.
  • •Comfort taking ownership of open-ended problems in collaboration with researchers and engineering teams.
  • •Use measurements to guide improvements in performance, reliability, and capacity management.
Experience:ML infrastructureDistributed systems
Skills:OwnershipDebuggingCollaborationSelf-motivationMeasurement-driven improvement
Tech Stack:Distributed systemsInferenceGradingML infrastructureGPU performanceCompute schedulingResource allocationHealth monitoringObservabilityAutomationCapacity managementOrchestration

Company Brief

OpenAI
Develops and deploys advanced generative AI models (including ChatGPT and DALLĀ·E) and AI infrastructure, providing APIs and consumer products to accelerate safe AGI for broad benefit.
Industry: AI & Machine Learning
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Scaleup
Valuation: Hectocorn (USD 100B+)
Funding: Series E+
Headquarters: San Francisco, United States
Founded: 2015
Glassdoor
Glassdoor: 4.4
WebsiteLinkedInGlassdoor