Software Engineer, Fleet Infrastructure

OpenAI
San Francisco, New York
Workplace: HybridFull timeUSD 230,000 - 490,000 annuallyFunction: Software EngineeringSkills: ["Problem-solving","Collaboration","Communication"]

Design, build, and operate infrastructure for OpenAI's GPU fleet enabling model training and deployment; focus on scheduling, cluster management, snapshot delivery, and CI/CD across hyperscale compute. Collaborate with researchers and product teams to match workloads, while ensuring high utilization and reliability. This hybrid SF role offers relocation assistance and works on one of the world’s largest GPU fleets.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
OpenAI
OpenAI
1 year ago

Software Engineer, Fleet Infrastructure

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 16 minutes agoStatus: Live

Job Summary

Design, build, and operate infrastructure for OpenAI's GPU fleet enabling model training and deployment; focus on scheduling, cluster management, snapshot delivery, and CI/CD across hyperscale compute. Collaborate with researchers and product teams to match workloads, while ensuring high utilization and reliability. This hybrid SF role offers relocation assistance and works on one of the world’s largest GPU fleets.
Location: San Francisco, New York
Workplace: Hybrid
Employment Type: Full time
Job Function: Software Engineering

Key Responsibilities

  • •Design, implement and operate components of our compute fleet including job scheduling, cluster management, snapshot delivery, and CI/CD systems.
  • •Interface with researchers and product teams to understand workload requirements.
  • •Collaborate with hardware, infrastructure, and business teams to provide a high utilization and high reliability service.
  • •Drive scalability and reliability improvements across the GPU fleet infrastructure.
  • •Partner with cross-functional teams to ensure deployments meet performance and safety requirements.

Pay and Benefits

Salary: USD 230,000 - 490,000 annually
Equity and Bonus:Equity
Perks:RelocationEquity

Key Requirements

  • •Experience with hyperscale compute systems
  • •Strong programming skills
  • •Experience working in public clouds, especially Azure
  • •Experience with Kubernetes
  • •Execution-focused mentality with a rigorous focus on user requirements
Experience:AI/MLGPUCloudKubernetes
Skills:Problem-solvingCollaborationCommunication
Tech Stack:KubernetesAzureCI/CDGPUBlob storage

Company Brief

OpenAI
Develops and deploys advanced generative AI models (including ChatGPT and DALL·E) and AI infrastructure, providing APIs and consumer products to accelerate safe AGI for broad benefit.
Industry: AI & Machine Learning
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Scaleup
Valuation: Hectocorn (USD 100B+)
Funding: Series E+
Headquarters: San Francisco, United States
Founded: 2015
Glassdoor
Glassdoor: 4.4
WebsiteLinkedInGlassdoor