Software Engineer, Fleet Management

OpenAI
San Francisco, New York
Workplace: HybridFull timeUSD 230,000 - 490,000 annuallyFunction: Software EngineeringSkills: ["Problem-solving","Collaboration","Automation","Reliability","Continuous improvement"]

Software Engineer for OpenAI’s Fleet team focused on building and operating large-scale infrastructure to manage hardware fleets. You’ll design systems spanning cloud and bare-metal environments, integrate hardware metrics with scheduling, and automate infrastructure to improve reliability and efficiency in a hybrid San Francisco-based role.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
OpenAI
OpenAI
9 months ago

Software Engineer, Fleet Management

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 1 hour agoStatus: Live

Job Summary

Software Engineer for OpenAI’s Fleet team focused on building and operating large-scale infrastructure to manage hardware fleets. You’ll design systems spanning cloud and bare-metal environments, integrate hardware metrics with scheduling, and automate infrastructure to improve reliability and efficiency in a hybrid San Francisco-based role.
Location: San Francisco, New York
Workplace: Hybrid
Employment Type: Full time
Job Function: Software Engineering

Key Responsibilities

  • •Design and build systems to manage both cloud and bare-metal fleets at scale.
  • •Develop tools that integrate low-level hardware metrics with high-level job scheduling and cluster management algorithms.
  • •Leverage LLMs to coordinate vendor operations and optimize infrastructure workflows.
  • •Automate infrastructure processes, reducing repetitive toil and improving system reliability.
  • •Collaborate with hardware, infrastructure, and research teams to ensure seamless integration across the stack.

Pay and Benefits

Salary: USD 230,000 - 490,000 annually
Equity and Bonus:Equity

Key Requirements

  • •Strong software engineering skills with experience in large-scale infrastructure environments.
  • •Broad knowledge of cluster-level systems (e.g., Kubernetes, CI/CD pipelines, Terraform, cloud providers).
  • •Deep expertise in server-level systems (e.g., systemd, containerization, Chef, Linux kernels, firmware management, host routing).
  • •Experience designing and building systems to manage both cloud and bare-metal fleets at scale.
  • •Passion for automation, reliability, and continuous improvement of infrastructure.
Experience:Large-scale infrastructureDistributed systemsCloud computing
Skills:Problem-solvingCollaborationAutomationReliabilityContinuous improvement
Tech Stack:KubernetesCI/CDTerraformCloudSystemdChefLinuxContainersLinux kernelsFirmwareHost routing

Company Brief

OpenAI
Develops and deploys advanced generative AI models (including ChatGPT and DALL·E) and AI infrastructure, providing APIs and consumer products to accelerate safe AGI for broad benefit.
Industry: AI & Machine Learning
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Scaleup
Valuation: Hectocorn (USD 100B+)
Funding: Series E+
Headquarters: San Francisco, United States
Founded: 2015
Glassdoor
Glassdoor: 4.4
WebsiteLinkedInGlassdoor