Software Engineer, Model Deployment- ChatGPT Engineering

OpenAI
London
Workplace: HybridFull timeFunction: Software EngineeringExperience: 5+ yearsSkills: ["Debugging","Systems design","Operational problem-solving","Communication","Cross-team collaboration"]

Design and operate software that manages large-scale GPU clusters powering ChatGPT inference. Build internal platforms, tooling, and AI-powered agents to automate fleet operations, reduce overhead, and improve developer productivity. Drive reliability and observability across thousands of GPUs by building systems for capacity planning, scheduling, fleet health monitoring, and incident response. Collaborate closely with research, platform, networking, and systems teams to improve compute utilization at the frontier of AI.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
OpenAI
OpenAI
2 months ago

Software Engineer, Model Deployment- ChatGPT Engineering

āœ“ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 4 hours agoStatus: Live

Job Summary

Design and operate software that manages large-scale GPU clusters powering ChatGPT inference. Build internal platforms, tooling, and AI-powered agents to automate fleet operations, reduce overhead, and improve developer productivity. Drive reliability and observability across thousands of GPUs by building systems for capacity planning, scheduling, fleet health monitoring, and incident response. Collaborate closely with research, platform, networking, and systems teams to improve compute utilization at the frontier of AI.
Location: London
Workplace: Hybrid
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Design, build, and operate software that manages large-scale GPU infrastructure supporting ChatGPT inference.
  • •Build internal platforms, tooling, and AI-powered agents to automate fleet operations and reduce operational overhead.
  • •Improve observability, reliability, and operational efficiency across thousands of GPUs.
  • •Develop systems for capacity planning, scheduling, fleet health monitoring, and incident response.
  • •Partner with infrastructure, research, platform, networking, and systems teams to continuously improve the compute platform and establish operational best practices.

Key Requirements

  • •5+ years of software engineering experience building production infrastructure.
  • •Strong programming skills in Go, Python, C++, Rust, or similar systems languages.
  • •Experience designing and operating highly available distributed systems.
  • •Experience with GPU infrastructure, high-performance computing, ML infrastructure, or large-scale compute platforms.
  • •Experience with Kubernetes, cloud infrastructure, Linux, networking, and observability tooling.
Experience:5+ yearsAI infrastructureGPU infrastructureHigh-performance computingML infrastructureDistributed systems
Skills:DebuggingSystems designOperational problem-solvingCommunicationCross-team collaboration
Tech Stack:GoPythonC++RustKubernetesLinuxContainer orchestrationCloud infrastructureNetworkingObservability toolingGPU clustersDistributed infrastructureCapacity planningIncident response

Company Brief

OpenAI
Develops and deploys advanced generative AI models (including ChatGPT and DALLĀ·E) and AI infrastructure, providing APIs and consumer products to accelerate safe AGI for broad benefit.
Industry: AI & Machine Learning
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Scaleup
Valuation: Hectocorn (USD 100B+)
Funding: Series E+
Headquarters: San Francisco, United States
Founded: 2015
Glassdoor
Glassdoor: 4.4
WebsiteLinkedInGlassdoor