Software Engineer, Model Inference

OpenAI
San Francisco
Workplace: OnsiteFull timeUSD 325,000 - 490,000 annuallyFunction: Software EngineeringExperience: 5+ yearsSkills: ["Problem-solving","Collaboration","Ownership","Communication","Teamwork"]

Join OpenAI’s Inference team to optimize world-leading AI models for high-volume, low-latency production and research environments. You will collaborate with ML researchers and engineers to push the latest technologies into production, improve inference performance, and build tools to diagnose bottlenecks while maximizing GPU utilization on Azure. Expect ownership of end-to-end problems and contributions to scalable distributed systems.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
OpenAI
OpenAI
1 year ago

Software Engineer, Model Inference

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 11 minutes agoStatus: Live

Job Summary

Join OpenAI’s Inference team to optimize world-leading AI models for high-volume, low-latency production and research environments. You will collaborate with ML researchers and engineers to push the latest technologies into production, improve inference performance, and build tools to diagnose bottlenecks while maximizing GPU utilization on Azure. Expect ownership of end-to-end problems and contributions to scalable distributed systems.
Location: San Francisco
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering

Key Responsibilities

  • •Collaborate with ML researchers, engineers, and product managers to bring latest technologies into production.
  • •Enable advanced research through engineering efforts and tooling.
  • •Introduce new techniques, tools, and architecture to improve performance, latency, throughput, and efficiency of the model inference stack.
  • •Build visibility tools to identify bottlenecks and instability and implement high-priority fixes.
  • •Optimize code and Azure VM fleet to utilize GPU RAM and compute resources efficiently.

Pay and Benefits

Salary: USD 325,000 - 490,000 annually
Equity and Bonus:Equity

Key Requirements

  • •5+ years of professional software engineering experience.
  • •Familiarity with PyTorch, NVIDIA GPUs and the tools that optimize them (e.g. NCCL, CUDA), plus HPC tech such as InfiniBand, MPI, NVLink.
  • •Experience architecting, building, observing, and debugging production distributed systems.
  • •Willingness to own problems end-to-end and quickly learn needed knowledge.
  • •Ability to optimize model inference performance for high-volume, low-latency environments.
Experience:5+ yearsAIMachine learningInferenceDistributed systems
Skills:Problem-solvingCollaborationOwnershipCommunicationTeamwork
Tech Stack:PyTorchCUDANVIDIANCCLMPIInfiniBandNVLinkAzureGPUGPU RAMFLOP

Company Brief

OpenAI
Develops and deploys advanced generative AI models (including ChatGPT and DALL·E) and AI infrastructure, providing APIs and consumer products to accelerate safe AGI for broad benefit.
Industry: AI & Machine Learning
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Scaleup
Valuation: Hectocorn (USD 100B+)
Funding: Series E+
Headquarters: San Francisco, United States
Founded: 2015
Glassdoor
Glassdoor: 4.4
WebsiteLinkedInGlassdoor