Training Performance Engineer

OpenAI
San Francisco
Workplace: HybridFull timeUSD 250,000 - 445,000 annuallyFunction: Education & TrainingSkills: ["Python","C++","Rust","CUDA","PyTorch","JAX","TensorFlow","GPU","Multi-GPU","HPC","Profiling","Throughput","Distributed training","Model training","Kernel efficiency","Scheduling","Communication","NCCL","MPI","UCX"]

Join OpenAI as a Training Performance Engineer to optimize large-scale distributed model training. You’ll profile training runs, boost GPU utilization, reduce bottlenecks across compute, communication and storage, and collaborate with runtime and systems teams to push throughput and uptime for frontier-scale models in a hybrid San Francisco, CA setting.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
OpenAI
OpenAI
10 months ago

Training Performance Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 1 hour agoStatus: Live

Job Summary

Join OpenAI as a Training Performance Engineer to optimize large-scale distributed model training. You’ll profile training runs, boost GPU utilization, reduce bottlenecks across compute, communication and storage, and collaborate with runtime and systems teams to push throughput and uptime for frontier-scale models in a hybrid San Francisco, CA setting.
Location: San Francisco
Workplace: Hybrid
Employment Type: Full time
Job Function: Education & Training

Key Responsibilities

  • •Profile end-to-end training runs to identify performance bottlenecks across compute, communication, and storage.
  • •Optimize GPU utilization and throughput for large-scale distributed model training.
  • •Collaborate with runtime and systems engineers to improve kernel efficiency, scheduling, and collective communication performance.
  • •Implement model graph transforms to improve end to end throughput.
  • •Build tooling to monitor and visualize MFU, throughput, and uptime across clusters.
  • •Partner with researchers to ensure new model architectures scale efficiently during pre-training.
  • •Contribute to infrastructure decisions that improve reliability and efficiency of large training jobs.

Pay and Benefits

Salary: USD 250,000 - 445,000 annually
Equity and Bonus:Equity

Key Requirements

  • •Strong programming skills in Python and C++ (Rust or CUDA a plus)
  • •Experience running distributed training jobs on multi-GPU systems or HPC clusters
  • •Familiarity with PyTorch, JAX, or TensorFlow and understanding of large-scale training loops
  • •Ability to profile end-to-end training runs and identify bottlenecks across compute, communication and storage
  • •Experience optimizing GPU utilization and throughput for large-scale distributed model training
Experience:AIMachine LearningDistributed Systems
Skills:PythonC++RustCUDAPyTorchJAXTensorFlowGPUMulti-GPUHPCProfilingThroughputDistributed trainingModel trainingKernel efficiencySchedulingCommunicationNCCLMPIUCX
Tech Stack:PythonC++RustCUDAPyTorchJAXTensorFlowNCCLMPIUCXGPUHPC

Company Brief

OpenAI
Develops and deploys advanced generative AI models (including ChatGPT and DALL·E) and AI infrastructure, providing APIs and consumer products to accelerate safe AGI for broad benefit.
Industry: AI & Machine Learning
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Scaleup
Valuation: Hectocorn (USD 100B+)
Funding: Series E+
Headquarters: San Francisco, United States
Founded: 2015
Glassdoor
Glassdoor: 4.4
WebsiteLinkedInGlassdoor