Software Engineer, Workload Enablement

OpenAI
San Francisco
Workplace: HybridFull timeUSD 293,000 - 455,000 annuallyFunction: Software EngineeringExperience: 5+ yearsEducation: bachelorsSkills: ["Communication","Problem-solving","Collaboration","Teamwork"]

OpenAI is hiring a Software Engineer to enable production workloads and end-to-end testing on new platforms. You will create test harnesses, port inference/training workloads to early-access systems, analyze performance and bottlenecks, and characterize end-to-end system behavior across compute, communications, storage, and the control plane. You’ll work with cross-functional teams to ensure stable, scalable platforms.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
OpenAI
OpenAI
4 months ago

Software Engineer, Workload Enablement

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 1 hour agoStatus: Live

Job Summary

OpenAI is hiring a Software Engineer to enable production workloads and end-to-end testing on new platforms. You will create test harnesses, port inference/training workloads to early-access systems, analyze performance and bottlenecks, and characterize end-to-end system behavior across compute, communications, storage, and the control plane. You’ll work with cross-functional teams to ensure stable, scalable platforms.
Location: San Francisco
Workplace: Hybrid
Employment Type: Full time
Job Function: Software Engineering

Key Responsibilities

  • •Port and validate key inference and training workloads on new platforms/SKUs as they arrive; drive correctness, performance, and stability to an internal readiness bar.
  • •Build a suite of benchmarks and stress tests that capture real end-to-end behavior of our workloads across CPU, GPU, memory, storage, network, thermals, and more.
  • •Deep-dive performance on distributed training/inference, including collective performance and tuning and overlap of compute/communication.
  • •Create repeatable test harnesses that run in CI/lab environments and produce actionable outputs (pass/fail, performance score, regression detection).
  • •Partner with systems and fleet bring-up engineers to ensure the platform is stable, performant, usable, and scalable (containerization, Kubernetes integration, telemetry hooks, failure triage loops)
  • •Work cross-functionally with vendors and internal stakeholders by producing clear bug reports, minimal repros, and prioritized issue lists.

Pay and Benefits

Salary: USD 293,000 - 455,000 annually
Equity and Bonus:Equity

Key Requirements

  • •BS in CS/EE (or equivalent practical experience)
  • •5+ years in ML systems, performance engineering, distributed systems, or HPC
  • •Strong hands-on experience with PyTorch and modern LLM training/inference stacks; large-scale distributed training concepts
  • •Proficiency in Python plus comfort reading/writing performance-critical code (C++/CUDA/HIP is a plus)
  • •Strong profiling/debugging skills (Nsight, rocprof, perf, flamegraphs) and experience with RDMA and debugging/optimizing comms libraries (NCCL or RCCL)
Experience:5+ yearsAIMachine learningDistributed systems
Education:Bachelor's
Skills:CommunicationProblem-solvingCollaborationTeamwork
Tech Stack:PyTorchLLM trainingNCCLRCCLRDMACUDAHIPPythonC++KubernetesNsightRocprofPerfFlamegraphs

Company Brief

OpenAI
Develops and deploys advanced generative AI models (including ChatGPT and DALL·E) and AI infrastructure, providing APIs and consumer products to accelerate safe AGI for broad benefit.
Industry: AI & Machine Learning
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Scaleup
Valuation: Hectocorn (USD 100B+)
Funding: Series E+
Headquarters: San Francisco, United States
Founded: 2015
Glassdoor
Glassdoor: 4.4
WebsiteLinkedInGlassdoor