Software Engineer, Model Runtime

OpenAI
San Francisco
Workplace: HybridFull timeUSD 266,000 - 445,000 annuallyFunction: Software EngineeringSkills: ["Quantitative reasoning","Profiling","Debugging","Cross-team collaboration","Abstraction design"]

Build the LLM model runtime inside the inference engine to execute frontier models at scale on OpenAI’s custom silicon. Design a production-grade runtime that handles scheduling, continuous batching, memory and KV-cache management, and distributed execution across chips, hosts, and racks. Optimize latency, throughput, utilization, and reliability, and partner across model, kernel, compiler, and architecture teams while adding profiling, observability, benchmarking, and performance modeling to drive continuous improvements.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
OpenAI
OpenAI
1 day ago

Software Engineer, Model Runtime

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 4 hours agoStatus: Live

Job Summary

Build the LLM model runtime inside the inference engine to execute frontier models at scale on OpenAI’s custom silicon. Design a production-grade runtime that handles scheduling, continuous batching, memory and KV-cache management, and distributed execution across chips, hosts, and racks. Optimize latency, throughput, utilization, and reliability, and partner across model, kernel, compiler, and architecture teams while adding profiling, observability, benchmarking, and performance modeling to drive continuous improvements.
Location: San Francisco
Workplace: Hybrid
Employment Type: Full time
Job Function: Software Engineering

Key Responsibilities

  • •Design and implement the LLM inference runtime for frontier models running on custom silicon.
  • •Build scheduling, continuous batching, memory management, KV-cache management, and execution orchestration for high-performance inference.
  • •Develop distributed execution strategies across chips, hosts, and racks, including model partitioning, communication, and synchronization.
  • •Optimize end-to-end latency, throughput, memory efficiency, and hardware utilization across model architectures and serving workloads.
  • •Create profiling, observability, benchmarking, and performance-modeling tools, and debug correctness, performance, and reliability issues across the full stack.

Pay and Benefits

Salary: USD 266,000 - 445,000 annually
Equity and Bonus:Equity

Key Requirements

  • •Strong systems programming experience in C++, Rust, Python, or similar performance-oriented environments.
  • •Experience building or optimizing runtimes, distributed systems, compilers, kernels, or model-serving infrastructure.
  • •Understanding of modern LLM inference (prefill/decode, batching, KV-cache tradeoffs, and model parallelism).
  • •Ability to reason quantitatively about latency, throughput, compute intensity, memory bandwidth, communication, and utilization.
  • •Comfort profiling and debugging performance across multiple layers of a hardware-software stack.
Experience:AILLM inferenceDistributed systemsModel servingCompilersKernels
Skills:Quantitative reasoningProfilingDebuggingCross-team collaborationAbstraction design
Tech Stack:C++RustPythonLLM inferencePrefillDecodeContinuous batchingKV-cacheDistributed executionModel partitioningCompilersKernelsSchedulingMemory managementProfilingObservabilityBenchmarkingVLLMSGLang

Company Brief

OpenAI
Develops and deploys advanced generative AI models (including ChatGPT and DALL·E) and AI infrastructure, providing APIs and consumer products to accelerate safe AGI for broad benefit.
Industry: AI & Machine Learning
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Scaleup
Valuation: Hectocorn (USD 100B+)
Funding: Series E+
Headquarters: San Francisco, United States
Founded: 2015
Glassdoor
Glassdoor: 4.4
WebsiteLinkedInGlassdoor