Software Engineer, Model Runtime
San Francisco
Workplace: HybridFull timeUSD 266,000 - 445,000 annuallyFunction: Software EngineeringSkills: ["Quantitative reasoning","Profiling","Debugging","Cross-team collaboration","Abstraction design"]Build the LLM model runtime inside the inference engine to execute frontier models at scale on OpenAI’s custom silicon. Design a production-grade runtime that handles scheduling, continuous batching, memory and KV-cache management, and distributed execution across chips, hosts, and racks. Optimize latency, throughput, utilization, and reliability, and partner across model, kernel, compiler, and architecture teams while adding profiling, observability, benchmarking, and performance modeling to drive continuous improvements.
Loading
Loading job details...
Preparing the role view and application actions.

