Inference Technical Lead, On-Device Transformers

OpenAI
San Francisco
Workplace: HybridFull timeUSD 445,000 - 445,000 annuallyFunction: Administration & Executive AssistanceSkills: ["GPU","CUDA","Kernels","Compilers","ML runtimes","Transformers","Edge deployment","Inference"]

Technical Lead role in OpenAI’s Future of Computing Research team, driving on-device/edge transformer deployments by selecting silicon platforms, co-designing model architectures for latency and memory constraints, and leading low-level inference stack development. Hybrid work in San Francisco with relocation assistance, collaborating with ML researchers and hardware teams to push frontier capabilities.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
OpenAI
OpenAI
4 months ago

Inference Technical Lead, On-Device Transformers

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 21 hours agoStatus: Live

Job Summary

Technical Lead role in OpenAI’s Future of Computing Research team, driving on-device/edge transformer deployments by selecting silicon platforms, co-designing model architectures for latency and memory constraints, and leading low-level inference stack development. Hybrid work in San Francisco with relocation assistance, collaborating with ML researchers and hardware teams to push frontier capabilities.
Location: San Francisco
Workplace: Hybrid
Employment Type: Full time
Job Function: Administration & Executive Assistance
Seniority: Manager level

Key Responsibilities

  • •Evaluate and select silicon platforms (GPUs, NPUs, and specialized accelerators) for on-device and edge deployment of OpenAI models.
  • •Work closely with research teams to co-design model architectures that meet real-world deployment constraints such as latency, memory, power, and bandwidth.
  • •Analyze and model system performance, identifying tradeoffs between model design, memory hierarchy, compute throughput, and hardware capabilities.
  • •Partner with hardware vendors and internal infrastructure teams to bring up new accelerators and ensure efficient execution of transformer workloads.
  • •Build and lead a team of engineers responsible for implementing the low-level inference stack, including kernel development and runtime systems.

Pay and Benefits

Salary: USD 445,000 annually
Equity and Bonus:Equity

Key Requirements

  • •Experience evaluating or deploying workloads on GPUs, NPUs, or other specialized accelerators.
  • •Understanding of transformer model performance, including attention, KV-cache behavior, and memory bandwidth requirements.
  • •Experience designing or optimizing high-performance compute systems, such as inference engines, distributed runtimes, or hardware-aware ML pipelines.
  • •Experience building or leading teams working on low-level performance-critical software such as CUDA kernels, compilers, or ML runtimes.
  • •Experience turning nascent research capabilities into deployable capabilities.
Skills:GPUCUDAKernelsCompilersML runtimesTransformersEdge deploymentInference
Tech Stack:GPUsNPUsCUDAKernelsCompilersML runtimesTransformersEdge deploymentInference

Company Brief

OpenAI
Develops and deploys advanced generative AI models (including ChatGPT and DALL·E) and AI infrastructure, providing APIs and consumer products to accelerate safe AGI for broad benefit.
Industry: AI & Machine Learning
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Scaleup
Valuation: Hectocorn (USD 100B+)
Funding: Series E+
Headquarters: San Francisco, United States
Founded: 2015
Glassdoor
Glassdoor: 4.4
WebsiteLinkedInGlassdoor