Inference Technical Lead, On-Device Transformers
San Francisco
Workplace: HybridFull timeUSD 445,000 - 445,000 annuallyFunction: Administration & Executive AssistanceSkills: ["GPU","CUDA","Kernels","Compilers","ML runtimes","Transformers","Edge deployment","Inference"]Technical Lead role in OpenAI’s Future of Computing Research team, driving on-device/edge transformer deployments by selecting silicon platforms, co-designing model architectures for latency and memory constraints, and leading low-level inference stack development. Hybrid work in San Francisco with relocation assistance, collaborating with ML researchers and hardware teams to push frontier capabilities.
Loading
Loading job details...
Preparing the role view and application actions.

