Staff Applied AI Inference Engineer
San Francisco
Workplace: OnsiteFull timeUSD 215,000 - 260,000 annuallyFunction: Data Science & Machine LearningSkills: ["Communication","Problem-solving","Ownership","Urgency"]Own the end-to-end AI inference stack to make large language models faster, cheaper, and more reliable in production. Profile latency and cost drivers, apply modern optimization techniques, and dive into the serving stack down to CUDA kernels. Partner with customer engineering teams to tailor deployments, move workloads from proof of concept to monitored services, and deliver measurable performance gains using Python-first development.
Loading
Loading job details...
Preparing the role view and application actions.

