Member of Technical Staff - Inference Research
Modal
New York, San Francisco
Workplace: OnsiteFull timeUSD 150,000 - 350,000 annuallyFunction: Research & Scientific (R&D)Skills: ["Independent ownership","Collaboration","Research execution"]Build hands-on inference research to improve cost per token and tail latency for real customer LLM workloads. Own end-to-end research bets spanning speculative decoding, disaggregated prefill/decode, quantization, KV-cache and memory management, and autoscaling for spiky serverless traffic. Train and iterate speculators from production traffic, partner with customers and Forward Deployed Engineers to deploy and tune models, and collaborate with external labs to turn frontier serving techniques into products.

