Research Engineer / Scientist - Storage for LLM
ByteDance
San Jose
Workplace: OnsiteFull timeFunction: Research & Scientific (R&D)Education: phdSkills: ["Collaboration","Research","Performance optimization"]Build and maintain a high-performance KV cache layer for LLM inference, focusing on distributed storage and GPU-aware caching. Design systems that improve latency, throughput, and cost-efficiency by optimizing reuse of transformer attention key-value states and prompt embeddings. Collaborate with inference and serving teams to integrate caching into token streaming, batched decoding, and model parallelism. Evaluate or extend open-source KV stores and monitor performance to iterate on caching algorithms.

