Research Engineer / Scientist - Storage for LLM
ByteDance
Seattle
Workplace: OnsiteFull timeFunction: Research & Scientific (R&D)Education: phdSkills: ["Collaboration","Systems thinking","Problem-solving","Low-latency optimization"]Build and maintain a high-performance KV cache layer for large language model (LLM) inference, improving latency, throughput, and cost efficiency. Design distributed storage and caching across GPUs/nodes, optimize eviction and low-latency access for long-context inputs, and integrate cache reuse into token streaming, batched decoding, and model parallelism. Develop cache consistency protocols and evaluate/extend open-source GPU-aware KV caching solutions.

