Senior Research Engineer / Scientist - Storage for LLM
ByteDance
Seattle
Workplace: OnsiteFull timeFunction: Research & Scientific (R&D)Education: phdSkills: ["Collaboration","Innovation","Systems thinking"]Design and maintain a high-performance KV cache layer for LLM inference, improving latency, throughput, and cost by optimizing reuse of attention key/value states and prompt embeddings. Build a distributed caching system across GPUs/nodes with low-latency access, eviction, consistency, and synchronization for multi-tenant workloads. Collaborate with inference and serving teams to integrate with token streaming, batched decoding, and model parallelism.

