Senior Research Engineer / Scientist - Storage for LLM

ByteDance
Seattle
Workplace: OnsiteFull timeFunction: Research & Scientific (R&D)Education: phdSkills: ["Collaboration","Innovation","Systems thinking"]

Design and maintain a high-performance KV cache layer for LLM inference, improving latency, throughput, and cost by optimizing reuse of attention key/value states and prompt embeddings. Build a distributed caching system across GPUs/nodes with low-latency access, eviction, consistency, and synchronization for multi-tenant workloads. Collaborate with inference and serving teams to integrate with token streaming, batched decoding, and model parallelism.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
1 month ago

Senior Research Engineer / Scientist - Storage for LLM

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 30 days agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Design and maintain a high-performance KV cache layer for LLM inference, improving latency, throughput, and cost by optimizing reuse of attention key/value states and prompt embeddings. Build a distributed caching system across GPUs/nodes with low-latency access, eviction, consistency, and synchronization for multi-tenant workloads. Collaborate with inference and serving teams to integrate with token streaming, batched decoding, and model parallelism.
Location: Seattle
Workplace: Onsite
Employment Type: Full time
Job Function: Research & Scientific (R&D)
Seniority: Mid level

Key Responsibilities

  • •Design and implement a distributed KV cache system to store and retrieve intermediate transformer states across GPUs or nodes.
  • •Optimize low-latency access and eviction policies for long-context inputs, token streams, and reused embeddings.
  • •Integrate the cache with token streaming pipelines, batched decoding, and model parallelism in collaboration with inference/serving teams.
  • •Develop cache consistency and synchronization protocols for multi-tenant, multi-request environments.
  • •Implement memory-aware sharding, eviction (windowed LRU, TTL), and replication strategies and iterate on caching algorithms to reduce inference compute costs and response time.

Key Requirements

  • •PhD in Computer Science, Applied Mathematics, Electrical Engineering, or a related technical field.
  • •Strong understanding of transformer internals and how KV caching impacts autoregressive decoding.
  • •Experience with distributed systems, memory management, and low-latency serving (RPC, gRPC, CUDA-aware networking).
  • •Familiarity with high-performance compute environments (NVIDIA GPUs, TensorRT, Triton Inference Server).
  • •Proficiency in systems-level languages such as C++, Rust, Go, or CUDA.
  • •Prior experience building inference-serving systems for LLMs (e.g., vLLM, SGLang, FasterTransformer, DeepSpeed, Hugging Face Text Generation Inference).
Experience:AILarge language modelsDistributed systemsLLM inferenceHigh-performance computing
Education:PhD / Doctorate
Skills:CollaborationInnovationSystems thinking
Tech Stack:C++RustGoCUDARPCGRPCNVIDIA GPUsTensorRTTriton Inference ServerVLLMSGLangFasterTransformerDeepSpeedHugging Face Text Generation InferenceHBMNUMANVLinkNCCLGDRGDS

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn