Research Engineer / Scientist - Storage for LLM

ByteDance
San Jose
Workplace: OnsiteFull timeFunction: Research & Scientific (R&D)Education: phdSkills: ["Collaboration","Research","Performance optimization"]

Build and maintain a high-performance KV cache layer for LLM inference, focusing on distributed storage and GPU-aware caching. Design systems that improve latency, throughput, and cost-efficiency by optimizing reuse of transformer attention key-value states and prompt embeddings. Collaborate with inference and serving teams to integrate caching into token streaming, batched decoding, and model parallelism. Evaluate or extend open-source KV stores and monitor performance to iterate on caching algorithms.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
1 month ago

Research Engineer / Scientist - Storage for LLM

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 30 days agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Build and maintain a high-performance KV cache layer for LLM inference, focusing on distributed storage and GPU-aware caching. Design systems that improve latency, throughput, and cost-efficiency by optimizing reuse of transformer attention key-value states and prompt embeddings. Collaborate with inference and serving teams to integrate caching into token streaming, batched decoding, and model parallelism. Evaluate or extend open-source KV stores and monitor performance to iterate on caching algorithms.
Location: San Jose
Workplace: Onsite
Employment Type: Full time
Job Function: Research & Scientific (R&D)
Seniority: Mid level

Key Responsibilities

  • •Design and implement a distributed KV cache system to store and retrieve intermediate states (attention keys/values) across GPUs or nodes.
  • •Optimize low-latency access and eviction policies for long-context inputs, token streams, and reused embeddings.
  • •Integrate the cache with token streaming pipelines, batched decoding, and model parallelism in collaboration with inference/serving teams.
  • •Develop cache consistency and synchronization protocols for multi-tenant, multi-request environments.
  • •Implement memory-aware sharding, eviction (windowed LRU, TTL), and replication strategies; monitor performance and iterate on caching algorithms.

Key Requirements

  • •PhD in Computer Science, Applied Mathematics, Electrical Engineering, or related technical field.
  • •Strong understanding of transformer internals and how KV caching impacts autoregressive decoding.
  • •Experience with distributed systems, memory management, and low-latency serving (RPC, gRPC, CUDA-aware networking).
  • •Familiarity with high-performance compute environments (NVIDIA GPUs, TensorRT, Triton Inference Server).
  • •Proficiency in systems languages such as C++, Rust, Go, or CUDA.
Experience:LLM inferenceDistributed systemsAI systemsGPU computing
Education:PhD / Doctorate
Skills:CollaborationResearchPerformance optimization
Tech Stack:Vector databasesMulti-modal databasesNL2SQLNL2ChartKV cacheDistributed KV cacheC++RustGoCUDARPCGRPCCUDA-aware networkingNVIDIA GPUsTensorRTTriton Inference ServerVLLMSGLangFasterTransformerDeepSpeed

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn