Research Engineer - LLM/VLM Inference Optimization (Seed Infra)
ByteDance
San Jose
Workplace: OnsiteFull timeFunction: Research & Scientific (R&D)Education: bachelorsSkills: ["Performance analysis","Optimization","Profiling","Collaboration","Systems programming"]Build and optimize high-performance inference systems for large-scale LLMs and VLMs. Work across inference engines, serving frameworks, and end-to-end deployment pipelines, applying compiler-level optimizations, parallel computing, graph fusion, CUDA kernel development, low-precision computation, and streaming/speculative decoding. Collaborate with research teams to diagnose bottlenecks and improve latency, throughput, and serving cost while contributing to model toolchains and the technical ecosystem.

