Research Engineer - LLM/VLM Inference Optimization (Seed Infra)
ByteDance
Seattle
Workplace: OnsiteFull timeFunction: Research & Scientific (R&D)Education: bachelorsSkills: ["Performance optimization","Performance analysis","Debugging","Collaboration"]Build and optimize high-performance inference systems for large-scale LLMs and VLMs, including inference engines, serving frameworks, and end-to-end deployment pipelines. Apply compiler-level optimizations, parallel computing, graph fusion, low-precision computation, streaming inference, speculative decoding, and high-concurrency request tuning. Work with research teams to analyze bottlenecks and improve production serving performance, toolchains, and the broader inference ecosystem.

