Research Engineer - LLM/VLM Inference Optimization (Seed Infra)
Seattle
Workplace: OnsiteFull timeFunction: Research & Scientific (R&D)Education: bachelorsSkills: ["Performance optimization","Performance analysis","Debugging","Collaboration"]Build and optimize high-performance inference systems for large-scale LLMs and VLMs, including inference engines, serving frameworks, and end-to-end deployment pipelines. Apply compiler-level optimizations, parallel computing, graph fusion, low-precision computation, streaming inference, speculative decoding, and high-concurrency request tuning. Work with research teams to analyze bottlenecks and improve production serving performance, toolchains, and the broader inference ecosystem.
Loading
Loading job details...
Preparing the role view and application actions.

