Research Engineer - LLM/VLM Inference Optimization (Seed Infra)
San Jose
Workplace: OnsiteFull timeFunction: Research & Scientific (R&D)Education: bachelorsSkills: ["Performance analysis","Optimization","Profiling","Collaboration","Systems programming"]Build and optimize high-performance inference systems for large-scale LLMs and VLMs. Work across inference engines, serving frameworks, and end-to-end deployment pipelines, applying compiler-level optimizations, parallel computing, graph fusion, CUDA kernel development, low-precision computation, and streaming/speculative decoding. Collaborate with research teams to diagnose bottlenecks and improve latency, throughput, and serving cost while contributing to model toolchains and the technical ecosystem.
Loading
Loading job details...
Preparing the role view and application actions.

