Research Engineer - LLM/VLM Inference Optimization (Seed Infra)

ByteDance
San Jose
Workplace: OnsiteFull timeFunction: Research & Scientific (R&D)Education: bachelorsSkills: ["Performance analysis","Optimization","Profiling","Collaboration","Systems programming"]

Build and optimize high-performance inference systems for large-scale LLMs and VLMs. Work across inference engines, serving frameworks, and end-to-end deployment pipelines, applying compiler-level optimizations, parallel computing, graph fusion, CUDA kernel development, low-precision computation, and streaming/speculative decoding. Collaborate with research teams to diagnose bottlenecks and improve latency, throughput, and serving cost while contributing to model toolchains and the technical ecosystem.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
1 month ago

Research Engineer - LLM/VLM Inference Optimization (Seed Infra)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 30 days agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Build and optimize high-performance inference systems for large-scale LLMs and VLMs. Work across inference engines, serving frameworks, and end-to-end deployment pipelines, applying compiler-level optimizations, parallel computing, graph fusion, CUDA kernel development, low-precision computation, and streaming/speculative decoding. Collaborate with research teams to diagnose bottlenecks and improve latency, throughput, and serving cost while contributing to model toolchains and the technical ecosystem.
Location: San Jose
Workplace: Onsite
Employment Type: Full time
Job Function: Research & Scientific (R&D)
Seniority: Mid level

Key Responsibilities

  • •Design, develop, and optimize high-performance inference systems for large-scale LLMs and VLMs, including inference engines, serving frameworks, and deployment pipelines.
  • •Develop state-of-the-art model inference engines using compiler-level optimizations, parallel computing, graph fusion, efficient CUDA kernel development, low-precision computation, streaming inference, speculative decoding, and high-concurrency request optimization.
  • •Collaborate with other research teams to identify performance bottlenecks, conduct in-depth performance analysis, and optimize large models.
  • •Contribute to the development of model toolchains and the broader technical ecosystem.

Key Requirements

  • •Bachelor’s degree or above in Computer Science, Electrical Engineering, Software Engineering, or a related field.
  • •Strong proficiency in C/C++ and Python with solid algorithms, data structures, and systems programming; familiarity with containerization and server-side debugging.
  • •Hands-on experience with at least one mainstream ML framework (e.g., PyTorch or TensorFlow).
  • •Experience deploying or optimizing LLM/VLM inference at production scale with measurable impact on latency, throughput, or serving cost.
  • •Familiarity with GPU architecture and experience optimizing compute-intensive operators (e.g., FlashAttention, GEMM, GEMV, Conv2D).
Experience:AI foundation modelsLLMVLMDistributed inferenceProduction ML
Education:Bachelor's
Skills:Performance analysisOptimizationProfilingCollaborationSystems programming
Tech Stack:C/C++PythonPyTorchTensorFlowCUDACUDA/OpenCLFlashAttentionGEMMGEMVConv2DContainerizationCUDA kernelTensorRTTritonCUTLASSOpenCLCPU/GPU architecturesModel toolchainsServer-side debugging

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn