Research Engineer - LLM/VLM Inference Optimization (Seed Infra)

ByteDance
Seattle
Workplace: OnsiteFull timeFunction: Research & Scientific (R&D)Education: bachelorsSkills: ["Performance optimization","Performance analysis","Debugging","Collaboration"]

Build and optimize high-performance inference systems for large-scale LLMs and VLMs, including inference engines, serving frameworks, and end-to-end deployment pipelines. Apply compiler-level optimizations, parallel computing, graph fusion, low-precision computation, streaming inference, speculative decoding, and high-concurrency request tuning. Work with research teams to analyze bottlenecks and improve production serving performance, toolchains, and the broader inference ecosystem.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
1 month ago

Research Engineer - LLM/VLM Inference Optimization (Seed Infra)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 30 days agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Build and optimize high-performance inference systems for large-scale LLMs and VLMs, including inference engines, serving frameworks, and end-to-end deployment pipelines. Apply compiler-level optimizations, parallel computing, graph fusion, low-precision computation, streaming inference, speculative decoding, and high-concurrency request tuning. Work with research teams to analyze bottlenecks and improve production serving performance, toolchains, and the broader inference ecosystem.
Location: Seattle
Workplace: Onsite
Employment Type: Full time
Job Function: Research & Scientific (R&D)

Key Responsibilities

  • •Design, develop, and optimize high-performance inference systems for large-scale LLMs and VLMs, including inference engines, serving frameworks, and deployment pipelines.
  • •Build state-of-the-art inference engines using performance optimization techniques such as compiler-level optimizations, parallel computing, graph fusion, CUDA kernel development, low-precision computation, streaming inference, speculative decoding, and high-concurrency request optimization.
  • •Collaborate with other research teams to identify performance bottlenecks, conduct performance analysis, and optimize large models.
  • •Develop model toolchains and contribute to the broader technical ecosystem supporting inference.
  • •Optimize end-to-end deployment for large models to improve production performance outcomes.

Key Requirements

  • •Bachelor's degree or above in Computer Science, Electrical Engineering, Software Engineering, or a related field.
  • •Strong proficiency in C/C++ and Python, with fundamentals in algorithms, data structures, and systems programming; familiarity with containerization and server-side debugging.
  • •Hands-on experience with at least one mainstream machine learning framework (e.g., PyTorch, TensorFlow).
  • •Experience deploying or optimizing LLM/VLM inference at production scale with demonstrated impact on latency, throughput, or serving cost.
  • •Familiarity with GPU architecture and experience optimizing compute-intensive operators (e.g., FlashAttention, GEMM, GEMV, Conv2D).
Education:Bachelor's
Skills:Performance optimizationPerformance analysisDebuggingCollaboration
Tech Stack:C/C++PythonC++PyTorchTensorFlowCUDAOpenCLTensorRTTritonCUTLASSFlashAttentionGEMMGEMVConv2DContainerizationStreaming inferenceSpeculative decoding

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn