Large Language Model Inference System Engineer Graduate (Applied Machine Learning) - 2027 Start

ByteDance
San Jose
Workplace: OnsiteFull timeFunction: Data Science & Machine LearningSkills: ["Coding","Performance analysis","Distributed systems interest","Systems thinking"]

Develop and optimize an end-to-end large model inference (MaaS) system for ultra-large-scale, heterogeneous GPU clusters, improving inference performance, stability, and total cost. Use system-level techniques such as disaggregated multi-role inference and distributed KV cache systems, plus heterogeneous and elastic computing and multi-tenant co-located inference. Work within the machine learning platform team supporting training and inference for recommendation, advertising, vision, speech, and NLP scenarios.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
1 day ago

Large Language Model Inference System Engineer Graduate (Applied Machine Learning) - 2027 Start

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 7 hours agoStatus: Live

Job Summary

Develop and optimize an end-to-end large model inference (MaaS) system for ultra-large-scale, heterogeneous GPU clusters, improving inference performance, stability, and total cost. Use system-level techniques such as disaggregated multi-role inference and distributed KV cache systems, plus heterogeneous and elastic computing and multi-tenant co-located inference. Work within the machine learning platform team supporting training and inference for recommendation, advertising, vision, speech, and NLP scenarios.
Location: San Jose
Workplace: Onsite
Employment Type: Full time
Job Function: Data Science & Machine Learning
Seniority: Graduate level

Key Responsibilities

  • •Participate in engineering development of the Volcano Ark MaaS inference system.
  • •Optimize large model inference performance, cost, and stability in ultra-large-scale heterogeneous inference clusters.
  • •Reduce large model inference costs using system-level approaches including disaggregated multi-role inference, distributed KV cache systems, heterogeneous inference, elastic computing, and multi-tenant co-located inference.

Key Requirements

  • •Currently completing or recently completed a Bachelor’s or Master’s degree in Computer Science or a related discipline.
  • •Strong command of algorithms, design patterns, and data structures, with solid knowledge of operating systems and computer architecture.
  • •Proficient in one or more programming languages such as C++ or Python with good coding style.
  • •Understands GPU hardware architecture and is familiar with high-performance computing software stacks such as CUDA, including GPU performance analysis.
  • •Strong interest in distributed systems and large-scale heterogeneous inference, including studying performance bottlenecks.
Education:
Skills:CodingPerformance analysisDistributed systems interestSystems thinking
Tech Stack:C++PythonCUDAGPUKubernetesRayRDMAKV CacheGPU performance analysis

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn