Inference Systems Backend Engineer - ARK Large Model Platform (Singapore)

ByteDance
Singapore
Workplace: OnsiteFull timeFunction: Software EngineeringSkills: ["Problem-solving","Teamwork","Communication","Self-motivation"]

Build and optimize Volcano Engine’s large model training and inference systems, including computation optimization, tuning thousand-GPU clusters, distributed LLM inference, and large-scale traffic scheduling. Tackle high-concurrency reliability and scalability challenges for workloads measured in hundreds of billions of tokens. Advance training/inference architectures (e.g., subgraph matching, compiler optimization, quantization), integrate heterogeneous GPU/NPU/TPU hardware, and improve utilization via elastic scheduling and GPU oversubscription while partnering with algorithm teams.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
1 month ago

Inference Systems Backend Engineer - ARK Large Model Platform (Singapore)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 30 days agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Build and optimize Volcano Engine’s large model training and inference systems, including computation optimization, tuning thousand-GPU clusters, distributed LLM inference, and large-scale traffic scheduling. Tackle high-concurrency reliability and scalability challenges for workloads measured in hundreds of billions of tokens. Advance training/inference architectures (e.g., subgraph matching, compiler optimization, quantization), integrate heterogeneous GPU/NPU/TPU hardware, and improve utilization via elastic scheduling and GPU oversubscription while partnering with algorithm teams.
Location: Singapore
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering

Key Responsibilities

  • •Develop and performance-optimize Volcano Engine large model training and inference systems, including computation optimization, thousand-GPU cluster tuning, distributed LLM inference, and large-scale inference traffic scheduling.
  • •Solve technical challenges for high concurrency, high reliability, and high scalability to support daily training and inference traffic at hundreds of billions of tokens.
  • •Research and introduce forward-looking training and inference architectures such as subgraph matching, compiler optimization, and model quantization.
  • •Integrate heterogeneous hardware with training and inference frameworks across GPUs, NPUs, and TPUs.
  • •Improve compute utilization across globally distributed ultra-large-scale GPU clusters using elastic scheduling, GPU oversubscription, and task orchestration; collaborate with algorithm teams to optimize algorithms and systems.

Key Requirements

  • •Proficiency in C/C++ and Python development under Linux environments, with experience in large-scale ML systems or search, advertising, and recommendation systems.
  • •Familiarity with at least one machine learning framework such as TensorFlow, PyTorch, or MXNet (or similar in-house frameworks).
  • •Familiarity with at least one large model training or inference framework, including vLLM, TensorRT-LLM, SGLang, or Megatron-LM.
  • •Strong problem-solving ability and capability to work independently while collaborating effectively as a team.
  • •Strong sense of responsibility, good learning ability, communication skills, and self-motivation.
Experience:Large-scale ML
Education:
Skills:Problem-solvingTeamworkCommunicationSelf-motivation
Tech Stack:CC++PythonLinuxTensorFlowPyTorchMXNetVLLMTensorRT-LLMSGLangMegatron-LMCUDACuDNNVolcanoEngineGPUNPUTPU

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn