LLM System Backend Engineer Graduate (AML Ark) - 2027 Start

ByteDance
Singapore
Full timeFunction: Software EngineeringSkills: []

Build backend systems for Volcano Ark’s foundation model and Agent platform. Work on both training and inference tracks: develop serverless post-training workflows (SFT and RL), design elastic multi-tenant training across data centers and heterogeneous hardware, and create an inference (Model-as-a-Service) platform that improves performance, cost efficiency, and reliability across GPU clusters. Optimize architectures such as distributed KV caches, elastic scheduling, and disaggregated inference.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
3 days ago

LLM System Backend Engineer Graduate (AML Ark) - 2027 Start

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 12 hours agoStatus: Live

Job Summary

Build backend systems for Volcano Ark’s foundation model and Agent platform. Work on both training and inference tracks: develop serverless post-training workflows (SFT and RL), design elastic multi-tenant training across data centers and heterogeneous hardware, and create an inference (Model-as-a-Service) platform that improves performance, cost efficiency, and reliability across GPU clusters. Optimize architectures such as distributed KV caches, elastic scheduling, and disaggregated inference.
Location: Singapore
Employment Type: Full time
Job Function: Software Engineering
Seniority: Graduate level

Key Responsibilities

  • •For the training track, develop the Volcano Ark training platform enabling serverless post-training (SFT and RL) for internal and external users.
  • •Design elastic training solutions for complex multi-tenant workloads across multiple data centers and heterogeneous hardware while optimizing throughput, resource utilization, and stability.
  • •Build reinforcement learning infrastructure to improve training efficiency and create developer-friendly APIs for RL training workflows.
  • •For the inference track, develop the Volcano Ark MaaS inference platform to optimize LLM inference performance, cost efficiency, and reliability across heterogeneous GPU clusters.
  • •Reduce inference costs through system optimizations such as disaggregated inference architectures, distributed KV cache systems, elastic compute scheduling, and multi-tenant co-located inference.

Key Requirements

  • •Completing or recently completed a Bachelor's or Master's degree in computing or a related discipline.
  • •Proficiency in one or more of Python, Rust, or C++ with strong coding practices and clean software design.
  • •Hands-on experience designing training frameworks or optimizing large-scale training systems, with LLM training infrastructure or complex distributed systems experience preferred.
  • •Solid understanding of GPU architectures and familiarity with high-performance computing software stacks such as CUDA, including GPU performance profiling and bottleneck analysis.
  • •Preferred: experience optimizing systems architecture for heterogeneous hardware and high-performance networking, including Kubernetes (K8s) and Ray.
Education:
Tech Stack:PythonRustC++CUDAKubernetesK8sRayRDMA

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn