Inference Systems Backend Engineer - ARK Large Model Platform (Singapore)
ByteDance
Singapore
Workplace: OnsiteFull timeFunction: Software EngineeringSkills: ["Problem-solving","Teamwork","Communication","Self-motivation"]Build and optimize Volcano Engine’s large model training and inference systems, including computation optimization, tuning thousand-GPU clusters, distributed LLM inference, and large-scale traffic scheduling. Tackle high-concurrency reliability and scalability challenges for workloads measured in hundreds of billions of tokens. Advance training/inference architectures (e.g., subgraph matching, compiler optimization, quantization), integrate heterogeneous GPU/NPU/TPU hardware, and improve utilization via elastic scheduling and GPU oversubscription while partnering with algorithm teams.

