Backend Engineer - AML Framework Development (Search, Ads, and Recommendation Direction)

ByteDance
Singapore
Workplace: OnsiteFull timeFunction: Software EngineeringEducation: bachelorsSkills: []

Build and optimize the AML large-model inference engine for ByteDance’s ads, search, and e-commerce ranking systems. You’ll improve end-to-end GPU performance via operator fusion, compilation optimization, memory access tuning, and asynchronous scheduling, while benchmarking against vLLM and TensorRT-LLM. Own distributed parallel inference design and implementation (tensor/pipeline/sequence and MoE expert parallelism) to maximize throughput, reduce latency, and improve multi-card efficiency across GPU/NPU hardware.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
1 month ago

Backend Engineer - AML Framework Development (Search, Ads, and Recommendation Direction)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 29 days agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Build and optimize the AML large-model inference engine for ByteDance’s ads, search, and e-commerce ranking systems. You’ll improve end-to-end GPU performance via operator fusion, compilation optimization, memory access tuning, and asynchronous scheduling, while benchmarking against vLLM and TensorRT-LLM. Own distributed parallel inference design and implementation (tensor/pipeline/sequence and MoE expert parallelism) to maximize throughput, reduce latency, and improve multi-card efficiency across GPU/NPU hardware.
Location: Singapore
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Iterate the large-model inference engine architecture and optimize end-to-end GPU performance using techniques like operator fusion and compilation-level improvements.
  • •Adapt inference solutions across GPU/NPU hardware architectures to improve inference engine universality and hardware adaptability.
  • •Design, develop, and optimize distributed parallel inference solutions, including tensor/pipeline/sequence parallelism and MoE expert parallelism to address multi-card efficiency challenges.
  • •Deeply optimize GPU memory access, computing pipelines, and asynchronous scheduling to eliminate inference bottlenecks and improve throughput and latency.
  • •Benchmark and iterate on cutting-edge inference technologies and implementations (e.g., global large-model inference, cache optimization) against mainstream frameworks such as vLLM and TensorRT-LLM.

Key Requirements

  • •Bachelor’s degree in Computer Science or equivalent with 3+ years of relevant experience.
  • •Strong low-level foundation in C/C++ and Python, with CUDA programming skills and knowledge of GPU hardware architecture, memory models, scheduling, and communication mechanisms.
  • •Ability to implement and optimize deep-learning operators for inference (e.g., matrix operations, normalization, activation), including operator reconstruction, memory access optimization, vectorization, and precision alignment.
  • •Knowledge of end-to-end deep learning inference compilation (graph optimization, operator fusion, constant folding, memory reuse, scheduling optimization, quantization compilation) to reduce memory usage and inference latency.
  • •Proficiency using GPU performance analysis tools like Nsight and Profiler to identify bottlenecks and deliver iterative software-hardware optimization solutions for low-latency, high-concurrency inference.
Experience:AI infrastructureMachine learning inferenceDistributed computingAdsSearch
Education:Bachelor's
Tech Stack:C/C++PythonCUDAGPUNsightProfilerVLLMTensorRT-LLMSGLangStreamMoETensor parallelismPipeline parallelismSequence parallelismOperator fusionQuantization

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn