Backend Inference Runtime Engineer Graduate (AML Inference) - 2027 Start

ByteDance
Singapore
Workplace: OnsiteFull timeFunction: Software EngineeringSkills: ["Cross-team collaboration","Communication","Presentation","Document writing","Stress tolerance"]

Build and iterate the architecture of large-model inference runtimes and optimize end-to-end GPU performance, improving throughput and reducing latency through operator fusion, compilation optimizations, and GPU memory/access scheduling. Adapt inference engines across GPU/NPU hardware, and design distributed parallel strategies (tensor, pipeline, sequence, MoE expert parallelism) to efficiently run ultra-large models. Benchmark against vLLM and TensorRT-LLM and implement performance and cost innovations.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
3 days ago

Backend Inference Runtime Engineer Graduate (AML Inference) - 2027 Start

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 12 hours agoStatus: Live

Job Summary

Build and iterate the architecture of large-model inference runtimes and optimize end-to-end GPU performance, improving throughput and reducing latency through operator fusion, compilation optimizations, and GPU memory/access scheduling. Adapt inference engines across GPU/NPU hardware, and design distributed parallel strategies (tensor, pipeline, sequence, MoE expert parallelism) to efficiently run ultra-large models. Benchmark against vLLM and TensorRT-LLM and implement performance and cost innovations.
Location: Singapore
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering
Seniority: Graduate level

Key Responsibilities

  • •Iterate large-model inference engine architecture and optimize end-to-end GPU performance (operator fusion, compilation optimization, GPU memory access, compute pipeline, and Stream asynchronous scheduling) to remove bottlenecks and reduce latency.
  • •Adapt the inference engine across GPU/NPU hardware architectures to improve universality and hardware adaptability.
  • •Design, develop, and optimize distributed parallel solutions for large model inference, implementing tensor parallelism, pipeline parallelism, sequence parallelism, and MoE expert parallelism to improve scalability and efficiency.
  • •Follow cutting-edge global large model inference and distributed parallelism work, including cache optimization and benchmarking against vLLM and TensorRT-LLM.
  • •Continuously iterate on performance and cost advantages of the inference system and build core technical barriers for the team.

Key Requirements

  • •Completing or recently completed a Bachelor's or Master's degree in computing or a related discipline.
  • •Strong low-level foundations with proficiency in C/C++ and Python, plus CUDA programming and familiarity with GPU hardware architecture, memory models, scheduling, and communication mechanisms.
  • •Deep learning operator development/optimization skills, including matrix operations, normalization, activation functions, operator reconstruction, memory access optimization, vectorization acceleration, and precision alignment.
  • •Knowledge of deep learning inference compilation (graph optimization, operator fusion, constant folding, memory reuse, scheduling optimization, quantization compilation) to reduce GPU memory use and inference latency.
  • •Experience using GPU performance analysis tools like Nsight and Profiler to identify bottlenecks and deliver systematic software–hardware optimization for high-concurrency, low-latency inference. (Preferred: distributed parallelism and optimization of vLLM/SGLang/TensorRT-LLM.)
Education:
Skills:Cross-team collaborationCommunicationPresentationDocument writingStress tolerance
Tech Stack:C/C++PythonCUDAGPUGPU hardware architectureGPU memory modelsDeep learningOperator fusionCUDA programmingComputational graph optimizationConstant foldingMemory reuseScheduling optimizationQuantization compilationNsightProfilerVLLMTensorRT-LLMSGLangTensor parallelism

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn