Backend Inference Runtime Engineer Graduate (AML Inference) - 2027 Start

ByteDance
San Jose
Workplace: OnsiteFull timeFunction: Software EngineeringSkills: ["Cross-team collaboration","Communication","Presentation","Document writing","Responsibility"]

Build and optimize the inference runtime for large-model serving. You’ll iterate on the inference engine architecture and end-to-end GPU performance via operator fusion, compilation optimizations, memory-access tuning, and asynchronous scheduling to reduce bottlenecks and latency. Work on distributed parallel strategies (tensor, pipeline, sequence, and MoE expert parallelism), adapt to GPU/NPU hardware, and benchmark/improve against frameworks like vLLM and TensorRT-LLM.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
2 hours ago

Backend Inference Runtime Engineer Graduate (AML Inference) - 2027 Start

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 2 hours agoStatus: Live

Job Summary

Build and optimize the inference runtime for large-model serving. You’ll iterate on the inference engine architecture and end-to-end GPU performance via operator fusion, compilation optimizations, memory-access tuning, and asynchronous scheduling to reduce bottlenecks and latency. Work on distributed parallel strategies (tensor, pipeline, sequence, and MoE expert parallelism), adapt to GPU/NPU hardware, and benchmark/improve against frameworks like vLLM and TensorRT-LLM.
Location: San Jose
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering
Seniority: Graduate level

Key Responsibilities

  • •Iterate on the large-model inference engine architecture and optimize end-to-end GPU performance to improve single-card throughput and reduce inference latency.
  • •Adapt the inference engine to different GPU/NPU hardware architectures by improving universality and hardware adaptability.
  • •Design, develop, and optimize distributed parallel solutions for large model inference, including tensor, pipeline, sequence, and MoE expert parallelism.
  • •Address multi-card splitting, cross-card communication overhead, load imbalance, and low parallel efficiency to improve parallel efficiency.
  • •Track cutting-edge inference technologies, benchmark against vLLM and TensorRT-LLM, and continuously optimize performance and cost advantages.

Key Requirements

  • •Currently completing or recently completed a Bachelor’s/Master’s degree in Software Development, Computer Science, Computer Engineering, or a related technical discipline.
  • •Strong low-level foundations with proficiency in C/C++ and Python, plus CUDA programming and familiarity with GPU hardware architecture and memory/computation/communication principles.
  • •Experience with deep-learning inference operators and GPU adaptation/optimization for matrix operations, normalization, and activation functions, including operator reconstruction and memory-access optimizations.
  • •Understanding of deep-learning inference compilation and techniques such as computational graph optimization, operator fusion, constant folding, memory reuse, scheduling optimization, and quantization compilation.
  • •Ability to use GPU performance analysis tools (e.g., Nsight and Profiler) to identify bottlenecks and implement systematic software–hardware optimization solutions.
Skills:Cross-team collaborationCommunicationPresentationDocument writingResponsibility
Tech Stack:CC++PythonCUDAGPUNPUOperator fusionCompilation optimizationGPU memory modelsComputing schedulingAsynchronous schedulingVLLMTensorRT-LLMNsightProfilerTensor parallelismPipeline parallelismSequence parallelismMoE expert parallelismDistributed parallelism

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn