Backend Inference Runtime Engineer Graduate (AML Inference) - 2027 Start
Singapore
Workplace: OnsiteFull timeFunction: Software EngineeringSkills: ["Cross-team collaboration","Communication","Presentation","Document writing","Stress tolerance"]Build and iterate the architecture of large-model inference runtimes and optimize end-to-end GPU performance, improving throughput and reducing latency through operator fusion, compilation optimizations, and GPU memory/access scheduling. Adapt inference engines across GPU/NPU hardware, and design distributed parallel strategies (tensor, pipeline, sequence, MoE expert parallelism) to efficiently run ultra-large models. Benchmark against vLLM and TensorRT-LLM and implement performance and cost innovations.
Loading
Loading job details...
Preparing the role view and application actions.

