Backend Inference Runtime Engineer Graduate (AML Inference) - 2027 Start
San Jose
Workplace: OnsiteFull timeFunction: Software EngineeringSkills: ["Cross-team collaboration","Communication","Presentation","Document writing","Responsibility"]Build and optimize the inference runtime for large-model serving. You’ll iterate on the inference engine architecture and end-to-end GPU performance via operator fusion, compilation optimizations, memory-access tuning, and asynchronous scheduling to reduce bottlenecks and latency. Work on distributed parallel strategies (tensor, pipeline, sequence, and MoE expert parallelism), adapt to GPU/NPU hardware, and benchmark/improve against frameworks like vLLM and TensorRT-LLM.
Loading
Loading job details...
Preparing the role view and application actions.

