ML Runtime Optimization Engineer
Applied Industrial Technologies
Sunnyvale
Workplace: OnsiteFull timeUSD 159,053 - 199,295 annuallyFunction: Software EngineeringExperience: 3+ yearsEducation: bachelorsSkills: ["PyTorch","JAX","ONNX","TensorRT","CUDA","XLA","Triton","Embedded programming"]Develop and optimize ML model performance on embedded runtimes across the full ML stack (PyTorch, JAX, ONNX, TensorRT, CUDA, XLA, Triton) for ADAS/AD applications on multiple embedded compute platforms. Drive pruning/quantization, profiling, and architecture optimization in memory-constrained environments. Collaborate with ML and software teams to deliver efficient, real-time inference on customer hardware.

