ML Systems Engineer — Inference Acceleration
Paris
Workplace: OnsiteFull timeFunction: IT Operations (Systems/Network Admin)Skills: ["Ownership","Execution","Communication","Problem-solving"]Optimize and serve modern AI models on Arago’s custom accelerator by working across kernels, model execution, multi-device distribution, runtime, and inference serving. Develop custom fused kernels and execution strategies, design mappings across devices, and build inference-serving techniques such as continuous batching and paged KV caches. Partner with hardware, compiler, and runtime teams to co-design abstractions and improve accelerator performance.
Loading
Loading job details...
Preparing the role view and application actions.

