Research Engineer, Model Inference & Serving - London
Paris, London, United Kingdom
Workplace: HybridFull timeFunction: QA, Test & Release EngineeringEducation: mastersSkills: ["Collaboration","Communication","Presentation","Teamwork","Problem-solving"]Join a highly skilled Inference team focused on scalable, low-latency inference pipelines for agentic AI. You’ll optimize model performance across memory, throughput and latency using distributed computing, quantization, and caching; develop GPU kernels; collaborate with research on model architectures; review cutting-edge papers; and advance state-of-the-art inference techniques in a hybrid Paris/London setting.
Loading
Loading job details...
Preparing the role view and application actions.

