Inference
Paris, London
Workplace: HybridFull timeFunction: Software EngineeringExperience: 8+ yearsSkills: ["System-level thinking","Performance optimization","Debugging","Reliability focus","Regression diagnosis"]Build low-latency inference pipelines for on-device robotics, enabling real-time next-token and diffusion-based control loops. Design and optimize distributed inference systems on GPU clusters to increase throughput and efficiency using large-batch serving. Write and integrate low-level CUDA/Triton kernels, then tune workloads for both latency and throughput via batching, scheduling, quantization, caching, and graph compilation. Create monitoring and debugging tools to ensure reliability and rapid regression diagnosis.
Loading
Loading job details...
Preparing the role view and application actions.

