Inference
Genesis AI
Paris, London
Workplace: HybridFull timeFunction: Software EngineeringExperience: 8+ yearsSkills: ["System-level thinking","Performance optimization","Debugging","Reliability focus","Regression diagnosis"]Build low-latency inference pipelines for on-device robotics, enabling real-time next-token and diffusion-based control loops. Design and optimize distributed inference systems on GPU clusters to increase throughput and efficiency using large-batch serving. Write and integrate low-level CUDA/Triton kernels, then tune workloads for both latency and throughput via batching, scheduling, quantization, caching, and graph compilation. Create monitoring and debugging tools to ensure reliability and rapid regression diagnosis.

