AI Inference Engineer
Fuse Energy
London, Dubai
Workplace: RemoteFull timeFunction: Data Science & Machine LearningSkills: ["Systems thinking","High-stakes ownership","Architecture ownership","Collaboration"]Define and build how AI inference workloads are served at scale, from first principles. Own the inference serving architecture and strategy, designing the serving stack (request routing, batching, scheduling, autoscaling) for high-throughput, latency-sensitive workloads. Drive model-level optimization (quantisation, distillation, speculative decoding) in partnership with GPU/CUDA teams, select serving frameworks and orchestration, and translate performance commitments into capacity plans. Establish standards, tooling, and benchmarks as the function grows.

