AI Inference Engineer

Fuse Energy
London
Workplace: RemoteFull timeFunction: Data Science & Machine LearningSkills: ["Systems thinking","High-stakes ownership","Architecture ownership","Collaboration"]

Define and build how AI inference workloads are served at scale, from first principles. Own the inference serving architecture and strategy, designing the serving stack (request routing, batching, scheduling, autoscaling) for high-throughput, latency-sensitive workloads. Drive model-level optimization (quantisation, distillation, speculative decoding) in partnership with GPU/CUDA teams, select serving frameworks and orchestration, and translate performance commitments into capacity plans. Establish standards, tooling, and benchmarks as the function grows.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Fuse Energy
Fuse Energy
1 month ago

AI Inference Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 1 hour agoStatus: Live

Job Summary

Define and build how AI inference workloads are served at scale, from first principles. Own the inference serving architecture and strategy, designing the serving stack (request routing, batching, scheduling, autoscaling) for high-throughput, latency-sensitive workloads. Drive model-level optimization (quantisation, distillation, speculative decoding) in partnership with GPU/CUDA teams, select serving frameworks and orchestration, and translate performance commitments into capacity plans. Establish standards, tooling, and benchmarks as the function grows.
Location: London
Workplace: Remote
Employment Type: Full time
Job Function: Data Science & Machine Learning
Seniority: Mid level

Key Responsibilities

  • •Define inference serving strategy and architecture from first principles.
  • •Design and build the serving stack including request routing, batching, scheduling, and autoscaling for throughput and latency.
  • •Own model-level optimization strategy for serving, selecting techniques to improve throughput and cost per token in partnership with CUDA/GPU engineers.
  • •Make core architecture calls on serving frameworks and orchestration (e.g., vLLM, TensorRT-LLM, SGLang, Triton Inference Server, or equivalents).
  • •Translate throughput, latency, and uptime commitments into technical specifications and serving capacity plans, and act as a direct owner for inference performance and reliability.

Pay and Benefits

Perks:EquityMeal Allowance

Key Requirements

  • •4+ years building or operating large-scale inference serving systems (or equivalent strong project/industry experience).
  • •Deep, hands-on experience with inference serving frameworks and optimization techniques like batching, KV-cache management, quantisation, and speculative decoding.
  • •Strong systems thinking to reason end-to-end from incoming request to served response across a large cluster.
  • •Comfort working with GPU/CUDA engineers to integrate low-level performance work into a serving system.
  • •Track record of making high-stakes architecture calls and owning outcomes, including operating without a playbook in an early-stage founding role.
Skills:Systems thinkingHigh-stakes ownershipArchitecture ownershipCollaboration
Tech Stack:CUDAGPUVLLMTensorRT-LLMSGLangTriton Inference ServerQuantisationDistillationSpeculative decodingKV-cache managementRequest routingBatchingSchedulingAutoscalingKubernetesSlurm

Company Brief

Fuse Energy
Develops and delivers renewable energy solutions, focusing on clean power projects and services to enable decarbonization and grid flexibility for commercial and utility customers.
Industry: Renewable Energy
Website