Research Engineer, Model Inference & Serving - London

H Company
Paris, London, United Kingdom
Workplace: HybridFull timeFunction: QA, Test & Release EngineeringEducation: mastersSkills: ["Collaboration","Communication","Presentation","Teamwork","Problem-solving"]

Join a highly skilled Inference team focused on scalable, low-latency inference pipelines for agentic AI. You’ll optimize model performance across memory, throughput and latency using distributed computing, quantization, and caching; develop GPU kernels; collaborate with research on model architectures; review cutting-edge papers; and advance state-of-the-art inference techniques in a hybrid Paris/London setting.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
H Company
H Company
10 months ago

Research Engineer, Model Inference & Serving - London

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 3 hours agoStatus: Live

Job Summary

Join a highly skilled Inference team focused on scalable, low-latency inference pipelines for agentic AI. You’ll optimize model performance across memory, throughput and latency using distributed computing, quantization, and caching; develop GPU kernels; collaborate with research on model architectures; review cutting-edge papers; and advance state-of-the-art inference techniques in a hybrid Paris/London setting.
Location: Paris, London, United Kingdom
Workplace: Hybrid
Employment Type: Full time
Job Function: QA, Test & Release Engineering

Key Responsibilities

  • •Develop scalable, low-latency and cost effective inference pipelines
  • •Optimize model performance: memory usage, throughput, and latency using distributed computing, model compression, quantization and caching mechanisms
  • •Develop specialized GPU kernels for performance-critical tasks like attention mechanisms, matrix multiplications, etc.
  • •Collaborate with research teams on model architectures to enhance efficiency during inference
  • •Review state-of-the-art papers to improve memory usage, throughput and latency (Flash attention, Paged Attention, Continuous batching, etc.)

Key Requirements

  • •MS or PhD in Computer Science, Machine Learning or related fields
  • •Proficiency in Python, Rust or C/C++
  • •Experience in GPU programming (CUDA, OpenAI Triton, Metal, etc.)
  • •Experience in model compression and quantization techniques
  • •Collaborative mindset and strong communication skills
Experience:AIMachine LearningInference
Education:Master's
Skills:CollaborationCommunicationPresentationTeamworkProblem-solving
Tech Stack:PythonRustC/C++CUDAOpenAI TritonMetalModel compressionQuantizationVLLMTensorRT-LLMSGLangLlama.cppNCCLPyTorchONNX Runtime

Company Brief

H Company
Develops agentic AI foundation models and deployable AI agents (e.g., Surfer H, Runner H) to automate web and enterprise tasks, improving productivity for large organisations through visual-language and planning capabilities.
Industry: AI & Machine Learning
Company Size: Medium (51 to 250 employees)
Growth: Early Stage Startup
Funding: Seed
Headquarters: Paris, France
Founded: 2023
Glassdoor
Glassdoor: 2.5
WebsiteLinkedInGlassdoor