Research Engineer, Model Inference & Serving - Paris

H Company
Paris, London
Workplace: HybridFull timeFunction: Software EngineeringEducation: mastersSkills: ["Collaboration","Communication","Teamwork","Problem-solving","Adaptability"]

A research-oriented software engineering role focused on building and optimizing low-latency AI inference pipelines for agentic models. You will work on GPU-accelerated kernels, model compression techniques, and collaborate with research teams to push Memory/Throughput/Latency improvements in a hybrid Paris/London environment.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
H Company
H Company
5 months ago

Research Engineer, Model Inference & Serving - Paris

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 8 hours agoStatus: Live

Job Summary

A research-oriented software engineering role focused on building and optimizing low-latency AI inference pipelines for agentic models. You will work on GPU-accelerated kernels, model compression techniques, and collaborate with research teams to push Memory/Throughput/Latency improvements in a hybrid Paris/London environment.
Location: Paris, London
Workplace: Hybrid
Employment Type: Full time
Job Function: Software Engineering

Key Responsibilities

  • •Develop scalable, low-latency and cost effective inference pipelines
  • •Optimize model performance for memory usage, throughput and latency
  • •Develop specialized GPU kernels for attention mechanisms, matrix multiplications, etc.
  • •Collaborate with research teams on model architectures to enhance efficiency during inference
  • •Review state-of-the-art papers to improve memory usage, throughput and latency

Key Requirements

  • •MS or PhD in Computer Science, Machine Learning or related fields
  • •Proficient in Python, Rust or C/C++
  • •Experience in GPU programming such as CUDA, Open AI Triton, Metal
  • •Experience in model compression and quantization techniques
  • •Soft skills: Collaborative mindset, strong communication and presentation skills, eager to explore new challenges
Experience:AIMLInference
Education:Master's
Skills:CollaborationCommunicationTeamworkProblem-solvingAdaptability
Tech Stack:PythonRustC/C++CUDAOpen AI TritonMetalModel compressionQuantizationLlama.cppPyTorchONNX RuntimeTensorRT-LLMVLLMNCCLGgml

Company Brief

H Company
Develops agentic AI foundation models and deployable AI agents (e.g., Surfer H, Runner H) to automate web and enterprise tasks, improving productivity for large organisations through visual-language and planning capabilities.
Industry: AI & Machine Learning
Company Size: Medium (51 to 250 employees)
Growth: Early Stage Startup
Funding: Seed
Headquarters: Paris, France
Founded: 2023
Glassdoor
Glassdoor: 2.5
WebsiteLinkedInGlassdoor