ML Systems Engineer — Inference Acceleration

Arago
Paris
Workplace: OnsiteFull timeFunction: IT Operations (Systems/Network Admin)Skills: ["Ownership","Execution","Communication","Problem-solving"]

Optimize and serve modern AI models on Arago’s custom accelerator by working across kernels, model execution, multi-device distribution, runtime, and inference serving. Develop custom fused kernels and execution strategies, design mappings across devices, and build inference-serving techniques such as continuous batching and paged KV caches. Partner with hardware, compiler, and runtime teams to co-design abstractions and improve accelerator performance.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Arago
Arago
1 day ago

ML Systems Engineer — Inference Acceleration

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 11 hours agoStatus: Live

Job Summary

Optimize and serve modern AI models on Arago’s custom accelerator by working across kernels, model execution, multi-device distribution, runtime, and inference serving. Develop custom fused kernels and execution strategies, design mappings across devices, and build inference-serving techniques such as continuous batching and paged KV caches. Partner with hardware, compiler, and runtime teams to co-design abstractions and improve accelerator performance.
Location: Paris
Workplace: Onsite
Employment Type: Full time
Job Function: IT Operations (Systems/Network Admin)

Key Responsibilities

  • •Analyze AI workloads to identify kernel-, runtime-, memory-, and system-level bottlenecks on Arago’s accelerator.
  • •Develop and optimize custom kernels, fused operators, and execution strategies to maximize device utilization.
  • •Design efficient multi-device mappings for models and operators, including communication and synchronization strategies.
  • •Develop inference-serving techniques such as continuous batching, paged KV caches, prefix/context caching, chunked prefill, and prefill/decode interleaving.
  • •Build profiling and benchmarking infrastructure across kernels, full models, and serving workloads.

Pay and Benefits

Equity and Bonus:Equity
Perks:EquityHealth InsurancePensionPaid LeavePublic Holidays

Key Requirements

  • •Strong experience in high-performance ML inference, GPU/accelerator programming, or ML systems engineering.
  • •Deep understanding of computer architecture and accelerator/GPU execution, including memory hierarchies, parallelism, and performance bottlenecks.
  • •Experience developing and optimizing custom kernels using CUDA, Triton, ROCm/HIP, or similar low-level environments.
  • •Experience with operator fusion, tiling, scheduling, data movement optimization, and profiling of compute- and memory-bound workloads.
  • •Hands-on inference-serving systems experience (e.g., vLLM, SGLang, TensorRT-LLM), including KV-cache management and prefill/decode scheduling.
Experience:AIMachine learningInferenceGPU accelerationDistributed systems
Skills:OwnershipExecutionCommunicationProblem-solving
Languages:English
Tech Stack:CUDATritonROCmHIPC++PythonVLLMSGLangTensorRT-LLMKV-cacheContinuous batchingPaged attentionPrefillDecode schedulingMLIRMLIR dialects

Company Brief

Arago
Develops energy-efficient photonic AI processors (codenamed “JEF”) that use light-based computation to drastically reduce energy consumption for AI workloads, targeting data centers and large-scale inference deployments.
Industry: Deep Tech
Company Size: Small (11 to 50 employees)
Growth: Early Stage Startup
Funding: Seed
Headquarters: Paris, France
Founded: 2024
Glassdoor
Glassdoor: 3.9
WebsiteGlassdoor