ML Infrastructure Engineer

White Circle
Paris, London
Workplace: HybridFull timeUSD 180,000 - 350,000 annuallyFunction: DevOps, Cloud & InfrastructureSkills: ["Ownership mindset","Problem-solving","Collaboration"]

Build the systems behind LLM post-training and RL workflows, including scalable pipelines for smoke tuning, evaluation, inference, and agentic development. Design data control systems and tune end-to-end training/inference performance across compute, memory, networking, storage, and checkpointing. Investigate how infrastructure affects learning dynamics and stability, while improving experiment iteration tooling such as artifacts, dashboards, failure inspection, and cost visibility.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
White Circle
White Circle
2 months ago

ML Infrastructure Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 5 hours agoStatus: Live

Job Summary

Build the systems behind LLM post-training and RL workflows, including scalable pipelines for smoke tuning, evaluation, inference, and agentic development. Design data control systems and tune end-to-end training/inference performance across compute, memory, networking, storage, and checkpointing. Investigate how infrastructure affects learning dynamics and stability, while improving experiment iteration tooling such as artifacts, dashboards, failure inspection, and cost visibility.
Location: Paris, London
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure

Key Responsibilities

  • •Build robust, flexible, and scalable RL and post-training pipelines, including smoke tuning runs and approach ablations.
  • •Design data control systems governing model inputs and training data flows through rollouts, replay, filtering, evaluation, and policy updates.
  • •Tune training and inference end-to-end for high throughput across networking, memory, compute scheduling, data loading, storage, checkpointing, and I/O.
  • •Investigate how infrastructure choices affect learning dynamics, evaluation quality, model behavior, and training stability.
  • •Build infrastructure for model iteration and agentic development environments, including experiment runs, artifacts, evals, dashboards, failure inspection, reproducibility, and cost visibility.

Pay and Benefits

Salary: USD 180,000 - 350,000 annually
Equity and Bonus:Equity
Perks:Health InsurancePaid LeaveRelocation

Key Requirements

  • •Designed, built, or maintained distributed RL/post-training systems at scale, including rollouts, replay buffers, reward signals, data filtering, policy updates, evaluation loops, and failure analysis.
  • •Comfortable with deep learning frameworks such as PyTorch or JAX.
  • •Proficient in Python, including concurrency, asynchronous programming, multiprocessing, and performance optimization.
  • •Ability to debug distributed GPU workloads across CUDA runtime, container runtime, drivers, NCCL (or equivalent), networking, storage, scheduling, and checkpointing.
  • •Experience with inference stacks (e.g., vLLM, SGLang, TensorRT-LLM, Dynamo) and profiling tools across the stack (e.g., py-spy, PyTorch profiler, Nsight, perf, tracing, metrics, logs).
Experience:LLMReinforcement learningAI safetyDistributed systems
Skills:Ownership mindsetProblem-solvingCollaboration
Languages:English
Tech Stack:PyTorchJAXPythonCUDANCCLVLLMSGLangTensorRT-LLMDynamoKubernetesSlurmRayUCXNVSHMEMRDMAInfiniBandRoCEEFARustC++

Company Brief

White Circle
Builds AI-driven products and services to help businesses automate workflows, extract insights from data, and improve decision-making using machine learning and natural language processing technologies.
Industry: AI & Machine Learning
Website