Research Engineer (Evals)

White Circle
Paris, London
Workplace: HybridFull timeUSD 150,000 - 250,000 annuallyFunction: Research & Scientific (R&D)Skills: ["Python","Problem-solving","Collaboration","Communication"]

Research Engineer to build and maintain an internal benchmark suite for single/multi-turn content guardrails and agentic safety, while studying agent behaviours in the wild. Collaborate with product teams to evaluate core functionality of flagship models, extend evals to new features and verticals, and contribute to AI safety research using Python production code.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
White Circle
White Circle
2 months ago

Research Engineer (Evals)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 8 hours agoStatus: Live

Job Summary

Research Engineer to build and maintain an internal benchmark suite for single/multi-turn content guardrails and agentic safety, while studying agent behaviours in the wild. Collaborate with product teams to evaluate core functionality of flagship models, extend evals to new features and verticals, and contribute to AI safety research using Python production code.
Location: Paris, London
Workplace: Hybrid
Employment Type: Full time
Job Function: Research & Scientific (R&D)

Key Responsibilities

  • •Own and maintain our internal benchmark suite, covering single/multi-turn content guardrails and agentic safety.
  • •Build benchmarks that distinguish specific model capabilities.
  • •Work with the product team to build evals covering core functionality of our flagship models.
  • •Build benchmarks for new features coming out of the research team.
  • •Adapt and extend evals to new verticals and changing product data.

Pay and Benefits

Salary: USD 150,000 - 250,000 annually
Equity and Bonus:Equity

Key Requirements

  • •Strong experience coding in Python and shipping production code
  • •Experience building/maintaining benchmarks or eval suites for ML/LLM systems
  • •Familiarity with large language models (LLMs) and frontier models
  • •Ability to design and implement efficient LLN inference setups with parallel calls, retries, and rate-limiting
  • •Automated red-teaming experience is a plus and a strong plus for evaluating AI safety benchmarks
Experience:AI SafetyAI researchLLMsBenchmarks
Skills:PythonProblem-solvingCollaborationCommunication
Languages:English
Tech Stack:PythonLLMsBenchmarkingInferenceParallel processingAPIs

Company Brief

White Circle
Builds AI-driven products and services to help businesses automate workflows, extract insights from data, and improve decision-making using machine learning and natural language processing technologies.
Industry: AI & Machine Learning
Website