Machine Learning Researcher, RL & Agentic Systems

Protege
United States
Workplace: RemoteFull timeFunction: Data Science & Machine LearningEducation: mastersSkills: ["Research","Evaluation","Benchmarking","Dataset design","Task design"]

Design and evaluate datasets, tasks, and environments for benchmarking agentic AI systems. Develop evaluation frameworks, measure dataset quality, and benchmark model behavior in RL- and agentic settings. Collaborate with research, engineering, and product to drive high-quality, real-world data for frontier AI.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Protege
Protege
3 months ago

Machine Learning Researcher, RL & Agentic Systems

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 23 hours agoStatus: Live

Job Summary

Design and evaluate datasets, tasks, and environments for benchmarking agentic AI systems. Develop evaluation frameworks, measure dataset quality, and benchmark model behavior in RL- and agentic settings. Collaborate with research, engineering, and product to drive high-quality, real-world data for frontier AI.
Location: United States
Workplace: Remote
Employment Type: Full time
Job Function: Data Science & Machine Learning

Key Responsibilities

  • •Design and build datasets, tasks, environments, and evaluation assets for benchmarking agentic systems and multi-step model behavior.
  • •Translate real-world workflows into structured tasks, interaction traces, trajectories, stateful environments, and verifiable outcomes that can be used to evaluate advanced AI systems.
  • •Develop frameworks that assess diversity, realism, coverage, fidelity, informativeness, and downstream usefulness of datasets for agentic systems.
  • •Build quality scorecards and evaluation methods that make dataset strengths, weaknesses, and failure modes legible across teams.
  • •Collaborate closely with research and engineering teams to identify data bottlenecks, improve evaluation methodology, and shape internal best practices around task-grounded AI training data.

Key Requirements

  • •PhD or equivalent Master’s Degree + 4+ years of industry experience in machine learning, computer science, statistics, engineering, mathematics, economics, or related quantitative fields.
  • •Strong understanding of AI model training pipelines, evaluation methodologies, and the role of data in shaping model performance.
  • •Experience with large, unstructured, or semi-structured datasets used to train or evaluate ML systems.
  • •Experience with reinforcement learning, sequential decision-making, agentic systems, tool-using models, or multi-step model evaluation.
  • •Experience designing tasks, benchmarks, environments, simulations, or evaluation frameworks for real-world model behavior.
Experience:AIMLData science
Education:Master's
Skills:ResearchEvaluationBenchmarkingDataset designTask design
Languages:English

Company Brief

Protege
Provides a privacy‑first platform that connects data holders with vetted AI developers to license and deliver high‑quality, multimodal training and evaluation datasets across healthcare, media, audio, and motion capture.
Industry: Data Infrastructure
Company Size: Small (11 to 50 employees)
Growth: Growth Stage Startup
Funding: Series A
Headquarters: New York City, United States
Founded: 2024
WebsiteLinkedIn