Research Scientist, Benchmarks & Evaluations
Protege
United States
Workplace: RemoteFull timeFunction: Data Science & Machine LearningEducation: phdSkills: ["Communication","Written communication","Problem-solving","Collaboration","Critical thinking"]Lead design and evaluation of benchmarks for frontier AI models, designing tasks that meaningfully differentiate capabilities, validating them with human baselines, and publishing findings that set Protege’s standard for AI evaluation datasets. Collaborate with DataLab, engineering, and partner labs to translate research into deployable evaluation data and product features.

