Principal Research Scientist - Evaluations

Canva
Sydney
Workplace: HybridFull timeFunction: Data Science & Machine LearningSkills: ["Leadership","Mentoring","Communication","Research-minded judgment"]

Define and lead Canva Research’s evaluation strategy for generative quality across design and multimodal content, from human-rater benchmarks to automated judges. Own measurement principles that align metrics with user experience and downstream product outcomes, including correlation, bias, saturation, and production monitoring. Build a shared evaluation layer across regions, mentor researchers, and partner with Design Generation, Foundation Models, and Agents teams to shape training and inference decisions.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Canva
Canva
2 days ago

Principal Research Scientist - Evaluations

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 5 hours agoStatus: Live

Job Summary

Define and lead Canva Research’s evaluation strategy for generative quality across design and multimodal content, from human-rater benchmarks to automated judges. Own measurement principles that align metrics with user experience and downstream product outcomes, including correlation, bias, saturation, and production monitoring. Build a shared evaluation layer across regions, mentor researchers, and partner with Design Generation, Foundation Models, and Agents teams to shape training and inference decisions.
Location: Sydney
Workplace: Hybrid
Employment Type: Full time · Permanent
Job Function: Data Science & Machine Learning
Seniority: Sr. Manager level

Key Responsibilities

  • •Establish evaluation gates that inform launch decisions and make judgement calls when signals are ambiguous.
  • •Diagnose anomalous evaluation results during production training runs, separating model regressions from infrastructure artifacts and communicating outcomes quickly.
  • •Partner with Design Generation, Foundation Models, and Agents teams so evaluation shapes training and inference rather than being reported after the fact.
  • •Work with designers, creators, and product teams to convert subjective creative judgment into measurable criteria.
  • •Mentor senior research scientists and engineers and raise evaluation craft standards across the group.

Pay and Benefits

Perks:EquityParental LeaveWellness Stipend

Key Requirements

  • •Track record defining evaluation frameworks and standards from the ground up, ideally for GenAI evaluation and measurement systems that changed team decisions.
  • •Experience linking evaluation metrics to downstream user or business outcomes and diagnosing metric divergences.
  • •Experience turning subjective human judgment into reliable evaluation signals via rubric design, human data pipelines, and model training.
  • •Strong grounding in multimodal generative models (diffusion, transformers, VLMs, and MLLMs) and where evaluation can break.
  • •Experience with reward modelling, preference learning, or alignment methods involving human feedback.
Experience:Generative AIMultimodalGenAI evaluation
Skills:LeadershipMentoringCommunicationResearch-minded judgment
Languages:English
Tech Stack:DiffusionTransformersVLMsMLLMsRLHFRLAIF

Company Brief

Canva
Canva is an online visual communications platform that provides drag-and-drop design tools, templates, and media assets for creating presentations, social media graphics, videos, documents and more, serving individuals and large enterprises globally.
Industry: Enterprise Software
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Scaleup
Valuation: Decacorn (USD 10B+)
Funding: Series E+
Headquarters: Surry Hills, New South Wales, Australia
Founded: 2012
Glassdoor
Glassdoor: 3.9
WebsiteLinkedInGlassdoor