PhD Research Scientist Intern - Reinforcement Learning for Diffusion Modelling

Canva
Vienna
Workplace: HybridInternshipFunction: Data Science & Machine LearningEducation: phdSkills: ["Communication","Collaboration","Technical writing","Prioritization","Multi-threading"]

Join an AI research internship and work on a live, industry-scale project with Canva’s AI team. You’ll design and validate rubric-guided evaluation for multi-layer VLM outputs, build VLM-based automated human-aligned evaluation, and convert evaluators into reward functions for reinforcement learning. You’ll then distill judges into lightweight reward models for layered image scoring and collaborate with research, engineering, and product toward production.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Canva
Canva
5 hours ago

PhD Research Scientist Intern - Reinforcement Learning for Diffusion Modelling

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 5 hours agoStatus: Live

Job Summary

Join an AI research internship and work on a live, industry-scale project with Canva’s AI team. You’ll design and validate rubric-guided evaluation for multi-layer VLM outputs, build VLM-based automated human-aligned evaluation, and convert evaluators into reward functions for reinforcement learning. You’ll then distill judges into lightweight reward models for layered image scoring and collaborate with research, engineering, and product toward production.
Location: Vienna
Workplace: Hybrid
Employment Type: Internship
Job Function: Data Science & Machine Learning
Seniority: Intern level

Key Responsibilities

  • •Design and validate a rubric-guided, per-layer VLM judge for RGBA layer decomposition, calibrated against human evaluations.
  • •Build VLM-based methods for automatic, human-aligned evaluation of multi-layer designs.
  • •Turn VLM-based evaluators into reward functions to train generative models with reinforcement learning.
  • •Distill evaluators into lightweight reward models that score layered images from learned representations at reduced inference cost.
  • •Collaborate with research, engineering, and product teams to move findings toward production and the layered-generation roadmap.

Key Requirements

  • •Currently completing a PhD, ideally third year or later.
  • •Strong diffusion or flow-matching background with hands-on policy-gradient RL for generative models (e.g., GRPO, PPO, DPO).
  • •Experience fine-tuning VLMs (e.g., LoRA) and designing prompts or rubrics for evaluation tasks.
  • •Reward modelling experience including preference optimization, pseudo-labelling, and distillation.
  • •Ability to communicate technical work clearly in writing and presentations, and to work closely with researchers and engineers.
Experience:Generative AIReinforcement learningDiffusion modelsMultimodalVLMAI research
Education:PhD / Doctorate
Skills:CommunicationCollaborationTechnical writingPrioritizationMulti-threading
Languages:English
Tech Stack:Reinforcement learningDiffusion modellingFlow-matchingPolicy-gradient RLGRPOPPODPOVLMLoRAReward modellingPreference optimisationPseudo-labellingDistillationPyTorchFSDPDeepSpeedMulti-GPU trainingRubric-guided evaluationReward functionsReward hacking

Company Brief

Canva
Canva is an online visual communications platform that provides drag-and-drop design tools, templates, and media assets for creating presentations, social media graphics, videos, documents and more, serving individuals and large enterprises globally.
Industry: Enterprise Software
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Scaleup
Valuation: Decacorn (USD 10B+)
Funding: Series E+
Headquarters: Surry Hills, New South Wales, Australia
Founded: 2012
Glassdoor
Glassdoor: 3.9
WebsiteLinkedInGlassdoor