PhD Research Scientist Intern - Reinforcement Learning for Diffusion Modelling

Canva
London
Workplace: HybridInternshipFunction: Data Science & Machine LearningEducation: phdSkills: ["Communication","Presentation skills","Collaboration","Prioritization"]

Work on an industry-scale AI research internship, applying reinforcement learning to diffusion-based generation. You’ll design rubric-guided VLM judging for RGBA decomposition, build VLM-based evaluation methods, and convert evaluators into reward functions for training generative models. Collaborate with researchers, engineers, and product teams to move findings toward production, while contributing results to the research community through publication.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Canva
Canva
13 hours ago

PhD Research Scientist Intern - Reinforcement Learning for Diffusion Modelling

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 5 hours agoStatus: Live

Job Summary

Work on an industry-scale AI research internship, applying reinforcement learning to diffusion-based generation. You’ll design rubric-guided VLM judging for RGBA decomposition, build VLM-based evaluation methods, and convert evaluators into reward functions for training generative models. Collaborate with researchers, engineers, and product teams to move findings toward production, while contributing results to the research community through publication.
Location: London
Workplace: Hybrid
Employment Type: Internship · 4 months
Job Function: Data Science & Machine Learning
Seniority: Intern level

Key Responsibilities

  • •Design and validate a rubric-guided, per-layer VLM judge for RGBA layer decomposition, calibrated against human evaluations.
  • •Build VLM-based methods for automatic, human-aligned evaluation of multi-layer designs.
  • •Turn VLM-based evaluators into reward functions to train generative models in a reinforcement learning setting.
  • •Distill judges into lightweight reward models that score layered images from learned representations at lower inference cost.
  • •Collaborate with research, engineering, and product teams to move findings toward production and the layered-generation roadmap, and contribute results for publication when appropriate.

Key Requirements

  • •Currently completing a PhD, ideally third year or later.
  • •Strong diffusion or flow-matching background, with hands-on policy-gradient RL for generative models (GRPO, PPO, DPO or similar).
  • •Experience fine-tuning VLMs (e.g. with LoRA) and designing prompts or rubrics for evaluation tasks.
  • •Reward modelling experience, preference optimisation, pseudo-labelling, and distillation.
  • •Ability to read and reproduce a recent paper quickly, and communicate technical work clearly in writing and presentations.
Experience:Generative modellingReinforcement learningMultimodal models
Education:PhD / Doctorate
Skills:CommunicationPresentation skillsCollaborationPrioritization
Languages:English
Tech Stack:PyTorchLoRAGRPOPPODPOFSDPDeepSpeed

Company Brief

Canva
Canva is an online visual communications platform that provides drag-and-drop design tools, templates, and media assets for creating presentations, social media graphics, videos, documents and more, serving individuals and large enterprises globally.
Industry: Enterprise Software
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Scaleup
Valuation: Decacorn (USD 10B+)
Funding: Series E+
Headquarters: Surry Hills, New South Wales, Australia
Founded: 2012
Glassdoor
Glassdoor: 3.9
WebsiteLinkedInGlassdoor