Senior Research Data Engineer

Canva
London
Workplace: HybridFull timeFunction: Research & Scientific (R&D)Skills: ["Communication","Collaboration","Ownership","Iteration","Problem-solving"]

Build and own the data foundations for multimodal agent research. Design end-to-end pipelines for collecting, curating, validating, and versioning text, image, and multimodal data, then deliver it into scalable training workflows. Partner with research scientists to translate requirements into data specifications, create evaluation datasets and benchmarks, and develop tooling for annotation, synthetic data, and preference collection (RLHF/DPO). Ensure quality, reliability, and reproducibility through monitoring and testing.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Canva
Canva
2 days ago

Senior Research Data Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 11 minutes agoStatus: Live

Job Summary

Build and own the data foundations for multimodal agent research. Design end-to-end pipelines for collecting, curating, validating, and versioning text, image, and multimodal data, then deliver it into scalable training workflows. Partner with research scientists to translate requirements into data specifications, create evaluation datasets and benchmarks, and develop tooling for annotation, synthetic data, and preference collection (RLHF/DPO). Ensure quality, reliability, and reproducibility through monitoring and testing.
Location: London
Workplace: Hybrid
Employment Type: Full time
Job Function: Research & Scientific (R&D)
Seniority: Mid level

Key Responsibilities

  • •Design and build data pipelines for agent training, including collection, filtering, deduplication, formatting, and versioning across text, image, and multimodal sources.
  • •Build and maintain infrastructure for scalable data loading, storage, and retrieval (e.g., S3, distributed systems, streaming pipelines).
  • •Collaborate with research scientists to turn research requirements into data specifications and iterate as experiments evolve.
  • •Create evaluation datasets and benchmarks by curating task distributions that surface real failure modes.
  • •Own data quality and reliability through validation frameworks, drift/contamination monitoring, comprehensive test coverage, and thorough dataset documentation (provenance and limitations).

Key Requirements

  • •Strong software engineering skills in Python, with experience building production-grade data pipelines for ML workflows.
  • •Experience with ML data workflows, including large-scale data processing/loading (e.g., Ray), data versioning, and training data formatting (tokenization, batching, sharding).
  • •Hands-on experience supporting data pipelines for large-scale distributed ML training runs.
  • •Practical prompt engineering experience for reliable LLM/VLM outputs.
  • •Experience with annotation tooling and human-in-the-loop data collection (Label Studio or internal systems), plus strong communication with researchers.
Skills:CommunicationCollaborationOwnershipIterationProblem-solving
Languages:English
Tech Stack:PythonS3AWSRayLabel StudioRLHFDPOLLMVLMStreaming pipelinesTokenizationBatchingSharding

Company Brief

Canva
Canva is an online visual communications platform that provides drag-and-drop design tools, templates, and media assets for creating presentations, social media graphics, videos, documents and more, serving individuals and large enterprises globally.
Industry: Enterprise Software
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Scaleup
Valuation: Decacorn (USD 10B+)
Funding: Series E+
Headquarters: Surry Hills, New South Wales, Australia
Founded: 2012
Glassdoor
Glassdoor: 3.9
WebsiteLinkedInGlassdoor