Senior Research Data Engineer

Canva
Vienna
Workplace: HybridFull timeFunction: Research & Scientific (R&D)Skills: ["Communication","Collaboration","Ownership","Iteration"]

Own the data foundations that power multimodal agent research, building pipelines, datasets, and tooling from collection and curation through preprocessing, quality assurance, and delivery into training workflows. Partner with research scientists to translate requirements into data specs, create evaluation benchmarks, develop human-annotation and synthetic/preference-data tooling (RLHF/DPO), and ensure trustworthy, reproducible datasets through monitoring and test coverage.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Canva
Canva
2 days ago

Senior Research Data Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 9 minutes agoStatus: Live

Job Summary

Own the data foundations that power multimodal agent research, building pipelines, datasets, and tooling from collection and curation through preprocessing, quality assurance, and delivery into training workflows. Partner with research scientists to translate requirements into data specs, create evaluation benchmarks, develop human-annotation and synthetic/preference-data tooling (RLHF/DPO), and ensure trustworthy, reproducible datasets through monitoring and test coverage.
Location: Vienna
Workplace: Hybrid
Employment Type: Full time
Job Function: Research & Scientific (R&D)
Seniority: Mid level

Key Responsibilities

  • •Design and build data pipelines for agent training, including collection, filtering, deduplication, formatting, and versioning across text, image, and multimodal sources.
  • •Build and maintain infrastructure for efficient data loading, storage, and retrieval at scale (S3, distributed systems, streaming pipelines).
  • •Collaborate with research scientists to translate research requirements into concrete data specifications and iterate as experiments evolve.
  • •Create evaluation datasets and benchmarks by curating task distributions that surface real failure modes.
  • •Own data quality by building validation frameworks, monitoring for drift/contamination, and ensuring datasets are trustworthy and reproducible.

Key Requirements

  • •Strong software engineering skills in Python and experience building production-grade data pipelines for ML workflows.
  • •Experience with ML data workflows, including large-scale data processing/loading (Ray or similar), data versioning, and training data formatting considerations.
  • •Hands-on experience supporting data pipelines for large-scale distributed ML training runs.
  • •Familiarity with annotation tooling and human-in-the-loop data collection (Label Studio or internal systems).
  • •Ability to scope ambiguous problem statements with researchers and translate needs into actionable data plans.
Skills:CommunicationCollaborationOwnershipIteration
Languages:English
Tech Stack:PythonRayAWSS3Streaming pipelinesLLMVLMTokenizationBatchingShardingLabel StudioRLHFDPO

Company Brief

Canva
Canva is an online visual communications platform that provides drag-and-drop design tools, templates, and media assets for creating presentations, social media graphics, videos, documents and more, serving individuals and large enterprises globally.
Industry: Enterprise Software
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Scaleup
Valuation: Decacorn (USD 10B+)
Funding: Series E+
Headquarters: Surry Hills, New South Wales, Australia
Founded: 2012
Glassdoor
Glassdoor: 3.9
WebsiteLinkedInGlassdoor