Research Data Engineer

Relation
London
Workplace: OnsiteFull timeFunction: Research & Scientific (R&D)Education: bachelorsSkills: ["Communication","Collaboration","Ownership","Stakeholder management","Resilience"]

Design and build scalable data systems that power predictive models of cellular behavior. You’ll ingest multi-modal scientific data, optimize storage layouts and access patterns for analytical and ML training, and evolve cloud-native lake/lakehouse infrastructure. Own data versioning, lineage, and quality monitoring, build workflow orchestration for production and large batch jobs, and partner daily with data scientists, ML scientists, and research engineers to enable efficient model experimentation.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Relation
Relation
1 month ago

Research Data Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 17 hours agoStatus: Live

Job Summary

Design and build scalable data systems that power predictive models of cellular behavior. You’ll ingest multi-modal scientific data, optimize storage layouts and access patterns for analytical and ML training, and evolve cloud-native lake/lakehouse infrastructure. Own data versioning, lineage, and quality monitoring, build workflow orchestration for production and large batch jobs, and partner daily with data scientists, ML scientists, and research engineers to enable efficient model experimentation.
Location: London
Workplace: Onsite
Employment Type: Full time
Job Function: Research & Scientific (R&D)
Seniority: Mid level

Key Responsibilities

  • •Design, build, and maintain scalable data pipelines to ingest multi-modal scientific data.
  • •Optimize data movement, storage layouts, and access patterns for analytical and ML workloads.
  • •Stand up and evolve cloud-native data lake/lakehouse infrastructure.
  • •Implement data versioning, lineage, and quality monitoring.
  • •Build and operate workflow orchestration for production pipelines and large-scale batch jobs.

Key Requirements

  • •A degree in Computer Science, Engineering, or a related quantitative discipline, with significant industry experience in data engineering, MLOps, or data platform roles.
  • •Excellent Python engineering skills.
  • •Deep experience with cloud-native data infrastructure (AWS S3 / GCS) and Infrastructure-as-Code (Terraform or equivalent).
  • •Experience designing data pipelines and storage layouts for large, heterogeneous datasets.
  • •Experience building scalable analytical data processing workflows using frameworks like Spark, Polars, Dask, or DuckDB, with an understanding of performance trade-offs.
Education:Bachelor's
Skills:CommunicationCollaborationOwnershipStakeholder managementResilience
Tech Stack:PythonAWS S3GCSTerraformSparkPolarsDaskDuckDBAirflowDagsterPrefectDockerK8sParquetZarrTileDBHDF5Lakehouse

Company Brief

Relation
Relation (Relation Therapeutics) is a UK-based, technology-enabled biopharmaceutical company that combines single-cell multi-omics, patient-derived tissue, functional assays and machine learning to discover and develop novel therapeutics across immunology, metabolic and bone diseases.
Industry: Biotech
Company Size: Medium (51 to 250 employees)
Growth: Early Stage Startup
Valuation: USD 100M to 250M
Funding: Seed
Headquarters: London, United Kingdom
Founded: 2019
Glassdoor
Glassdoor: 3.8
WebsiteLinkedInGlassdoor