Data Engineer

Cusp.AI
Cambridge, Amsterdam, London, Berlin
Workplace: HybridFull timeFunction: Software EngineeringExperience: 3+ yearsSkills: ["Python","SQL","Airflow","Prefect","Dagster","Flyte","Docker","Kubernetes","CI/CD","Data pipelines"]

Design, build, and maintain scalable data pipelines that ingest, clean, and standardize diverse chemical datasets to support ML research. Collaborate with ML researchers and chemists to ensure data quality and accuracy, enabling self-serve access for the scientific team. Leverage Python, databases, and orchestration/CI/CD tools to support real-time AI-driven materials discovery.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Cusp.AI
Cusp.AI
3 months ago

Data Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 1 hour agoStatus: Live

Job Summary

Design, build, and maintain scalable data pipelines that ingest, clean, and standardize diverse chemical datasets to support ML research. Collaborate with ML researchers and chemists to ensure data quality and accuracy, enabling self-serve access for the scientific team. Leverage Python, databases, and orchestration/CI/CD tools to support real-time AI-driven materials discovery.
Location: Cambridge, Amsterdam, London, Berlin
Workplace: Hybrid
Employment Type: Full time
Job Function: Software Engineering

Key Responsibilities

  • •Design and build robust data pipelines for materials science datasets, experimental results, and computational chemistry outputs.
  • •Develop processes to integrate diverse data sources including materials databases, literature, patent filings, and laboratory instruments.
  • •Create automated workflows for processing crystallographic data, molecular structures, and materials properties.
  • •Build scalable systems to handle high-throughput computational chemistry calculations and experimental data.
  • •Partner closely with the scientific and research teams to implement automated quality checks for crystal structure data, chemical compositions, and experimental measurements.

Pay and Benefits

Equity and Bonus:Equity
Perks:EquityLearning Budget

Key Requirements

  • •3+ years in data engineering, preferably in scientific or research environments; able to work autonomously and provide guidance on best practices
  • •High proficiency in Python and databases with experience in large-scale data processing; regular programming, not just scripting
  • •Advanced user of workflow orchestration tools (Airflow, Prefect, Dagster, Flyte or similar)
  • •Solid experience with containerisation (Docker, Kubernetes) and CI/CD practices
  • •Direct experience handling large/complex datasets and interest in scientific packages
Experience:3+ yearsMaterials scienceAIData engineering
Skills:PythonSQLAirflowPrefectDagsterFlyteDockerKubernetesCI/CDData pipelines
Languages:English
Tech Stack:PythonSQLAirflowPrefectDagsterFlyteDockerKubernetes

Company Brief

Cusp.AI
CuspAI builds AI foundation models and simulation tools to accelerate discovery and design of novel materials for applications like batteries, semiconductors, water treatment, and carbon capture, combining generative models with physics-based simulation.
Industry: Materials Science
Company Size: Small (11 to 50 employees)
Growth: Growth Stage Startup
Valuation: USD 500M to 1B
Funding: Series A
Headquarters: Cambridge, United Kingdom
Founded: 2024
WebsiteLinkedInGlassdoor