Senior Data Engineer, Data Lakehouse Infrastructure

TRM Labs
United States
Workplace: RemoteFull timeUSD 190,000 - 220,000 annuallyFunction: Solutions Engineering & Sales EngineeringExperience: 5+ yearsSkills: ["Python","SQL","Spark","Airflow","GCP","BigQuery","Dataproc","Kafka","Snowflake","Trino","Iceberg","Hudi","SparkSQL","Dataflow","Composer"]

Senior data engineer responsible for designing, implementing, and scaling core lakehouse data infrastructure on GCP, including data modeling, ingestion, query performance optimization, and metadata management. Work across teams to build high-performance, scalable data pipelines and governance frameworks using Spark, Trino, Iceberg, Hudi, Snowflake, and related tools.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
TRM Labs
TRM Labs
1 year ago

Senior Data Engineer, Data Lakehouse Infrastructure

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 18 hours agoStatus: Live

Job Summary

Senior data engineer responsible for designing, implementing, and scaling core lakehouse data infrastructure on GCP, including data modeling, ingestion, query performance optimization, and metadata management. Work across teams to build high-performance, scalable data pipelines and governance frameworks using Spark, Trino, Iceberg, Hudi, Snowflake, and related tools.
Location: United States
Workplace: Remote
Employment Type: Full time
Job Function: Solutions Engineering & Sales Engineering

Key Responsibilities

  • •Architect and scale a high-performance data lakehouse on GCP, leveraging technologies like StarRocks, Apache Iceberg, GCS, BigQuery, Dataproc, and Kafka.
  • •Design, build, and optimize distributed query engines such as Trino, Spark, or Snowflake to support complex analytical workloads.
  • •Implement metadata management in open table formats like Iceberg and data discovery frameworks for governance and observability using Iceberg compatible catalogs.
  • •Develop and orchestrate robust ETL/ELT pipelines using Apache Airflow, Spark, and GCP-native tools (e.g., Dataflow, Composer).
  • •Collaborate across departments, partnering with data scientists, backend engineers, and product managers to design and implement

Pay and Benefits

Salary: USD 190,000 - 220,000 annually

Key Requirements

  • •5+ years of experience in data or software engineering, with a focus on distributed data systems and cloud-native architectures.
  • •Proven experience building and scaling data platforms on GCP, including storage, compute, orchestration, and monitoring.
  • •Strong command of one or more query engines such as Trino, Presto, Spark, or Snowflake.
  • •Experience with modern table formats like Apache Hudi, Iceberg, or Delta Lake.
  • •Exceptional programming skills in Python, as well as adeptness in SQL or SparkSQL.
Experience:5+ yearsData engineeringCloudGCPDistributed data systems
Skills:PythonSQLSparkAirflowGCPBigQueryDataprocKafkaSnowflakeTrinoIcebergHudiSparkSQLDataflowComposer
Languages:English
Tech Stack:PythonSQLSparkTrinoIcebergHudiSnowflakeAirflowGCPBigQueryDataprocKafkaDataflowComposer

Company Brief

TRM Labs
Provides blockchain intelligence and crypto risk management solutions to help financial institutions, governments, and crypto firms detect fraud, trace illicit activity, and comply with regulatory requirements across digital asset ecosystems.
Industry: RegTech
Company Size: Large (251 to 1,000 employees)
Growth: Growth Stage Startup
Headquarters: San Francisco, United States
Founded: 2018
WebsiteLinkedIn