Staff Data Engineer

Harbor Compliance
United States
Workplace: RemoteFull timeUSD 172,000 - 215,000 annuallyFunction: Software EngineeringSkills: ["Ownership","Independent execution","Cross-functional collaboration","Comfort with ambiguity","Building without existing infrastructure"]

Build Harbor Compliance’s first end-to-end, AI-ready data platform by designing near real-time streaming pipelines, CDC ingestion, and event-driven architectures. Implement and operate vector database and embedding infrastructure to support semantic search and retrieval-augmented generation. Own ETL/ELT workflows, data reliability and observability, and data governance while partnering with BI, Product, Engineering, and Finance to deliver low-latency, self-serve and executive-ready data products.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Harbor Compliance
Harbor Compliance
1 day ago

Staff Data Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 12 hours agoStatus: Live

Job Summary

Build Harbor Compliance’s first end-to-end, AI-ready data platform by designing near real-time streaming pipelines, CDC ingestion, and event-driven architectures. Implement and operate vector database and embedding infrastructure to support semantic search and retrieval-augmented generation. Own ETL/ELT workflows, data reliability and observability, and data governance while partnering with BI, Product, Engineering, and Finance to deliver low-latency, self-serve and executive-ready data products.
Location: United States
Workplace: Remote
Employment Type: Full time
Job Function: Software Engineering
Seniority: Mid level

Key Responsibilities

  • •Design, build, and own near real-time data pipelines (CDC, streaming ingestion, event-driven architectures) for the platform’s backbone.
  • •Evaluate and implement vector database infrastructure and embedding pipelines to enable AI-augmented use cases (semantic search, RAG, AI agents).
  • •Build ELT/ETL pipelines ingesting data from HubSpot, financial systems, and HRIS for real-time and batch use cases.
  • •Architect the underlying warehouse/lakehouse as a system of record and build lightweight transformation layers (e.g., dbt) for analytics-ready datasets.
  • •Own pipeline reliability and observability (monitoring, automated failure alerting, and lineage tracking) and implement data governance practices.

Pay and Benefits

Salary: USD 172,000 - 215,000 annually
Perks:Health InsurancePaid LeaveParental Leave401kLearning Budget

Key Requirements

  • •7+ years of hands-on data engineering experience, including depth in streaming/event-driven systems beyond batch pipelines.
  • •Designing and building near real-time pipelines from scratch in production (e.g., Kafka, Kinesis, Flink, Debezium/CDC).
  • •Hands-on production experience with vector databases and embeddings (e.g., Zilliz, Pinecone, Weaviate, pgvector, Milvus).
  • •Advanced proficiency in SQL and Python.
  • •Working knowledge of cloud warehouse/lakehouse platforms (Snowflake, BigQuery, or Databricks) and dbt.
Experience:Data engineeringStreamingEvent-drivenB2B SaaSAIData platform
Skills:OwnershipIndependent executionCross-functional collaborationComfort with ambiguityBuilding without existing infrastructure
Tech Stack:KafkaKinesisFlinkDebeziumCDCVector databasesZillizPineconeWeaviatePgvectorMilvusSQLPythonSnowflakeBigQueryDatabricksDbtELTETLFivetran

Company Brief

Harbor Compliance
Provides compliance and licensing solutions for businesses and nonprofits, offering registration, licensing, renewals, and ongoing regulatory support through a combination of software and expert services.
Industry: RegTech
Company Size: Large (251 to 1,000 employees)
Growth: Established Company
Funding: Bootstrapped
Headquarters: Portland, United States
Founded: 2012
WebsiteLinkedIn