Software Engineer, Distributed Data Systems

Exa
San Francisco
Workplace: OnsiteFull timeUSD 150,000 - 300,000 annuallyFunction: Software EngineeringSkills: []

Data Engineer at Exa will design and build scalable data infrastructure powering web crawling, embedding model training, and real-time search. You’ll architect lakehouse-based pipelines, manage large-scale streaming, and optimize production-grade data systems across petabytes of data in a hands-on, autonomous role in a San Francisco in-person team.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Exa
Exa
8 months ago

Software Engineer, Distributed Data Systems

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 11 hours agoStatus: Live

Job Summary

Data Engineer at Exa will design and build scalable data infrastructure powering web crawling, embedding model training, and real-time search. You’ll architect lakehouse-based pipelines, manage large-scale streaming, and optimize production-grade data systems across petabytes of data in a hands-on, autonomous role in a San Francisco in-person team.
Location: San Francisco
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering

Key Responsibilities

  • •Architect and build the data infrastructure that powers crawling billions of pages, training embedding models, and serving real-time search.
  • •Design systems that scale to hundreds of petabytes with high reliability and low downtime.
  • •Develop and operate large-scale distributed data processing pipelines and streaming systems.
  • •Collaborate on embedding training infrastructure and data layer design for scalable analytics.
  • •Scale and optimize storage and processing components to support petabyte-scale data workflows.

Pay and Benefits

Salary: USD 150,000 - 300,000 annually
Equity and Bonus:Equity

Key Requirements

  • •Deep understanding of lakehouse architectures (Delta Lake, Iceberg, Hudi) and when to use them.
  • •Experience building and operating large-scale distributed data processing pipelines.
  • •Hands-on experience with streaming data systems (Kafka, Flink, or similar).
  • •Familiarity with Ray, Spark, or ClickHouse at production scale.
  • •An obsessive focus on reliability and building systems that don't page you at 3am.
Experience:AISearchData engineering
Tech Stack:Delta LakeIcebergHudiKafkaFlinkRaySparkClickHouseRAPIDSCuDFRust

Eligibility

Work Authorization:Sponsorship available.

Company Brief

Exa
Builds an embeddings-based neural search engine and web search API for AI applications, offering crawling, retrieval, and deep research tools to serve developers and AI agents with up-to-date web knowledge.
Industry: API Platforms
Company Size: Small (11 to 50 employees)
Growth: Scaleup
Valuation: USD 500M to 1B
Funding: Series B
Headquarters: San Francisco, United States
Founded: 2021
WebsiteLinkedIn