Research Engineer, Data Lake Infrastructure & Data Analytics Graduate (AML-Ark-US) - 2027 Start (PhD)

ByteDance
Seattle
Workplace: OnsiteFull timeFunction: Data Analytics & Business IntelligenceEducation: phdSkills: []

Build and maintain the data lake that powers large model and agent development, including storage design, query performance, and cost efficiency. Develop large-scale batch and streaming pipelines to transform raw model/agent logs into analysis-ready datasets. Create analytics and metrics platforms so algorithm and product teams can quickly find actionable insights. Own data quality and governance and partner to support a continuously improving data flywheel through data-driven feedback loops.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
ByteDance
ByteDance
2 days ago

Research Engineer, Data Lake Infrastructure & Data Analytics Graduate (AML-Ark-US) - 2027 Start (PhD)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 14 hours agoStatus: Live

Job Summary

Build and maintain the data lake that powers large model and agent development, including storage design, query performance, and cost efficiency. Develop large-scale batch and streaming pipelines to transform raw model/agent logs into analysis-ready datasets. Create analytics and metrics platforms so algorithm and product teams can quickly find actionable insights. Own data quality and governance and partner to support a continuously improving data flywheel through data-driven feedback loops.
Location: Seattle
Workplace: Onsite
Employment Type: Full time
Job Function: Data Analytics & Business Intelligence
Seniority: Graduate level

Key Responsibilities

  • •Build and maintain the data lake that powers model and agent development, including storage design, query performance, and cost efficiency.
  • •Develop large-scale batch and streaming pipelines that turn raw model and agent logs into analysis-ready datasets.
  • •Build analytics and metrics platforms that help algorithm and product teams find actionable insights quickly.
  • •Own data quality and governance and partner with algorithm teams to support a continuously improving data flywheel behind system improvement.
  • •Support ongoing data-driven feedback loops that inform model improvement, system optimization, and product decisions.

Key Requirements

  • •Completing or recently completed a PhD in Computer Science, Software Engineering, Data Science, or a related field.
  • •Strong fundamentals in data structures, algorithms, and distributed systems.
  • •Strong programming skills in at least one mainstream language, plus proficiency in SQL.
  • •Hands-on experience with a big data engine or lakehouse storage system.
  • •Demonstrated ability through substantial projects, internships, research, or open-source work.
Experience:LLMsAI agentsBig dataLakehouseStreaming pipelinesOpen source
Education:PhD / Doctorate in Computer Science, Software Engineering, Data Science, or a related field
Tech Stack:SQLData lakeLakehouseColumnar storageQuery optimizationBatch pipelinesStreaming pipelinesOLAPDistributed systemsCloud-native infrastructure

Company Brief

ByteDance
Develops consumer internet and content platforms, including TikTok and other apps for short-form video, news, and entertainment. It also builds advertising, commerce, and creator tools that connect audiences, brands, and publishers across global markets.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Established Company
Headquarters: Beijing, China
Founded: 2012
WebsiteLinkedIn