Data Engineer Graduate (TikTok Recommendation Ecosystem Architecture) - 2027 Start (PhD)

TikTok
Singapore
Workplace: OnsiteFull timeFunction: Architecture & Urban PlanningEducation: phdSkills: ["Collaboration","Debugging","Performance tuning"]

Build real-time and offline data architecture for large-scale recommendation systems. You’ll design scalable streaming Lakehouse infrastructure that powers feature pipelines, model training, and real-time inference, while collaborating with ML platform teams to support PyTorch-based workflows. Own core distributed storage and processing components—covering unified training/inference state, file formats, stream compaction, and metadata management—using technologies like Flink and modern Lakehouse tables.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
TikTok
TikTok
1 day ago

Data Engineer Graduate (TikTok Recommendation Ecosystem Architecture) - 2027 Start (PhD)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 12 hours agoStatus: Live

Job Summary

Build real-time and offline data architecture for large-scale recommendation systems. You’ll design scalable streaming Lakehouse infrastructure that powers feature pipelines, model training, and real-time inference, while collaborating with ML platform teams to support PyTorch-based workflows. Own core distributed storage and processing components—covering unified training/inference state, file formats, stream compaction, and metadata management—using technologies like Flink and modern Lakehouse tables.
Location: Singapore
Workplace: Onsite
Employment Type: Full time
Job Function: Architecture & Urban Planning
Seniority: Graduate level

Key Responsibilities

  • •Design and implement real-time and offline data architecture for large-scale recommendation systems.
  • •Build scalable, high-performance streaming Lakehouse systems for feature pipelines, model training, and real-time inference.
  • •Collaborate with ML platform teams to support PyTorch-based model training workflows and design data formats/access patterns for large-scale samples and features.
  • •Own core components of a distributed storage and processing stack, from file format to stream compaction and metadata management.
  • •Build unified infrastructure integrating the training data base and training/inference state system for multimodal foundation models in search, recommendation, and advertising scenarios.

Key Requirements

  • •Completing or recently completing a PhD in Artificial Intelligence, Software Development, Computer Science, Computer Engineering, or a related discipline.
  • •Experience building large-scale distributed systems in storage, stream processing, or ML infrastructure.
  • •Hands-on understanding of Apache Flink internals, including state management, connectors, or UDFs.
  • •Familiarity with Lakehouse technologies (Apache Paimon, Iceberg, Delta Lake, or Hudi) around incremental ingestion, schema evolution, and snapshot isolation.
Experience:Large-scale distributed systemsStream processingML infrastructure
Education:PhD / Doctorate
Skills:CollaborationDebuggingPerformance tuning
Tech Stack:Apache FlinkPyTorchApache PaimonIcebergDelta LakeHudiLakehouseData lakesCachingDistributed computingGPU IOStreaming LakehouseFeature pipelinesParquetORCLanceJavaScalaC++HBase

Company Brief

TikTok
Short-form video platform that lets users create, share, and discover entertainment content through algorithmic recommendations. It also offers advertising and creator tools for brands, influencers, and businesses.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Headquarters: Singapore, Singapore
Founded: 2016
WebsiteLinkedIn