Software Engineer Graduate (Data Arch - Data Ecosystem ) - 2026 (PhD)

TikTok
San Jose
Workplace: OnsiteFull timeFunction: Software EngineeringEducation: phdSkills: ["Collaboration","Debugging","Performance tuning","Problem-solving"]

Join the TikTok Data Ecosystem team to build storage and computing infrastructure for offline recommendation data serving over a billion users. You’ll design real-time and offline data architecture, develop scalable streaming Lakehouse systems for feature pipelines, and collaborate with ML platform teams on PyTorch training workflows and data formats. Own core components of the distributed storage and processing stack, including file formats, stream compaction, and metadata management.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
TikTok
TikTok
1 month ago

Software Engineer Graduate (Data Arch - Data Ecosystem ) - 2026 (PhD)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 1 hour agoStatus: Live

Job Summary

Join the TikTok Data Ecosystem team to build storage and computing infrastructure for offline recommendation data serving over a billion users. You’ll design real-time and offline data architecture, develop scalable streaming Lakehouse systems for feature pipelines, and collaborate with ML platform teams on PyTorch training workflows and data formats. Own core components of the distributed storage and processing stack, including file formats, stream compaction, and metadata management.
Location: San Jose
Workplace: Onsite
Employment Type: Full time
Job Function: Software Engineering
Seniority: Graduate level

Key Responsibilities

  • •Design and implement real-time and offline data architecture for large-scale recommendation systems.
  • •Build scalable, high-performance streaming Lakehouse systems supporting feature pipelines, model training, and real-time inference.
  • •Collaborate with ML platform teams to support PyTorch-based model training workflows and define data formats and access patterns.
  • •Own core components of the distributed storage and processing stack, including file format, stream compaction, and metadata management.

Key Requirements

  • •Completing or recently completed a PhD in Software Development, Computer Science, Computer Engineering, or a related technical discipline.
  • •Experience building large-scale distributed systems, ideally in storage, stream processing, or ML infrastructure.
  • •Solid understanding of Apache Flink internals, including hands-on state management, connectors, or UDFs.
  • •Familiarity with Lakehouse technologies such as Apache Paimon, Iceberg, Delta Lake, or Hudi, including incremental ingestion, schema evolution, and snapshot isolation.
  • •Ability to commit to an onboarding date by end of year 2026.
Experience:Distributed systemsStorageStream processingML infrastructureLakehouseRecommendation systems
Education:PhD / Doctorate in Software Development, Computer Science, Computer Engineering, or a related technical discipline
Skills:CollaborationDebuggingPerformance tuningProblem-solving
Tech Stack:Apache FlinkPyTorchApache PaimonIcebergDelta LakeHudiLakehouseParquetORCLanceJavaScalaC++UDFsState managementConnectorsStream compactionMetadata management

Company Brief

TikTok
Short-form video platform that lets users create, share, and discover entertainment content through algorithmic recommendations. It also offers advertising and creator tools for brands, influencers, and businesses.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Headquarters: Singapore, Singapore
Founded: 2016
WebsiteLinkedIn