Research Scientist - Data and State Acceleration - Global Frontier Tech Recruitment Program - 2027 Start (PhD)

TikTok
San Jose
Workplace: OnsiteFull timeFunction: Data Science & Machine LearningEducation: phdSkills: ["Collaboration","Debugging","Performance tuning"]

Build real-time and offline data architectures for large-scale recommendation systems, powering feature pipelines, model training, and real-time inference. Work on streaming Lakehouse systems and unified infrastructure that connects training data with training/inference state for multimodal foundation models. Partner with ML platform teams to support PyTorch training workflows, optimize data formats/access patterns, and own core distributed storage and processing components including file formats, compaction, and metadata management.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
TikTok
TikTok
1 month ago

Research Scientist - Data and State Acceleration - Global Frontier Tech Recruitment Program - 2027 Start (PhD)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 21 hours agoStatus: Live

Job Summary

Build real-time and offline data architectures for large-scale recommendation systems, powering feature pipelines, model training, and real-time inference. Work on streaming Lakehouse systems and unified infrastructure that connects training data with training/inference state for multimodal foundation models. Partner with ML platform teams to support PyTorch training workflows, optimize data formats/access patterns, and own core distributed storage and processing components including file formats, compaction, and metadata management.
Location: San Jose
Workplace: Onsite
Employment Type: Full time
Job Function: Data Science & Machine Learning
Seniority: Graduate level

Key Responsibilities

  • •Design and implement real-time and offline data architecture for large-scale recommendation systems
  • •Build scalable, high-performance streaming Lakehouse systems for feature pipelines, model training, and real-time inference
  • •Collaborate with ML platform teams to support PyTorch-based training workflows and define efficient data formats/access patterns
  • •Own core components of the distributed storage and processing stack, including file format, stream compaction, and metadata management
  • •Integrate training data with training/inference state for multimodal foundation models in search, recommendation, and advertising scenarios

Key Requirements

  • •Completing or recently completed a PhD in Software Development, Computer Science, Computer Engineering, or a related technical discipline
  • •Experience building large-scale distributed systems, preferably in storage, stream processing, or ML infrastructure
  • •Understanding of Apache Flink internals, including hands-on experience with state management, connectors, or UDFs
  • •Familiarity with Lakehouse technologies such as Apache Paimon, Iceberg, Delta Lake, or Hudi (incremental ingestion, schema evolution, snapshot isolation)
  • •Proficiency in Java/Scala/C++ and strong debugging/performance tuning ability
Education:PhD / Doctorate in Software Development, Computer Science, Computer Engineering, or related technical discipline
Skills:CollaborationDebuggingPerformance tuning
Tech Stack:PyTorchApache FlinkState managementConnectorsUDFsLakehouseApache PaimonIcebergDelta LakeHudiData lakesCachingDistributed computingGPU IOStreaming LakehouseData formatsFile formatStream compactionMetadata managementParquet

Company Brief

TikTok
Short-form video platform that lets users create, share, and discover entertainment content through algorithmic recommendations. It also offers advertising and creator tools for brands, influencers, and businesses.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Headquarters: Singapore, Singapore
Founded: 2016
WebsiteLinkedIn