Sr. Research Engineer/Scientist (all levels), World Models

TikTok
San Jose
Workplace: OnsiteFull timeFunction: Research & Scientific (R&D)Education: mastersSkills: ["Cross-functional collaboration","Communication"]

Work on TikTok’s Vision-Applied Research team to build next-generation world models. Develop methods and infrastructure to train large-scale generative models using simulated and real-world multimodal datasets, with a focus on long-horizon temporal consistency, realistic physics, and complex object dynamics. Create interactive multi-modal data generation and long-context training pipelines for video generation, enabling real-time user and agent interaction with the model.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
TikTok
TikTok
1 month ago

Sr. Research Engineer/Scientist (all levels), World Models

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 3 minutes agoStatus: Live

Job Summary

Work on TikTok’s Vision-Applied Research team to build next-generation world models. Develop methods and infrastructure to train large-scale generative models using simulated and real-world multimodal datasets, with a focus on long-horizon temporal consistency, realistic physics, and complex object dynamics. Create interactive multi-modal data generation and long-context training pipelines for video generation, enabling real-time user and agent interaction with the model.
Location: San Jose
Workplace: Onsite
Employment Type: Full time
Job Function: Research & Scientific (R&D)
Seniority: Mid level

Key Responsibilities

  • •Develop large-scale, diverse, interactive multi-modal data generation pipelines.
  • •Build training pipelines for long-context interactive video generation models.
  • •Advance video generation models for long-horizon temporal consistency and realistic physical dynamics.
  • •Enable accurate object interactions and causal relationships from large-scale multimodal data.
  • •Support real-time interaction so users and agents can engage with the world model.

Key Requirements

  • •M.S. or Ph.D. in Computer Vision, Computer Graphics, Machine Learning, or equivalent experience.
  • •3+ years research experience in GenAI, multimodal foundation models, or Embodied AI.
  • •Demonstrated ability to communicate complex technical concepts and collaborate within cross-functional research teams.
  • •Experience in video generation and synthesis, efficient real-time diffusion models, 3D/physics-based simulation, and/or reinforcement learning for agentic interaction.
  • •Track record of first-author publications in venues including CVPR, ICLR, NeurIPS, SIGGRAPH, and ICML.
Experience:Generative AIMultimodal foundation modelsEmbodied AIVideo generationDiffusion modelsReinforcement learning3D simulation
Education:Master's in Computer Vision, Computer Graphics, Machine Learning
Skills:Cross-functional collaborationCommunication
Tech Stack:Generative modelsMultimodal datasetsVideo generationDiffusion modelsReinforcement learning3D/physics-based simulation

Company Brief

TikTok
Short-form video platform that lets users create, share, and discover entertainment content through algorithmic recommendations. It also offers advertising and creator tools for brands, influencers, and businesses.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Headquarters: Singapore, Singapore
Founded: 2016
WebsiteLinkedIn