Research Engineer/Scientist (all levels), World Models

TikTok
San Jose
Workplace: OnsiteFull timeFunction: Research & Scientific (R&D)Skills: ["Communication","Cross-functional collaboration"]

Build next-generation world models for TikTok’s Vision-Applied Research team, focusing on long-horizon temporal consistency, realistic physics, and complex dynamics. Develop large-scale multimodal data generation and training pipelines for long-context interactive video generation models. Advance video generation to capture causal object interactions, and enable real-time interaction between users/agents and the model. Collaborate with cross-functional research teams and publish in top venues.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
TikTok
TikTok
1 month ago

Research Engineer/Scientist (all levels), World Models

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 5 hours agoStatus: Live

Job Summary

Build next-generation world models for TikTok’s Vision-Applied Research team, focusing on long-horizon temporal consistency, realistic physics, and complex dynamics. Develop large-scale multimodal data generation and training pipelines for long-context interactive video generation models. Advance video generation to capture causal object interactions, and enable real-time interaction between users/agents and the model. Collaborate with cross-functional research teams and publish in top venues.
Location: San Jose
Workplace: Onsite
Employment Type: Full time
Job Function: Research & Scientific (R&D)

Key Responsibilities

  • •Develop large-scale, diverse, interactive multimodal data generation pipelines.
  • •Build training pipelines for long-context interactive video generation models.
  • •Advance video generation models for long-horizon temporal consistency and realistic physical dynamics.
  • •Improve modeling of object interactions and causal relationships from large-scale multimodal data.
  • •Enable users and agents to interact with the world model in real time.

Key Requirements

  • •M.S. or Ph.D. in Computer Vision, Computer Graphics, Machine Learning, or equivalent experience.
  • •Extensive research experience in generative AI, multimodal foundation models, or embodied AI.
  • •Ability to communicate complex technical concepts and collaborate effectively in cross-functional research teams.
  • •Experience in video generation and synthesis, efficient real-time diffusion models, 3D/physics-based simulation, and/or reinforcement learning for agentic interaction.
  • •First-author publications in top venues such as CVPR, ICLR, NeurIPS, SIGGRAPH, and ICML.
Experience:Generative AIMultimodal foundation modelsEmbodied AIVideo generationDiffusion modelsReinforcement learningPhysics simulationWorld models
Education:
Skills:CommunicationCross-functional collaboration
Tech Stack:Generative AIMultimodal foundation modelsEmbodied AIWorld modelsMultimodal datasetsMultimodal data generationInteractive video generationLong-context video generationTemporal consistencyDiffusion modelsReal-time diffusion3D/physics-based simulationReinforcement learningAgentic environment interactionVideo synthesisObject interactionsCausal relationships

Company Brief

TikTok
Short-form video platform that lets users create, share, and discover entertainment content through algorithmic recommendations. It also offers advertising and creator tools for brands, influencers, and businesses.
Industry: Digital Media
Company Size: Enterprise (1,001+ employees)
Growth: Scaleup
Headquarters: Singapore, Singapore
Founded: 2016
WebsiteLinkedIn