Research Scientist, Post-Training — Video Generation

Pika
Palo Alto
Workplace: OnsiteFull timeUSD 185,000 - 400,000 annuallyFunction: Research & Scientific (R&D)Experience: 5+ yearsSkills: ["Communication","Collaboration","Problem-solving"]

Join Pika as a Multimodal LLM Researcher to design and build real-time, multimodal generation and agentic platforms. Lead research on LLMs/VLMs/Audio LMs, develop diffusion-based architectures, curate datasets, and collaborate with engineering to deploy production-ready, real-time systems. On-site at the Palo Alto HQ to empower millions of creators with next-gen creative tech.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Pika
Pika
4 months ago

Research Scientist, Post-Training — Video Generation

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 2 hours agoStatus: Live

Job Summary

Join Pika as a Multimodal LLM Researcher to design and build real-time, multimodal generation and agentic platforms. Lead research on LLMs/VLMs/Audio LMs, develop diffusion-based architectures, curate datasets, and collaborate with engineering to deploy production-ready, real-time systems. On-site at the Palo Alto HQ to empower millions of creators with next-gen creative tech.
Location: Palo Alto
Workplace: Onsite
Employment Type: Full time
Job Function: Research & Scientific (R&D)

Key Responsibilities

  • •Lead and contribute to research focused on real-time, multimodal generation (text, image, video, audio) and orchestration of agentic platform infrastructure
  • •Design and prototype novel algorithms and architectures for high-fidelity, real-time multimodal synthesis and interactive experiences
  • •Focus on real-time model inference and synthesis across modalities
  • •Work on diffusion model distillation or diffusion-based world models for multimodal applications
  • •Train and finetune autoregressive and diffusion models in LLM/VLM/Audio LM contexts with a focus on real-time performance

Pay and Benefits

Salary: USD 185,000 - 400,000 annually
Equity and Bonus:Equity

Key Requirements

  • •5+ years of relevant experience in large language models, vision-language models, or audio language models, including research during graduate studies
  • •First-author publications in top conferences/journals (e.g., NeurIPS, CVPR, ICML, ICCV, SIGGRAPH, Interspeech)
  • •Deep expertise in language modeling (LLM), vision-language modeling (VLM), or audio language modeling (Audio LM)
  • •Strong experience with generative models (autoregressive and diffusion) and their real-time deployment
  • •Hands-on experience curating or augmenting large multimodal datasets
Experience:5+ yearsMultimodal AIResearchAcademic publications
Skills:CommunicationCollaborationProblem-solving
Tech Stack:PythonPyTorchTensorFlowDiffusion modelsLLMVLMAudio LM

Company Brief

Pika
Pika (Pika Labs) builds AI-driven tools to generate and edit videos from text prompts and images, enabling creators to produce cinematic, 3D, anime and stylized videos with a web-based platform and APIs.
Industry: AI & Machine Learning
Company Size: Medium (51 to 250 employees)
Growth: Growth Stage Startup
Valuation: USD 250M to 500M
Funding: Series B
Headquarters: Palo Alto, United States
Founded: 2023
WebsiteLinkedIn