Research Scientist, Data

Pika
Palo Alto
Workplace: HybridFull timeUSD 185,000 - 400,000 annuallyFunction: Data Science & Machine LearningExperience: 5+ yearsSkills: ["Cross-functional collaboration","Problem-solving","Communication"]

Architect and scale large-scale data engineering systems that power training for advanced multimodal foundation models. Own data pipeline architecture and implementation for text, image, audio, and video workflows, partnering with research and engineering teams to curate, clean, and manage high-quality datasets. Build tools for ingestion, labeling, filtering, augmentation, and storage, while ensuring data quality, reliability, compliance, and efficient processing for distributed training.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Pika
Pika
2 months ago

Research Scientist, Data

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 18 hours agoStatus: Live

Job Summary

Architect and scale large-scale data engineering systems that power training for advanced multimodal foundation models. Own data pipeline architecture and implementation for text, image, audio, and video workflows, partnering with research and engineering teams to curate, clean, and manage high-quality datasets. Build tools for ingestion, labeling, filtering, augmentation, and storage, while ensuring data quality, reliability, compliance, and efficient processing for distributed training.
Location: Palo Alto
Workplace: Hybrid
Employment Type: Full time
Job Function: Data Science & Machine Learning
Seniority: Sr. Manager level

Key Responsibilities

  • •Own large-scale data pipeline architecture and implementation to support model training and research workflows across text, image, audio, and video.
  • •Partner with research and engineering teams to curate, clean, and manage diverse multimodal datasets for pre-training and mid-training.
  • •Develop scalable strategies and tools for data ingestion, labeling, filtering, augmentation, and storage.
  • •Ensure data quality, reliability, and compliance, managing privacy and ethical considerations across the data lifecycle.
  • •Optimize data processing, transformation, and delivery for large-scale distributed training pipelines and productionize new dataset creation methods.

Pay and Benefits

Salary: USD 185,000 - 400,000 annually
Equity and Bonus:Equity
Perks:Health Insurance401k

Key Requirements

  • •5+ years building and scaling machine learning data pipelines at staff or lead engineer level, ideally in research or model training environments.
  • •Strong background in data engineering and ML data curation for LLMs, VLMs, or other large-scale multimodal models.
  • •Expertise in distributed data systems (e.g., Spark, Hadoop, Ray) and efficient large dataset processing/ETL workflows.
  • •Proven ability to build robust, scalable, production-grade data infrastructure for ML pipelines.
  • •Strong programming skills (Python, SQL, PySpark) and familiarity with cloud data platforms (AWS, GCP, Azure), plus experience with labeling, filtering, deduplication, and QA.
Experience:5+ yearsMachine learningData engineeringML data curationMultimodalGenerative AIFoundation models
Skills:Cross-functional collaborationProblem-solvingCommunication
Tech Stack:PythonSQLPySparkSparkHadoopRayETLAWSGCPAzureLLMsVLMsDistributed data systemsData ingestionData labelingData filteringDeduplicationData quality assuranceDataset management

Company Brief

Pika
Pika (Pika Labs) builds AI-driven tools to generate and edit videos from text prompts and images, enabling creators to produce cinematic, 3D, anime and stylized videos with a web-based platform and APIs.
Industry: AI & Machine Learning
Company Size: Medium (51 to 250 employees)
Growth: Growth Stage Startup
Valuation: USD 250M to 500M
Funding: Series B
Headquarters: Palo Alto, United States
Founded: 2023
WebsiteLinkedIn