Principal Scientist - Data Pipeline Engineer

Adobe Systems
San Jose, Seattle, San Francisco
Workplace: OnsiteFull timeUSD 206,300 - 388,000 annuallyFunction: Research & Scientific (R&D)Experience: 10+ yearsEducation: phdSkills: ["Communication","Cross-team collaboration","Technical leadership","Problem-solving","Data curation judgment"]

Architect and scale multimodal data processing pipelines that power Adobe Firefly’s multimodal foundation models (image, video, and audio). Build distributed, GPU-accelerated systems that transform billions of assets into training-ready data at scale. Optimize ingestion, processing, and delivery by eliminating bottlenecks, improving inference throughput, and designing scalable storage, indexing, and serving infrastructure. Partner with modeling teams to drive data curation and training outcomes.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Adobe Systems
Adobe Systems
1 month ago

Principal Scientist - Data Pipeline Engineer

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 4 hours agoStatus: Live

Job Summary

Architect and scale multimodal data processing pipelines that power Adobe Firefly’s multimodal foundation models (image, video, and audio). Build distributed, GPU-accelerated systems that transform billions of assets into training-ready data at scale. Optimize ingestion, processing, and delivery by eliminating bottlenecks, improving inference throughput, and designing scalable storage, indexing, and serving infrastructure. Partner with modeling teams to drive data curation and training outcomes.
Location: San Jose, Seattle, San Francisco
Workplace: Onsite
Employment Type: Full time
Job Function: Research & Scientific (R&D)
Seniority: Mid level

Key Responsibilities

  • •Architect and optimize large-scale distributed pipelines that process billions of images, video, and audio assets through ML workflows into training-ready data.
  • •Scale inference throughput via batching, parallelism, and hardware utilization to reduce cost and increase training data generation speed.
  • •Identify and eliminate bottlenecks across ingestion, processing, and delivery from storage/I/O to compute scheduling.
  • •Design scalable data infrastructure to reliably store, index, and serve billions of data points across large-scale databases, distributed storage, and high-throughput compute.
  • •Drive data curation for model training by partnering with modeling teams and acting as a hands-on technical leader bridging data engineering and applied ML.

Pay and Benefits

Salary: USD 206,300 - 388,000 annually

Key Requirements

  • •10+ years of experience in data engineering, ML infrastructure, or distributed systems at large scale (billions of records or assets).
  • •Strong software engineering background with hands-on expertise in distributed systems and large-scale data processing frameworks such as Ray or Spark.
  • •Proficiency in Python and strong experience in a systems-level language (C++, Rust, Go, or Java) with debugging skills in distributed/ML runtime environments.
  • •Deep knowledge of databases and storage systems at scale, including data lakes, indexing, and retrieval across billions of data points.
  • •Strong ML background, especially optimizing GPU inference pipelines for VLMs, LLMs, and other large models, plus experience with data curation for model training.
Experience:10+ yearsLarge-scale data engineeringML infrastructureDistributed systemsMultimodalGPU inference
Education:PhD / Doctorate
Skills:CommunicationCross-team collaborationTechnical leadershipProblem-solvingData curation judgment
Tech Stack:PythonC++RustGoJavaRaySparkGPUCPUDistributed storageData lakesIndexingRetrievalVLMsLLMs

Company Brief

Adobe Systems
Provides creative, marketing, and document management software and cloud services, including Photoshop, Illustrator, Acrobat, and the Adobe Experience Cloud, serving creative professionals, enterprises, and governments worldwide.
Industry: SaaS
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: San Jose, United States
Founded: 1982
Glassdoor
Glassdoor: 4.0
WebsiteLinkedInGlassdoor