Member of Engineering (Synthetic Data Research)

Poolside
United States
Workplace: RemoteFull timeFunction: Education & TrainingSkills: ["Python","LLM","GPU clusters","Distributed data pipelines","Prompt engineering","Data quality"]

Hands-on ML engineering role focused on generating and refining large-scale synthetic datasets for pretraining AI models. You will design scalable data pipelines, leverage LLM research, and collaborate with Pretraining, Posttraining, Evals, and Product to ensure data quality and impact across training of large models.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Poolside
Poolside
7 months ago

Member of Engineering (Synthetic Data Research)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 23 hours agoStatus: Live

Job Summary

Hands-on ML engineering role focused on generating and refining large-scale synthetic datasets for pretraining AI models. You will design scalable data pipelines, leverage LLM research, and collaborate with Pretraining, Posttraining, Evals, and Product to ensure data quality and impact across training of large models.
Location: United States
Workplace: Remote
Employment Type: Full time
Job Function: Education & Training

Key Responsibilities

  • •Follow the latest research related to LLMs and synthetic data generation in particular. Be familiar with the most relevant open-source datasets and models.
  • •Design and implement complex pipelines that can generate large amounts of data while maintaining high diversity and optimizing the resources available.
  • •Closely work with other teams such as Pretraining, Posttraining, Evals and Product to ensure alignment on the quality of the models delivered.
  • •Continuously measure and refine the quality of the datasets being generated while validating the final data strategy through quantitative data ablation experiments.
  • •Deliver large, high-quality, and diverse synthetic datasets mixing natural language and code modalities to train best-in-class coding agents.

Pay and Benefits

Perks:Remote WorkHealth InsuranceEquipmentHome OfficeLearning Budget

Key Requirements

  • •Strong machine learning and engineering background.
  • •Experience with Large Language Models (LLM), including: Understanding of how LLMs learn; Data ablations and scaling laws; Post-training techniques; Training reasoning and agentic models.
  • •Experience with implementing cost-efficient, complex pipelines to generate synthetical datasets at scale optimizing for data quality, correctness, diversity, etc.
  • •Experience with evals tracking model capabilities (general knowledge, reasoning, math, coding, long-context, etc).
  • •Excellent programming skills in Python.
Experience:AIMachine learningData pipelines
Skills:PythonLLMGPU clustersDistributed data pipelinesPrompt engineeringData quality
Languages:English
Tech Stack:PythonLLMGPU clustersDistributed data pipelinesPrompt engineering

Company Brief

Poolside
Builds foundation models, coding agents, and enterprise systems to automate and accelerate software development; offers on-prem/VPC deployments, developer tooling, and forward-deployed engineering for high‑security environments.
Industry: AI & Machine Learning
Company Size: Medium (51 to 250 employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series B
Headquarters: San Francisco, United States
Founded: 2023
WebsiteLinkedIn