Member of Engineering (Pre-training / Data Research)

Poolside
United Kingdom, United States
Workplace: RemoteFull timeFunction: Education & TrainingSkills: ["Python","Llm","Gpu","Distributed systems","Data pipelines","Transformer","Tokenization","Data curation","Deduplication","Data mixing","Curriculum"]

Seeking an experienced ML/Data Engineer to improve pretraining datasets for Poolside models. You will design scalable data pipelines, generate diverse synthetic data, and run short experiments to optimize data quality, working closely with Pretraining, Posttraining, Evals, and Product. You’ll apply state-of-the-art dataset design research on large-scale GPU clusters in a distributed data workflow.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Poolside
Poolside
1 year ago

Member of Engineering (Pre-training / Data Research)

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 23 hours agoStatus: Live

Job Summary

Seeking an experienced ML/Data Engineer to improve pretraining datasets for Poolside models. You will design scalable data pipelines, generate diverse synthetic data, and run short experiments to optimize data quality, working closely with Pretraining, Posttraining, Evals, and Product. You’ll apply state-of-the-art dataset design research on large-scale GPU clusters in a distributed data workflow.
Location: United Kingdom, United States
Workplace: Remote
Employment Type: Full time
Job Function: Education & Training

Key Responsibilities

  • •Follow the latest research related to LLMs and data quality. Be familiar with the most relevant open-source datasets and models.
  • •Design and implement complex pipelines that can generate large amounts of data while maintaining high diversity and optimizing resources.
  • •Closely work with other teams such as Pretraining, Posttraining, Evals and Product to ensure short feedback loops on the quality of the models delivered.
  • •Suggest, conduct and analyze data ablations or training experiments to improve the quality of the datasets generated.

Pay and Benefits

Perks:Remote WorkHealth InsuranceEquipmentPaid LeaveHome OfficeWell-being

Key Requirements

  • •Strong ML and engineering background with experience in Large Language Models (LLMs) and data ablations.
  • •Excellent programming skills in Python.
  • •Experience with large-scale GPU clusters and distributed data pipelines.
  • •Familiarity with data curation, deduplication, data mixing, tokenization, curriculum and impact of data repetition.
  • •Ability to design and implement high-volume data generation pipelines and run time-bounded experiments.
Experience:AiMlNlpData
Skills:PythonLlmGpuDistributed systemsData pipelinesTransformerTokenizationData curationDeduplicationData mixingCurriculum
Languages:English
Tech Stack:PythonGPUDistributed data pipelinesLLMTransformerTokenizationData curationDeduplicationData mixingCurriculumModel trainingPretraining

Company Brief

Poolside
Builds foundation models, coding agents, and enterprise systems to automate and accelerate software development; offers on-prem/VPC deployments, developer tooling, and forward-deployed engineering for high‑security environments.
Industry: AI & Machine Learning
Company Size: Medium (51 to 250 employees)
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series B
Headquarters: San Francisco, United States
Founded: 2023
WebsiteLinkedIn