Research, Mid-Training

Cognition
San Francisco
Workplace: OnsiteFull timeFunction: Education & TrainingSkills: ["Problem-solving","Communication","Collaboration","Autonomy","Initiative"]

Join our applied AI lab to own late-stage model training decisions, shaping data mix, annealing schedules, context length, and synthetic data strategies. This cross-cutting role blends research and engineering to improve mid-training performance, with impact on Devin and other systems, operating at the seam between pre-training and post-training in a small, highly skilled team.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Cognition
Cognition
5 months ago

Research, Mid-Training

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 5 hours agoStatus: Live
Reposted: similar role first listed 5 months ago

Job Summary

Join our applied AI lab to own late-stage model training decisions, shaping data mix, annealing schedules, context length, and synthetic data strategies. This cross-cutting role blends research and engineering to improve mid-training performance, with impact on Devin and other systems, operating at the seam between pre-training and post-training in a small, highly skilled team.
Location: San Francisco
Workplace: Onsite
Employment Type: Full time
Job Function: Education & Training
Seniority: Mid level

Key Responsibilities

  • •Data Mix and Quality Uplift: Design and iterate on high-quality data mixtures for late-stage and annealing training runs; develop principled methods for sourcing, filtering, and weighting data.
  • •Capability Injection: Drive targeted improvements in coding, mathematics, and long-horizon reasoning through curated data strategies and training interventions.
  • •Synthetic Data Research: Develop and evaluate synthetic data pipelines that generate training signal at scale and assess their limits and failure modes.
  • •Annealing and Schedule Design: Research and optimize multi-stage learning rate schedules, warmup strategies, and compute allocation across training phases.
  • •Context Length Extension: Research and implement methods to extend effective context length with strategies for encoding, data construction, and evaluation.

Key Requirements

  • •Deep familiarity with the LLM training pipeline end to end: pre-training data, optimization, architecture, and how mid-training and post-training interact.
  • •Hands-on experience with continual pre-training, annealing, or late-stage data mixing for large models.
  • •Strong intuition for data quality: what makes a dataset useful for training, how to filter and curate at scale, and how data mix choices compound across evals.
  • •Experience developing or evaluating synthetic data pipelines for capability improvement.
  • •Proficiency in Python and deep learning frameworks (PyTorch); comfortable debugging distributed training at scale.
Experience:AIMachine learningLLMResearchSoftware engineering
Skills:Problem-solvingCommunicationCollaborationAutonomyInitiative
Languages:English
Tech Stack:PythonPyTorchDistributed trainingLLMData pipelines

Company Brief

Cognition
Builds Devin, an autonomous AI software engineer and AI-native developer tools (Windsurf, DeepWiki) to automate software engineering workflows for enterprise customers.
Industry: Developer Tools
Company Size: Medium (51 to 250 employees)
Revenue: USD 50M to 100M
Growth: Scaleup
Valuation: Unicorn (USD 1B+)
Funding: Series C
Headquarters: San Francisco, United States
Founded: 2023
Glassdoor
Glassdoor: 4.7
WebsiteLinkedInGlassdoor