Forward Deployed Engineer - LLM Post-training

Reflection AI
San Francisco, New York
Workplace: OnsiteFull timeFunction: Data Science & Machine LearningExperience: 3+ yearsSkills: ["Communication","Problem-solving","Ownership"]

Role as a core member of the Applied AI team responsible for fine-tuning open-weight language models for enterprise use, building evaluation harnesses, and deploying adapted models to production. You will work with customer data, design training runs (SFT, preference optimization, RLHF/DPO), and collaborate with research and customer teams to ensure measurable improvements and robust deployments.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Reflection AI
Reflection AI
4 months ago

Forward Deployed Engineer - LLM Post-training

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 17 hours agoStatus: Live

Job Summary

Role as a core member of the Applied AI team responsible for fine-tuning open-weight language models for enterprise use, building evaluation harnesses, and deploying adapted models to production. You will work with customer data, design training runs (SFT, preference optimization, RLHF/DPO), and collaborate with research and customer teams to ensure measurable improvements and robust deployments.
Location: San Francisco, New York
Workplace: Onsite
Employment Type: Full time
Job Function: Data Science & Machine Learning

Key Responsibilities

  • •Fine-tune Reflection's open-weight models for customer-specific use cases: prepare datasets, configure training runs (SFT, preference optimization, reinforcement fine-tuning), and iterate based on evals.
  • •Build and maintain evaluation infrastructure: design eval suites, curate test sets, establish baselines, and measure whether fine-tuned models actually improve on the tasks customers care about.
  • •Prepare training data from raw customer inputs: inspect data quality, clean and format datasets, identify adversarial or noisy samples, and build reproducible data pipelines.
  • •Debug and diagnose training and inference issues: interpret loss curves, catch data quality problems, and identify when training dynamics indicate something is wrong.
  • •Support end-to-end deployments of fine-tuned models across hybrid environments (public cloud, VPC, and on-premises), helping ensure inference performance and reliability in production.

Pay and Benefits

Perks:Health InsuranceDentalVisionLife InsuranceParental LeaveRelocationEquity

Key Requirements

  • •Applied ML experience with hands-on fine-tuning of language models, including preparing datasets, running training loops, evaluating results, and shipping a fine-tuned model; familiarity with SFT, DPO, RLHF, or similar techniques.
  • •Understanding of evaluation methodology: how to design evals, interpret training graphs, and assess true model improvement versus overfitting.
  • •Comfort with training infrastructure: GPUs, compute management, debugging common training failures; not afraid of stack traces from training loops.
  • •Strong software engineering fundamentals (Python); ability to write clean, reproducible code; experience with data pipelines and version control for datasets and experiments.
  • •3+ years of engineering experience with meaningful exposure to applied ML or ML engineering (e.g., MLE, Applied Scientist, Data Scientist who shipped models to production, or ML-focused SWE).
  • •Demonstrated ability and interest to work in customer-facing environments, translating domain requirements into training strategies.
  • •Self-starter with high agency and ownership, thriving in fast-paced startup environments.
Experience:3+ yearsApplied MLOpen weight modelsEnterprise AI
Skills:CommunicationProblem-solvingOwnership
Tech Stack:PythonGPUsSFTDPORLHFTraining pipelinesEvaluation harnesses

Company Brief

Reflection AI
Builds frontier autonomous AI systems focused on autonomous coding agents (product: Asimov) to create organizational superintelligence, founded by former DeepMind/Google researchers and hiring across SF, NYC, London.
Industry: AI & Machine Learning
Company Size: Small (11 to 50 employees)
Growth: Growth Stage Startup
Funding: Series A
Headquarters: San Francisco, United States
Founded: 2024
WebsiteLinkedInGlassdoor