Principal Research Engineer, Model Training & Post-Training

Inflection AI
Palo Alto
Workplace: OnsiteFull timeUSD 400,000 - 550,000 annuallyFunction: Education & TrainingEducation: phdSkills: ["Transformer models","Distributed training","SFT","RLHF","DPO","GRPO","RLAIF","Reward modeling","Tool-use fine-tuning","Data quality","Evaluation design"]

Lead the end-to-end model-improvement loop for large-scale foundation models, from data curation and training through evaluation, post-training, release criteria, and production feedback. Own architecture, distillation, and alignment methods; drive large-scale distributed training (1,000+ GPUs) and data strategies, ensuring high-quality, enterprise-ready models with strong reliability and cost-performance balance.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
Inflection AI
Inflection AI
2 months ago

Principal Research Engineer, Model Training & Post-Training

âś“ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 2 hours agoStatus: Live

Job Summary

Lead the end-to-end model-improvement loop for large-scale foundation models, from data curation and training through evaluation, post-training, release criteria, and production feedback. Own architecture, distillation, and alignment methods; drive large-scale distributed training (1,000+ GPUs) and data strategies, ensuring high-quality, enterprise-ready models with strong reliability and cost-performance balance.
Location: Palo Alto
Workplace: Onsite
Employment Type: Full time
Job Function: Education & Training
Seniority: Sr. Manager level

Key Responsibilities

  • •Own the model-improvement roadmap across capability, reliability, emotional intelligence, tool use, safety, latency, cost, and enterprise readiness.
  • •Lead training and post-training strategy, including supervised fine-tuning, RLHF, DPO, GRPO, RLAIF, reward modeling, preference optimization, tool-use fine-tuning, distillation, synthetic data, and related methods.
  • •Drive model architecture and optimization decisions across modern transformer-based and hybrid architectures, including both training-time and inference-time performance.
  • •Lead large-scale training efforts on distributed GPU clusters, including systems operating at the scale of 1,000+ GPUs.
  • •Define and execute data strategy across data curation, mixture design, deduplication, decontamination, human-in-the-loop pipelines, preference data, evaluation data, synthetic data, and production feedback loops.

Pay and Benefits

Salary: USD 400,000 - 550,000 annually
Equity and Bonus:Equity
Perks:Health InsuranceDentalVision

Key Requirements

  • •Experience leading or contributing to large-scale LLM, multimodal, or foundation-model training or post-training programs.
Experience:AIMachine learningLLMFoundation-model
Education:PhD / Doctorate
Skills:Transformer modelsDistributed trainingSFTRLHFDPOGRPORLAIFReward modelingTool-use fine-tuningData qualityEvaluation design
Languages:English
Tech Stack:TransformersPyTorchGPU clustersDistributed trainingRLHFSFT

Company Brief

Inflection AI
Develops conversational AI and personal-assistant technologies, including the Pi chatbot, to enable more natural human–computer interaction and simplify access to information and productivity tools.
Industry: AI & Machine Learning
Company Size: Medium (51 to 250 employees)
Growth: Early Stage Startup
Headquarters: Palo Alto, United States
Founded: 2022
WebsiteLinkedIn