AI Research Scientist, Reinforcement Learning (LLM) and Post-Training

AMD
Santa Clara
Workplace: HybridFull timeUSD 178,500 - 306,000 annuallyFunction: Data Science & Machine LearningEducation: phdSkills: ["Research","Analysis","Collaboration","Technical leadership","Evaluation"]

Develop and advance reinforcement learning methods for post-training large language models and code models used in engineering-adjacent tasks. Invent and analyze RL algorithms such as policy optimization, preference-based methods, exploration, credit assignment, and reward modeling, then run rigorous empirical studies. Design reward models, training recipes, and curricula, characterize failure modes, and collaborate with RL infrastructure to scale training and improve measurable task success while maintaining stability and safety.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
AMD
AMD
2 days ago

AI Research Scientist, Reinforcement Learning (LLM) and Post-Training

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 4 hours agoStatus: Live
Reposted: similar role first listed 1 month ago

Job Summary

Develop and advance reinforcement learning methods for post-training large language models and code models used in engineering-adjacent tasks. Invent and analyze RL algorithms such as policy optimization, preference-based methods, exploration, credit assignment, and reward modeling, then run rigorous empirical studies. Design reward models, training recipes, and curricula, characterize failure modes, and collaborate with RL infrastructure to scale training and improve measurable task success while maintaining stability and safety.
Location: Santa Clara
Workplace: Hybrid
Employment Type: Full time
Job Function: Data Science & Machine Learning

Key Responsibilities

  • •Research and develop RL methods for post-training LLMs and code models on structured engineering tasks with verifiable or preference-based feedback.
  • •Design reward models, curricula, and off-policy or on-policy training recipes for sparse, noisy, or expensive labels from experts and simulators.
  • •Characterize failure modes such as reward hacking, degenerate policies, and instability, and propose experimental mitigations.
  • •Collaborate with RL infrastructure engineers to scale training and define interfaces for rollout generation, logging, and reproducibility.
  • •Publish at top venues and provide internal technical leadership on the RL roadmap.

Pay and Benefits

Salary: USD 178,500 - 306,000 annually

Key Requirements

  • •Strong publication record in reinforcement learning or closely related machine learning areas.
  • •Hands-on experience training RL or preference-optimized models at non-trivial scale (GPUs, distributed jobs).
  • •Experience with LLM post-training, RLHF/RLAIF, or policy optimization for language or code agents.
  • •Fluency in RL theory and the practical path from ablation to production-scale training.
  • •PhD in Computer Science, Machine Learning, or a related field (strongly preferred).
Education:PhD / Doctorate in Computer Science, Machine Learning, or related field
Skills:ResearchAnalysisCollaborationTechnical leadershipEvaluation
Languages:En-us
Tech Stack:Reinforcement learningLLMLarge generative modelsPolicy optimizationPreference-based methodsExplorationCredit assignmentReward modelingReward misspecificationVariance reductionRLHFRLAIFGPUsDistributed jobsReward modelsOff-policy trainingOn-policy trainingReward hackingDegenerate policiesNeurIPS

Company Brief

AMD
Designs and produces semiconductor products including CPUs, GPUs, and adaptive SoCs for consumer, enterprise, and embedded markets, competing across PCs, data centers, and gaming industries.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1969
Glassdoor
Glassdoor: 3.9
WebsiteLinkedIn