Summer 2027 Master's AI Research, Reinforcement Learning and LLM Post-Training Intern

AMD
Santa Clara
Workplace: HybridInternshipFunction: Education & TrainingEducation: phdSkills: ["Reproducibility","Analysis","Collaboration","Documentation"]

Conduct reinforcement learning research for post-training language and code models, including policy optimization, preference learning, reward modeling, and exploration/credit-assignment methods. Build and run controlled, verifiable experiments using preference-based or simulator feedback, diagnose failure modes like reward hacking and policy degeneration, and develop evaluation methods aligned with engineering constraints. Partner with research and infrastructure teams on rollout generation, training, logging, and reproducibility while documenting results in technical reports and publications.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
AMD
AMD
1 day ago

Summer 2027 Master's AI Research, Reinforcement Learning and LLM Post-Training Intern

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 13 hours agoStatus: Live

Job Summary

Conduct reinforcement learning research for post-training language and code models, including policy optimization, preference learning, reward modeling, and exploration/credit-assignment methods. Build and run controlled, verifiable experiments using preference-based or simulator feedback, diagnose failure modes like reward hacking and policy degeneration, and develop evaluation methods aligned with engineering constraints. Partner with research and infrastructure teams on rollout generation, training, logging, and reproducibility while documenting results in technical reports and publications.
Location: Santa Clara
Workplace: Hybrid
Employment Type: Internship
Job Function: Education & Training
Seniority: Intern level

Key Responsibilities

  • •Research and prototype reinforcement learning methods for post-training language and code models.
  • •Explore policy optimization, preference learning, reward modeling, exploration, and credit-assignment techniques.
  • •Design and run controlled experiments using verifiable, preference-based, or simulator-generated feedback.
  • •Analyze failure modes such as reward hacking, policy degeneration, and training instability.
  • •Collaborate with research and infrastructure teams on rollout generation, training, logging, and reproducibility, and document findings for technical reports/publications.

Key Requirements

  • •Currently pursuing a PhD in Computer Science, Machine Learning, Electrical or Computer Engineering, or a related field.
  • •Knowledge of reinforcement learning and modern deep-learning methods.
  • •Experience implementing and evaluating machine-learning models using Python and frameworks such as PyTorch.
  • •Familiarity with LLM post-training and RLHF/RLAIF, preference optimization, or language and code agents.
  • •Experience conducting reproducible experiments and analyzing empirical results.
Education:PhD / Doctorate
Skills:ReproducibilityAnalysisCollaborationDocumentation
Languages:En-us
Tech Stack:PythonPyTorch

Company Brief

AMD
Designs and produces semiconductor products including CPUs, GPUs, and adaptive SoCs for consumer, enterprise, and embedded markets, competing across PCs, data centers, and gaming industries.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1969
Glassdoor
Glassdoor: 3.9
WebsiteLinkedIn