Summer 2027 PhD Gen AI and Reinforcement Learning Research Intern

AMD
Santa Clara
Workplace: HybridInternshipUSD 91,520 - 137,280 annuallyFunction: Education & TrainingEducation: phdSkills: ["Collaboration","Documentation","Research","Experimentation","Analytical thinking"]

Conduct research and prototype methods for post-training large generative models, including policy optimization, preference learning, reward modeling, and credit assignment. Develop approaches to improve reasoning, code generation, tool use, and agentic behavior, then design controlled, verifiable experiments to study failure modes like reward hacking and training instability. Create evaluations for quality, robustness, safety, and real-world task performance, and collaborate on training, logging, rollout generation, and reproducibility while documenting results for technical reports.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
AMD
AMD
1 day ago

Summer 2027 PhD Gen AI and Reinforcement Learning Research Intern

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 13 hours agoStatus: Live

Job Summary

Conduct research and prototype methods for post-training large generative models, including policy optimization, preference learning, reward modeling, and credit assignment. Develop approaches to improve reasoning, code generation, tool use, and agentic behavior, then design controlled, verifiable experiments to study failure modes like reward hacking and training instability. Create evaluations for quality, robustness, safety, and real-world task performance, and collaborate on training, logging, rollout generation, and reproducibility while documenting results for technical reports.
Location: Santa Clara
Workplace: Hybrid
Employment Type: Internship
Job Function: Education & Training
Seniority: Intern level

Key Responsibilities

  • •Research and prototype post-training methods for large generative models.
  • •Explore policy optimization, preference learning, reward modeling, exploration, and credit-assignment techniques.
  • •Develop methods to improve reasoning, code generation, tool use, and agentic behavior.
  • •Design and run controlled experiments using verifiable, preference-based, or model-generated feedback.
  • •Collaborate on training, rollout generation, logging, and reproducibility, and document findings for technical reports and publications.

Pay and Benefits

Salary: USD 91,520 - 137,280 annually

Key Requirements

  • •Currently pursuing a PhD in Computer Science, Machine Learning, Artificial Intelligence, Electrical or Computer Engineering, or a related field.
  • •Strong knowledge of reinforcement learning and modern deep-learning methods.
  • •Experience implementing and evaluating machine-learning models using Python and PyTorch.
  • •Familiarity with LLM post-training, RLHF/RLAIF, preference optimization, or language and multimodal agents.
  • •Experience conducting reproducible experiments and analyzing empirical results.
Experience:AIMachine learningReinforcement learningLarge language modelsMultimodal learningGenerative AIDistributed systemsGPU computing
Education:PhD / Doctorate
Skills:CollaborationDocumentationResearchExperimentationAnalytical thinking
Languages:En-us
Tech Stack:PythonPyTorchLLMRLHFRLAIFGPU computingDistributed systems

Company Brief

AMD
Designs and produces semiconductor products including CPUs, GPUs, and adaptive SoCs for consumer, enterprise, and embedded markets, competing across PCs, data centers, and gaming industries.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1969
Glassdoor
Glassdoor: 3.9
WebsiteLinkedIn