Summer 2027 PhD AI Research Infrastructure, RL Post-Training Intern

AMD
Santa Clara
Workplace: HybridInternshipFunction: Education & TrainingEducation: phdSkills: []

Build and optimize infrastructure for reinforcement learning-based post-training of large language and multimodal models. Develop scalable systems for rollout generation, inference, reward computation, and policy updates, while improving distributed training efficiency, reliability, and fault tolerance. Create tools for experiment configuration, checkpointing, logging, monitoring, and reproducibility, and help researchers turn experimental requirements into production-quality infrastructure.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
AMD
AMD
1 day ago

Summer 2027 PhD AI Research Infrastructure, RL Post-Training Intern

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 13 hours agoStatus: Live

Job Summary

Build and optimize infrastructure for reinforcement learning-based post-training of large language and multimodal models. Develop scalable systems for rollout generation, inference, reward computation, and policy updates, while improving distributed training efficiency, reliability, and fault tolerance. Create tools for experiment configuration, checkpointing, logging, monitoring, and reproducibility, and help researchers turn experimental requirements into production-quality infrastructure.
Location: Santa Clara
Workplace: Hybrid
Employment Type: Internship
Job Function: Education & Training
Seniority: Intern level

Key Responsibilities

  • •Develop and optimize infrastructure for RL-based post-training of large language and multimodal models.
  • •Build scalable systems for rollout generation, inference, reward computation, and policy updates.
  • •Improve distributed training efficiency, reliability, fault tolerance, and resource utilization.
  • •Design interfaces and tools that enable researchers to implement/evaluate new RL algorithms and manage experiments.
  • •Profile end-to-end training pipelines and resolve performance, memory, and communication bottlenecks.

Key Requirements

  • •Must be currently pursuing a PhD in Computer Science, Machine Learning, Artificial Intelligence, Computer Engineering, or a related field.
  • •Strong programming skills in Python and experience with PyTorch.
  • •Knowledge of reinforcement learning, LLM post-training, RLHF/RLAIF, or preference optimization.
  • •Experience with distributed training, multi-GPU workloads, or large-scale inference.
  • •Familiarity with parallelism strategies (data, tensor, pipeline, or sequence parallelism) and distributed training/research infrastructure.
Education:PhD / Doctorate in Computer Science, Machine Learning, Artificial Intelligence, Computer Engineering, or a related field
Languages:En-us
Tech Stack:PythonPyTorchRLHFRLAIFMultimodal modelsMulti-GPU workloadsDistributed trainingContainerizationCheckpointingLoggingMonitoringCloud computingCluster computingExperiment tracking

Company Brief

AMD
Designs and produces semiconductor products including CPUs, GPUs, and adaptive SoCs for consumer, enterprise, and embedded markets, competing across PCs, data centers, and gaming industries.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1969
Glassdoor
Glassdoor: 3.9
WebsiteLinkedIn