AI Research Scientist - Infrastructure Engineer, Reinforcement Learning
AMD
Santa Clara
Workplace: HybridFull timeUSD 178,500 - 306,000 annuallyFunction: DevOps, Cloud & InfrastructureEducation: bachelorsSkills: ["Communication","Collaboration","Reliability","Cost-quality tradeoff focus"]Own reinforcement learning infrastructure at scale, including distributed policy/value training, rollout generation, logging, checkpointing, and researcher-facing APIs across large GPU fleets. Improve RL scientists’ productivity by boosting throughput, fault tolerance, reproducibility, and observability, turning fragile experiments into reliable systems. Build and instrument high-throughput training stacks integrated with scheduling/storage, and drive reliability through on-call rotations, runbooks, and postmortems.

