AI Research Scientist - Infrastructure Engineer, Reinforcement Learning

AMD
Santa Clara
Workplace: HybridFull timeUSD 178,500 - 306,000 annuallyFunction: DevOps, Cloud & InfrastructureEducation: bachelorsSkills: ["Communication","Collaboration","Reliability","Cost-quality tradeoff focus"]

Own reinforcement learning infrastructure at scale, including distributed policy/value training, rollout generation, logging, checkpointing, and researcher-facing APIs across large GPU fleets. Improve RL scientists’ productivity by boosting throughput, fault tolerance, reproducibility, and observability, turning fragile experiments into reliable systems. Build and instrument high-throughput training stacks integrated with scheduling/storage, and drive reliability through on-call rotations, runbooks, and postmortems.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
AMD
AMD
2 months ago

AI Research Scientist - Infrastructure Engineer, Reinforcement Learning

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 7 hours agoStatus: Live

Job Summary

Own reinforcement learning infrastructure at scale, including distributed policy/value training, rollout generation, logging, checkpointing, and researcher-facing APIs across large GPU fleets. Improve RL scientists’ productivity by boosting throughput, fault tolerance, reproducibility, and observability, turning fragile experiments into reliable systems. Build and instrument high-throughput training stacks integrated with scheduling/storage, and drive reliability through on-call rotations, runbooks, and postmortems.
Location: Santa Clara
Workplace: Hybrid
Employment Type: Full time
Job Function: DevOps, Cloud & Infrastructure
Seniority: Sr. Manager level

Key Responsibilities

  • •Design and implement distributed RL training stacks (data parallel, pipeline parallel, or hybrid) integrated with AMD schedulers and storage.
  • •Build high-throughput rollout workers, trajectory stores, and reward computation pipelines with versioning and audit trails.
  • •Instrument jobs for debugging (NaNs, stragglers, OOMs), and implement autoscaling and preemption-safe checkpointing.
  • •Collaborate with research scientists on experiment templates, hyperparameter sweeps, and safe promotion paths for broader team use.
  • •Drive reliability through on-call rotations, runbooks, and postmortems for infra incidents impacting RL training.

Pay and Benefits

Salary: USD 178,500 - 306,000 annually

Key Requirements

  • •Bachelor’s degree required; Master’s or PhD preferred in Computer Science for research-heavy collaboration.
  • •Strong systems track record in machine learning platforms with deep systems expertise and demonstrated technical impact.
  • •Deep experience with PyTorch (or JAX), NCCL/MPI-style distributed training, and GPU cluster orchestration.
  • •Experience owning RL training infrastructure, LLM post-training pipelines, or large-scale experiment management.
  • •Proficiency in C++/Python performance tuning, I/O optimization, and containerized workloads.
Experience:Machine learningReinforcement learningDistributed trainingGPU clustersLLM post-training
Education:Bachelor's in Computer Science
Skills:CommunicationCollaborationReliabilityCost-quality tradeoff focus
Languages:English
Tech Stack:Reinforcement learningPyTorchJAXNCCLMPIGPU cluster orchestrationC++PythonAutoscalingCheckpointingContainerized workloadsLoggingGPU fleets

Company Brief

AMD
Designs and produces semiconductor products including CPUs, GPUs, and adaptive SoCs for consumer, enterprise, and embedded markets, competing across PCs, data centers, and gaming industries.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1969
Glassdoor
Glassdoor: 3.9
WebsiteLinkedIn