AI Research Scientist, Recursive Self Improvement, AI Safety and Reinforcement Learning

AMD
Santa Clara
Workplace: HybridFull timeUSD 178,500 - 306,000 annuallyFunction: Research & Scientific (R&D)Education: phdSkills: ["Critical thinking","Analytical reasoning","Collaboration","Communication","Measurement mindset"]

Research self-improving training loops for recursive self-improvement (RSI) under explicit governance, kill switches, and human oversight. Build theory- and systems-grounded evaluations to detect capability drift, Goodhart effects, and distributional shift in closed-loop training. Partner with RL scientists on intersections with policy optimization and preference learning, define red-team and monitoring protocols, and publish technical findings aligned with responsible deployment standards.

Loading

Loading job details...

Preparing the role view and application actions.

FursaFursa
AMD
AMD
2 months ago

AI Research Scientist, Recursive Self Improvement, AI Safety and Reinforcement Learning

✓ Verified Job

Canonical indexed version, validated from employer's careers page.

Source: Company careers pageValidated by: Fursa AI
Last checked: 7 hours agoStatus: Live

Job Summary

Research self-improving training loops for recursive self-improvement (RSI) under explicit governance, kill switches, and human oversight. Build theory- and systems-grounded evaluations to detect capability drift, Goodhart effects, and distributional shift in closed-loop training. Partner with RL scientists on intersections with policy optimization and preference learning, define red-team and monitoring protocols, and publish technical findings aligned with responsible deployment standards.
Location: Santa Clara
Workplace: Hybrid
Employment Type: Full time
Job Function: Research & Scientific (R&D)
Seniority: Sr. Manager level

Key Responsibilities

  • •Research self-improving training loops, including model-generated supervision, iterative distillation, self-critique, and automated curriculum updates with clear scope limits.
  • •Develop evaluations to measure capability drift, Goodhart effects, and distributional shift in closed-loop training.
  • •Partner with RL scientists on RSI-style objectives intersecting policy optimization and preference learning.
  • •Define red-team protocols and monitoring for RSI pilots, including documented rollback criteria before experiments affect shared infrastructure.
  • •Publish technical reports and align internal narratives with responsible deployment standards.

Pay and Benefits

Salary: USD 178,500 - 306,000 annually

Key Requirements

  • •Strong background in machine learning (ML), AI safety, reinforcement learning, or a related field, with publications or substantial work in iterative training and self-training.
  • •Experience with empirical safety evaluation, scalable oversight, or stress-testing of generative model training pipelines.
  • •Strong software skills for building controlled experimental harnesses and reproducible RSI microcosms.
  • •PhD in Computer Science, Machine Learning, or related field strongly preferred.
  • •Demonstrated ability to formalize assumptions, bound autonomy, and insist on counterfactual evaluation with concrete metrics.
Experience:Machine learningAI safetyReinforcement learningGenerative AI
Education:PhD / Doctorate in Computer Science, Machine Learning, or related field
Skills:Critical thinkingAnalytical reasoningCollaborationCommunicationMeasurement mindset
Tech Stack:Machine learningMLAI safetyReinforcement learningRecursive self-improvement (RSI)Synthetic dataSelf-playIterative distillationPolicy optimizationPreference learningCapability driftGoodhart effectsDistributional shiftRed-team protocolsMonitoringKill switchesHuman oversightGenerative model training pipelines

Company Brief

AMD
Designs and produces semiconductor products including CPUs, GPUs, and adaptive SoCs for consumer, enterprise, and embedded markets, competing across PCs, data centers, and gaming industries.
Industry: Electronics Manufacturing
Company Size: Enterprise (1,001+ employees)
Revenue: USD 1B+
Growth: Public Company
Valuation: Public Company (Market Cap in USD)
Funding: IPO / Publicly Listed
Headquarters: Santa Clara, United States
Founded: 1969
Glassdoor
Glassdoor: 3.9
WebsiteLinkedIn