Member of Technical Staff - Post-Training and RL
X AI
Palo Alto
Workplace: OnsiteFull timeUSD 180,000 - 600,000 annuallyFunction: Education & TrainingSkills: ["Reinforcement learning","Alignment","RLHF","Reward modeling","DPO"]Join a hands-on, technically rigorous team tackling post-training and reinforcement learning challenges. You’ll work on reward modeling, RLHF/DPO, and RL to improve reasoning and real-world capabilities, with clarity on your first project before offer. A meritocratic environment values initiative, excellence, and concise knowledge sharing to advance truth-seeking AI at scale.

