Member of Technical Staff, Reinforcement Learning
San Mateo
Workplace: OnsiteFull timeFunction: Education & TrainingExperience: 2+ yearsSkills: ["Reinforcement learning","PyTorch","Transformers","Diffusion models","RLHF","PPO","DPO","TensorRT","LLM serving"]Role focusing on designing and implementing reinforcement learning pipelines for diffusion-based LLMs, developing reward models, and aligning model behavior with human intent at scale. You will optimize RL training (PPO, DPO, RLHF), work on data preprocessing and evaluation, and advance techniques for controlled text generation and scalable AI systems in a cutting-edge research environment.

