Reward Models Explained
Jeremie discusses the complexities of process and outcome reward models, emphasizing their finicky nature and the challenges in obtaining quality data. He introduces a novel approach using pre-trained models with reinforcement learning, where thinking tags create a structured environment for generating outputs. This method allows for efficient reward generation, particularly with quantifiable datasets like math and coding, leading to more reliable model performance.In this clip
From this podcast

Last Week in AI
#198 - DeepSeek R1 & Janus, Qwen1M & 2.5VL, OpenAI Agents
Related Questions