Published Jun 26, 2024
RLHF Roundup: Trying to get good at PPO, charting RLHF's impact, RewardBench retrospective, and a reward model competition
Delve into the latest advancements in AI with Nathan Lambert as he explores a groundbreaking reward model competition, evaluates the pivotal role of RewardBench in AI model assessments, and unpacks the transformative impact of Reinforcement Learning from Human Feedback on language models.

Topics covered
Popular Clips
Episode Highlights
Related Episodes


A recipe for frontier model post-training
Answers 383 questions
The DPO debate: Do we need RL for RLHF?
Answers 383 questions
Why reward models are still key to understanding alignment
Answers 383 questions
OpenAI's Model (behavior) Spec, RLHF transparency, and personalization questions
Answers 383 questions
Evaluations: Trust, performance, and price (bonus, announcing RewardBench)
Answers 383 questions
RLHF: A thin line between useful and lobotomized
Answers 383 questions
Where 2024’s “open GPT4” can’t match OpenAI’s
Answers 383 questions
Model merging lessons in The Waifu Research Department
Answers 383 questions
The end of the "best open LLM"
Answers 383 questions
Interconnects year in review: 2023
Answers 383 questions
How to cultivate a high-signal AI feed
Answers 383 questions
OLMoE and the hidden simplicity in training better foundation models
Answers 383 questions
Phi 3 and Arctic: Outlier LMs are hints
Answers 383 questions
