Published Dec 6, 2023
The DPO debate: Do we need RL for RLHF?
Join Nathan Lambert as he delves into the debate surrounding the necessity of reinforcement learning (RL) for optimizing human feedback in AI, contrasting it with Direct Preference Optimization (DPO) while critically examining both the theoretical and practical aspects.

Topics covered
Popular Clips
Episode Highlights
Related Episodes


A recipe for frontier model post-training
Answers 383 questions
RLHF: A thin line between useful and lobotomized
Answers 383 questions
Why we disagree on what open-source AI should be
Answers 383 questions
OpenAI's Model (behavior) Spec, RLHF transparency, and personalization questions
Answers 383 questions
Why reward models are still key to understanding alignment
Answers 383 questions
Open Language Models (OLMos) and the LLM landscape
Answers 383 questions
Where 2024’s “open GPT4” can’t match OpenAI’s
Answers 383 questions
Interconnects year in review: 2023
Answers 383 questions
The end of the "best open LLM"
Answers 383 questions
The koan of an open-source LLM
Answers 383 questions
Local LLMs, some facts some fiction
Answers 383 questions
A realistic path to robotic foundation models
Answers 383 questions
