Published May 1, 2024
RLHF: A thin line between useful and lobotomized
Nathan Lambert explores the cutting-edge of Reinforcement Learning from Human Feedback, unveiling the potential of innovations like Kahneman-Tversky Optimization to revolutionize AI reasoning and capabilities. With a deep dive into the complexities of preference fine-tuning and chattiness dynamics, Lambert navigates the delicate balance between advanced AI performance and maintaining system integrity.

Topics covered
Popular Clips
Episode Highlights
Related Episodes


A recipe for frontier model post-training
Answers 383 questions
OpenAI's Model (behavior) Spec, RLHF transparency, and personalization questions
Answers 383 questions

Interviewing Ross Taylor on LLM reasoning, Llama fine-tuning, Galactica, agents
Answers 383 questions
The DPO debate: Do we need RL for RLHF?
Answers 383 questions
Where 2024’s “open GPT4” can’t match OpenAI’s
Answers 383 questions
Why reward models are still key to understanding alignment
Answers 383 questions
Open Language Models (OLMos) and the LLM landscape
Answers 383 questions
Local LLMs, some facts some fiction
Answers 383 questions
Phi 3 and Arctic: Outlier LMs are hints
Answers 383 questions
GPT-4o-mini changed ChatBotArena
Answers 383 questions
How to cultivate a high-signal AI feed
Answers 383 questions
