John Schulman (OpenAI Cofounder) - Reasoning, RLHF, & Plan for 2027 AGI

Topics covered
Popular Clips
Episode Highlights
Generalization
John Schulman discusses the fascinating ability of AI models to generalize across various tasks and domains. He highlights how models trained on English data can perform reasonably well in other languages, demonstrating a form of generalization that extends beyond the initial training data 1. Schulman also notes the rapid progression of AI capabilities since GPT-2, emphasizing the significant improvements achieved through post-training processes 2.
We've seen some version of this with multimodal data, where if you do text-only fine-tuning, you also get reasonable behavior with images.
---
These insights underscore the potential for AI to adapt and perform tasks beyond its original programming, suggesting a promising future for AI development.
Long Horizon RL
In exploring long horizon reinforcement learning (RL), Schulman emphasizes the importance of evaluating models for potential misbehavior and discontinuous jumps in capabilities 3. He suggests that while current models are not yet coherent over long periods, advancements in long horizon RL could unlock significant improvements in AI's ability to plan and execute complex tasks 4.
I would expect this capability would start to become clear when we start to look at long horizon tasks more.
---
Schulman also discusses the need for models to develop introspection and active learning skills, which could enhance their ability to learn and adapt during task execution 5.
Related Episodes


Paul Christiano - Preventing AI Takeover
Answers 383 questions

Demis Hassabis - Scaling, Superhuman AIs, AlphaZero atop LLMs, Rogue Nations Threat
Answers 383 questions

David Deutsch - AI, America, Fun, & Bayes
Answers 383 questions

Dario Amodei (Anthropic CEO) - $10 Billion Models, OpenAI, Scaling, & AGI in 2 years
Answers 383 questions

Holden Karnofsky - Transformative AI & Most Important Century
Answers 383 questions

Joe Carlsmith - Otherness and Control in The Age of AGI
Answers 383 questions

Steve Hsu - Intelligence, Embryo Selection, & The Future of Humanity
Answers 383 questions

Francois Chollet - LLMs won’t lead to AGI - $1,000,000 Prize to find true solution
Answers 383 questions

Sholto Douglas & Trenton Bricken - How to Build & Understand GPT-7's Mind
Answers 383 questions














