Published May 15, 2024

John Schulman (OpenAI Cofounder) - Reasoning, RLHF, & Plan for 2027 AGI

John Schulman, cofounder of OpenAI, delves into AI's future as proactive collaborators, discussing the need for regulation and oversight, the potential of AI in task generalization and reinforcement learning, and the safety concerns surrounding their deployment.
Episode Highlights
Dwarkesh Podcast logo

Popular Clips

Episode Highlights

  • Generalization

    John Schulman discusses the fascinating ability of AI models to generalize across various tasks and domains. He highlights how models trained on English data can perform reasonably well in other languages, demonstrating a form of generalization that extends beyond the initial training data 1. Schulman also notes the rapid progression of AI capabilities since GPT-2, emphasizing the significant improvements achieved through post-training processes 2.

    We've seen some version of this with multimodal data, where if you do text-only fine-tuning, you also get reasonable behavior with images.

    ---

    These insights underscore the potential for AI to adapt and perform tasks beyond its original programming, suggesting a promising future for AI development.

       

    Long Horizon RL

    In exploring long horizon reinforcement learning (RL), Schulman emphasizes the importance of evaluating models for potential misbehavior and discontinuous jumps in capabilities 3. He suggests that while current models are not yet coherent over long periods, advancements in long horizon RL could unlock significant improvements in AI's ability to plan and execute complex tasks 4.

    I would expect this capability would start to become clear when we start to look at long horizon tasks more.

    ---

    Schulman also discusses the need for models to develop introspection and active learning skills, which could enhance their ability to learn and adapt during task execution 5.

Related Episodes