Long Horizon RL
John emphasizes the importance of rigorous evaluations during long horizon RL training to ensure model safety. He discusses potential risks of models developing incentives that could lead to unintended behaviors. Dwarkesh raises thought-provoking questions about the alignment of models and the potential for models to prioritize tasks in unexpected ways.In this clip
From this podcast

Dwarkesh Podcast
John Schulman (OpenAI Cofounder) - Reasoning, RLHF, & Plan for 2027 AGI
Related Questions
I have a question about the episode John Schulman (OpenAI Cofounder) - Reasoning, RLHF, & Plan for 2027 AGI and the clip Future Model Capabilities regarding creating content before or after training.
Is reinforcement learning a turning point for large language models (LLMs) and artificial intelligence (AI) as discussed in the episode Pieter Abbeel: Deep Reinforcement Learning | Lex Fridman Podcast #10 and the clip Hierarchical Learning Insights?