Long Horizon RL

John emphasizes the importance of rigorous evaluations during long horizon RL training to ensure model safety. He discusses potential risks of models developing incentives that could lead to unintended behaviors. Dwarkesh raises thought-provoking questions about the alignment of models and the potential for models to prioritize tasks in unexpected ways.