Kathleen and Daniel discuss the concerning issues surrounding alignment, reinforcement learning with human feedback, and the ease of undoing fine tuning in frontier models. They highlight the challenges of mitigating short and long term risks and the need for more work in making large language models safe.