Model Editing Insights

Peter discusses the adaptation of model editing methods to focus on deleting factual information from AI models, moving beyond traditional benchmarks. He highlights the importance of addressing dangerous knowledge within models and introduces evaluation criteria aimed at ensuring models fail to provide previously known answers. The conversation also touches on the implications of sophisticated attacks on AI systems and the evolving landscape of machine learning safety.