Deception vs Corrigible

Evan discusses the differences between deception and corrigible behavior in AI models. He explains how deception can be simpler to achieve during training, while corrigible behavior requires building a robust pointer to the correct objective. The podcast explores the challenges of understanding simplicity in neural networks and its implications for training processes.