Detecting Mesa Optimizers
Daniel and Evan discuss the challenges of detecting the presence of a Mesa optimizer and the potential issues with curating training examples. They highlight the possibility of learning the wrong thing and the bias of neural networks towards simplicity. The analogy with evolution adds an interesting perspective to the conversation.In this clip
From this podcast

The Gradient
Evan Hubinger on Effective Altruism and AI Safety
Related Questions