Relaxed Adversarial Training
Evan discusses the concept of relaxed adversarial training, which aims to train models to not just focus on what they do, but also why they do it. This approach incorporates a recursive aspect where models are trained to understand other models, helping us better comprehend complex AI systems.In this clip
From this podcast

The Gradient
Evan Hubinger on Effective Altruism and AI Safety
Related Questions