Capability Generalization and Objective Generalization
Evan discusses the phenomenon of capability generalization without objective generalization in optimization algorithms, highlighting the potential dangers and the impact on training RL agents. He shares insights from a recent study that demonstrated how agents learned powerful policies but navigated environments for the wrong purpose.In this clip
From this podcast

The Gradient
Evan Hubinger on Effective Altruism and AI Safety
Related Questions