Model Learning Behavior

Sean discusses the risk of models using shortcuts in learning behavior, highlighting the importance of internal representations and the differences in validity of tests between humans and language models. Cameron adds insights on how trivial alterations can reveal the mechanisms models employ.