Model Learning Behavior
Sean discusses the risk of models using shortcuts in learning behavior, highlighting the importance of internal representations and the differences in validity of tests between humans and language models. Cameron adds insights on how trivial alterations can reveal the mechanisms models employ.In this clip
From this podcast

The Gradient
Cameron Jones & Sean Trott: Understanding, Grounding, and Reference in LLMs
Related Questions