Matt and Afra discuss the ongoing debate around diagnostic classifiers, highlighting the limitations of these classifiers in capturing specific linguistic information. Safra and Lopez propose alternative techniques to evaluate model representations, questioning the reliability of traditional probing tasks. The conversation sheds light on the need for caution when interpreting above-chance performance in probing tasks.