Data Set Assumptions
Gábor challenges the implicit assumptions in training and test set splits, highlighting the limitations of widely used datasets like PTB in NLP. Walid acknowledges the importance of dataset construction efforts but points out the serious constraints of the PTB corpus. The discussion prompts a call for more diverse and challenging datasets in NLP research.In this clip
From this podcast

NLP Highlights
40 - On the State of the Art of Evaluation in Neural Language Models, with Gábor Melis
Related Questions