Data Set Assumptions

Gábor challenges the implicit assumptions in training and test set splits, highlighting the limitations of widely used datasets like PTB in NLP. Walid acknowledges the importance of dataset construction efforts but points out the serious constraints of the PTB corpus. The discussion prompts a call for more diverse and challenging datasets in NLP research.