Published Aug 27, 2018
67 - GLUE: A Multi-Task Benchmark and Analysis Platform, with Sam Bowman
Sam Bowman delves into the GLUE benchmark, a pioneering framework for evaluating natural language understanding models, exploring its impact on model architecture, task selection, and the pivotal role of diagnostic datasets in assessing linguistic capabilities.

Topics covered
Popular Clips
Episode Highlights
Related Episodes

79 - The glass ceiling in NLP, with Natalie Schluter
Answers 383 questions

128 - Dynamic Benchmarking, with Douwe Kiela
Answers 383 questions

22 - Deep Multitask Learning for Semantic Dependency Parsing, with Noah Smith
Answers 383 questions
44 - Truly Low Resource NLP, with Anders Søgaard
Answers 383 questions

92 - Computational Humanities, with David Bamman
Answers 383 questions
63 - Neural Lattice Language Models, with Jacob Buckman
Answers 383 questions
52 - Sequence-to-Sequence Learning as Beam-Search Optimization, with Sam Wiseman
Answers 383 questions

81 - BlackboxNLP, with Afra Alishahi and Tal Linzen
Answers 383 questions

115 - AllenNLP, interviewing Matt Gardner
Answers 383 questions

64 - Neural Network Models for Sentence Pair Tasks, with Wuwei Lan and Wei Xu
Answers 383 questions

114 - Behavioral Testing of NLP Models, with Marco Tulio Ribeiro
Answers 383 questions106 - Ethical Considerations In NLP Research, with Emily Bender
Answers 383 questions
