Pretraining in Natural Language Processing

Richard discusses the power of pretraining in natural language processing, starting with word vectors and expanding to include encoders and decoders. He highlights the influential Elmo paper and the vision of a single model for all of NLP.