52 - Sequence-to-Sequence Learning as Beam-Search Optimization, with Sam Wiseman

Topics covered
Popular Clips
Episode Highlights
Gradient Challenges
Efficient gradient computation is a critical aspect of seq2seq models, especially when considering the limitations of past technologies. explains that before tools like Pytorch, the complexity of computing gradients was a significant concern. He notes that the backward pass can be made independent of the beam size, which enhances computational efficiency 1. elaborates on the beam search process, emphasizing the importance of learning to search better rather than just approximating the arcmax 2.
If I'm going to be searching, I might as well kind of learn to search better.
---
This approach ensures that updates occur each time a mistake is made, rather than relying solely on approximate methods.
  Â
Performance Insights
The runtime performance of seq2seq models varies significantly with different beam sizes and configurations. discusses how training and testing with different beam sizes can impact model performance, particularly in tasks like word ordering and dependency parsing 3. He highlights that models trained with larger beams perform poorly when evaluated with smaller beams due to early decision-making constraints.
If you train with a big beam, then your model doesn't have to be super confident early on.
---
Wiseman also notes that while the beam search process is computationally intensive, it scales sub-linearly with the beam size, making it feasible for practical applications.
Related Episodes

67 - GLUE: A Multi-Task Benchmark and Analysis Platform, with Sam Bowman
Answers 383 questions02 - Bidirectional Attention Flow for Machine Comprehension
Answers 383 questions09 - Learning to Generate Reviews and Discovering Sentiment
Answers 383 questions
63 - Neural Lattice Language Models, with Jacob Buckman
Answers 383 questions10 - A Syntactic Neural Model for General-Purpose Code Generation
Answers 383 questions

22 - Deep Multitask Learning for Semantic Dependency Parsing, with Noah Smith
Answers 383 questions
29 - Neural machine translation via binary code prediction, with Graham Neubig
Answers 383 questions05 - Transition-Based Dependency Parsing with Stack Long Short-Term Memory
Answers 383 questions

104 - Model Distillation, with Victor Sanh and Thomas Wolf
Answers 383 questions25 - Neural Semantic Parsing over Multiple Knowledge-bases
Answers 383 questions
