Published Mar 15, 2018

52 - Sequence-to-Sequence Learning as Beam-Search Optimization, with Sam Wiseman

Sam Wiseman delves into enhancing sequence-to-sequence models with pre-training and constrained training strategies, offering insights into efficient gradient computation, runtime performance, and the transformative potential of Beam Search Optimization in overcoming exposure and label bias in model training.
Episode Highlights
NLP Highlights logo

Popular Clips

Episode Highlights

  • Gradient Challenges

    Efficient gradient computation is a critical aspect of seq2seq models, especially when considering the limitations of past technologies. explains that before tools like Pytorch, the complexity of computing gradients was a significant concern. He notes that the backward pass can be made independent of the beam size, which enhances computational efficiency 1. elaborates on the beam search process, emphasizing the importance of learning to search better rather than just approximating the arcmax 2.

    If I'm going to be searching, I might as well kind of learn to search better.

    ---

    This approach ensures that updates occur each time a mistake is made, rather than relying solely on approximate methods.

       

    Performance Insights

    The runtime performance of seq2seq models varies significantly with different beam sizes and configurations. discusses how training and testing with different beam sizes can impact model performance, particularly in tasks like word ordering and dependency parsing 3. He highlights that models trained with larger beams perform poorly when evaluated with smaller beams due to early decision-making constraints.

    If you train with a big beam, then your model doesn't have to be super confident early on.

    ---

    Wiseman also notes that while the beam search process is computationally intensive, it scales sub-linearly with the beam size, making it feasible for practical applications.

Related Episodes