Hattie Zhou: Lottery Tickets and Algorithmic Reasoning in LLMs

Topics covered
Popular Clips
Questions from this episode
- Asked by 38 people
Episode Highlights
Fortuitous Forgetting
The concept of fortuitous forgetting in neural networks explores how forgetting can aid learning and generalization. explains that, similar to humans, forgetting undesirable information can help models relearn and improve performance 1. This iterative process involves removing unwanted data and retraining the model to enhance desirable traits. notes that this approach can lead to faster grokking, a phenomenon where models suddenly improve their test accuracy after prolonged training 2.
The hypothesis is that when you do this, every time you retrain, you're starting from a better initialization that has less of the undesirable information and more of the desirable information.
---
This method aligns with other neural network phenomena like double descent, where models simplify over time to improve generalization.
Iterative Processes
Iterative training processes are crucial for enhancing model performance and understanding learning dynamics. discusses how these methods, though expensive, can be effective in refining models by repeatedly improving upon previous iterations 3. She highlights examples where models generate labels, fine-tune on them, and iteratively enhance their performance. This approach is seen as a paradigm for model improvement, especially in large language models.
The iterative process is kind of taking something that works and just multiplying it and almost bootstrapping the process, and it's easy.
---
Additionally, Zhou mentions that during training, weights may converge to zero but later become significant, indicating the complex dynamics of learning and forgetting 4.
Related Episodes


Jonathan Frankle: From Lottery Tickets to LLMs
Answers 383 questions

Zachary Lipton: Where Machine Learning Falls Short
Answers 383 questions

Hugo Larochelle: Deep Learning as Science
Answers 383 questions

Kyunghyun Cho: Neural Machine Translation, Language, and Doing Good Science
Answers 383 questions

Vera Liao: AI Explainability and Transparency
Answers 383 questions

Joel Lehman: Open-Endedness and Evolution through Large Models
Answers 383 questions

Subbarao Kambhampati: Planning, Reasoning, and Interpretability in the Age of LLMs
Answers 383 questions

Sara Hooker: Cohere For AI, the Hardware Lottery, and DL Tradeoffs
Answers 383 questions

Been Kim: Interpretable Machine Learning
Answers 383 questions

Thomas Dietterich: From the Foundations
Answers 383 questions

Melanie Mitchell: Abstraction and Analogy in AI
Answers 383 questions

Greg Yang on Communicating Research, Tensor Programs, and µTransfer
Answers 383 questions

Chip Huyen: Machine Learning Tools and Systems
Answers 383 questions

Eric Jang on Robots Learning at Google and Generalization via Language
Answers 383 questions

Seth Lazar: Normative Philosophy of Computing
Answers 383 questions
