Zachary Lipton: Where Machine Learning Falls Short

Topics covered
Popular Clips
Episode Highlights
Pretraining Insights
Pretraining models on seemingly nonsensical data can yield surprising benefits. explains that even when models are pretrained on random phonemes and special tokens, they can achieve a significant portion of the performance benefits typically associated with large upstream corpora 1. This raises questions about the actual mechanisms behind pretraining benefits, as notes, "We get out a huge chunk, often like the majority of the benefit, without having used any upstream data."
We get out a huge chunk, often like the majority of the benefit, without having used any upstream data.
---
The research suggests that self-pretraining on downstream datasets can be as effective as traditional methods, challenging the notion that large, unsupervised datasets are necessary for effective model training 2.
Challenging Assumptions
The assumptions surrounding pretraining processes are being critically examined. argues that many probing studies make loose causal claims about why pretraining works, without providing a clear mechanism 3. He questions the focus on mutual information and the ease of prediction, suggesting that the real value of pretraining might lie elsewhere.
The danger of a lot of the probing work is that it leads people to think that, oh, this is why it works, as opposed to like, oh, here's a curious property of these representations.
---
Lipton emphasizes the need for a deeper understanding of what makes an initialization effective, rather than relying on superficial explanations 4.
Related Episodes


Kyunghyun Cho: Neural Machine Translation, Language, and Doing Good Science
Answers 383 questions

Hugo Larochelle: Deep Learning as Science
Answers 383 questions

Ted Underwood: Machine Learning and the Literary Imagination
Answers 383 questions

Hattie Zhou: Lottery Tickets and Algorithmic Reasoning in LLMs
Answers 383 questions

Been Kim: Interpretable Machine Learning
Answers 383 questions

Thomas Dietterich: From the Foundations
Answers 383 questions

Yann LeCun on his Start in Research and Self-Supervised Learning
Answers 383 questions

Eric Jang on Robots Learning at Google and Generalization via Language
Answers 383 questions

Melanie Mitchell: Abstraction and Analogy in AI
Answers 383 questions

Linus Lee: At the Boundary of Machine and Mind
Answers 383 questions

Scott Aaronson: Against AI Doomerism
Answers 383 questions

Nicholas Thompson: AI and Journalism
Answers 383 questions

Some Changes at The Gradient
Answers 383 questions

Luis Voloch: AI and Biology
Answers 383 questions

Joel Lehman: Open-Endedness and Evolution through Large Models
Answers 383 questions
