Published Oct 13, 2022

Zachary Lipton: Where Machine Learning Falls Short

Zachary Lipton discusses his journey from music to academia while critiquing current approaches in machine learning, urging a rethink of pretraining models and emphasizing the limitations of algorithms in addressing social issues.
Episode Highlights
The Gradient logo

Popular Clips

Episode Highlights

  • Pretraining Insights

    Pretraining models on seemingly nonsensical data can yield surprising benefits. explains that even when models are pretrained on random phonemes and special tokens, they can achieve a significant portion of the performance benefits typically associated with large upstream corpora 1. This raises questions about the actual mechanisms behind pretraining benefits, as notes, "We get out a huge chunk, often like the majority of the benefit, without having used any upstream data."

    We get out a huge chunk, often like the majority of the benefit, without having used any upstream data.

    ---

    The research suggests that self-pretraining on downstream datasets can be as effective as traditional methods, challenging the notion that large, unsupervised datasets are necessary for effective model training 2.

       

    Challenging Assumptions

    The assumptions surrounding pretraining processes are being critically examined. argues that many probing studies make loose causal claims about why pretraining works, without providing a clear mechanism 3. He questions the focus on mutual information and the ease of prediction, suggesting that the real value of pretraining might lie elsewhere.

    The danger of a lot of the probing work is that it leads people to think that, oh, this is why it works, as opposed to like, oh, here's a curious property of these representations.

    ---

    Lipton emphasizes the need for a deeper understanding of what makes an initialization effective, rather than relying on superficial explanations 4.

Related Episodes