Published Jun 24, 2022

Preetum Nakkiran: An Empirical Theory of Deep Learning

Preetum Nakkiran delves into the empirical theory of deep learning, offering a unique perspective on model behavior through concepts like distributional generalization and the double descent phenomenon, while sharing his personal research journey and emphasizing the importance of blending theoretical and empirical approaches to unravel AI's complexities.
Episode Highlights
The Gradient logo

Popular Clips

Episode Highlights

  • Empirical Theory

    , a research scientist at Apple, explores the concept of an empirical theory of deep learning, aiming to understand it as a natural system. He describes deep learning as a map from inputs, such as model architecture and training procedures, to outputs, which are the trained models 1. This approach seeks to identify the structure within this map, akin to studying any natural object 2. highlights Preetum's contributions to deep learning theory, including his work on the deep double descent paper 3.

       

    Generalization

    Generalization in deep learning traditionally focuses on scalar metrics like test loss, but argues for a broader understanding. He emphasizes that models are more than just their scalar performance metrics, and understanding them requires examining the entire function they represent 4. This involves considering the distribution of errors and the types of errors made, not just the overall accuracy. notes the importance of formalizing claims about generalization within a causal framework, highlighting the complexity of these concepts 5.

       

    Distributional Gen.

    Distributional generalization, a concept proposed by , examines the joint distribution of inputs and outputs in deep learning models. This approach seeks to understand model behavior beyond traditional metrics by focusing on the joint distribution of inputs and outputs on both train and test sets 6. highlights the challenges of introducing novel ideas in academia, sharing his experiences with paper rejections due to the unconventional nature of his work 7. He explains that understanding this joint distribution can reveal insights into model behavior, such as how noise is reproduced in overparameterized models 8.

Related Episodes