Feature Learning and Primerizations

Greg Yang explains the concept of feature learning and primerizations in neural networks. He introduces the idea of maximum update primerization (MUP) as the optimal way to scale neural networks and discusses its theoretical properties. This episode provides insights into the best parameterization to use when training large models.