Training Phases
Jonathan discusses the concept of training happening in distinct phases, revealing that networks find the same local optimum regardless of data order. This insight challenges traditional views on stochastic gradient descent and network optimization.In this clip
From this podcast

Generally Intelligent
Episode 13: Jonathan Frankle, MIT, on the lottery ticket hypothesis and the science of deep learning
Related Questions