Training Phases

Jonathan discusses the concept of training happening in distinct phases, revealing that networks find the same local optimum regardless of data order. This insight challenges traditional views on stochastic gradient descent and network optimization.