Greedy Hierarchical Variational Auto Encoders

Chelsea Finn discusses the development of a new model, the greedy hierarchical variational auto encoder, which addresses the underfitting problem in video prediction models. By training the model in a layer-wise fashion and freezing previous models, memory improvements and better video predictions were achieved. The discussion also touches on end-to-end fine tuning and the recent paper FitVid, which focuses on efficiently leveraging model capacity.