Latent Action Models

The discussion delves into the intriguing concept of latent action models, focusing on how representations are learned from playthrough videos to predict subsequent frames. Insights reveal a structured approach to defining action spaces through discrete codebooks, facilitating interaction with models. Additionally, the training process of the tokenizer, latent action model, and dynamics model is explored, highlighting their interdependencies and training methodologies.