Learning Where to Look

Kristen discusses the evolution of visual intelligence in agents, emphasizing the need for them to determine where to focus their attention in video rather than relying solely on static images. She highlights the limitations of traditional datasets, which are curated human photographs, and contrasts them with the dynamic, first-person perspective of continuous video. This shift necessitates a new approach to training systems that can intelligently identify and prioritize relevant content in real-time.