Amir discusses the various benchmark datasets used for evaluating video understanding models, including ActivityNet and mini kinetics. He highlights the importance of comparing different frame selection methods, including reinforcement learning and audio-based approaches. The conversation also touches on the potential for combining skip convolutions with frame selection techniques, emphasizing their complementary nature in addressing spatial redundancy in video processing.