Arvind discusses the impressive yet misleading nature of large data volumes, such as 150 billion hours of video, and how extracting text from this data reveals a smaller scale than expected. He reflects on past surprises in AI capabilities, particularly with models like GPT-2, and expresses skepticism about the potential for new emergent text capabilities from diverse datasets like YouTube videos, despite acknowledging the promise of multimodal advancements.