Vision and Language Embedding
Hao explains the process of embedding in both language and vision modalities, highlighting the use of object detection to convert images into meaningful sequences of features. He contrasts this with traditional grid embedding methods, emphasizing the growing dominance of object detection in vision-language tasks. The discussion sheds light on how positional embeddings are integrated into visual data, paralleling techniques used in language processing.In this clip
From this podcast

NLP Highlights
107 - Multi-Modal Transformers, with Hao Tan and Mohit Bansal
Related Questions