Data Set Dynamics
Waleed and Hao delve into the complexities of using both image captioning and visual question answering datasets for training cross-model transformers. They emphasize the importance of grounding, where the relationship between words and image components is key, rather than requiring complete descriptions. Mohit adds that the relationship between captions and questions exists on a spectrum, suggesting potential benefits in reformatting VQA questions into statement-like forms for enhanced model training.In this clip
From this podcast

NLP Highlights
107 - Multi-Modal Transformers, with Hao Tan and Mohit Bansal
Related Questions