Cross Attention Insights
Hao explains the architecture of encoders for both language and vision inputs, highlighting the use of transformer blocks. He clarifies that while cross attention is not a new concept, their approach uniquely stacks these layers to create high-level representations, enhancing the integration of vision and language data.In this clip
From this podcast

NLP Highlights
107 - Multi-Modal Transformers, with Hao Tan and Mohit Bansal
Related Questions