Cross Attention Insights

Hao explains the architecture of encoders for both language and vision inputs, highlighting the use of transformer blocks. He clarifies that while cross attention is not a new concept, their approach uniquely stacks these layers to create high-level representations, enhancing the integration of vision and language data.