Multimodal Models: Dolly Two vs. Imagen

Daniel and Christopher discuss the differences between Dolly Two and Imagen, two recent multimodal models. They explore how Imagen, with its pure text encoding, outperforms Dolly Two, which incorporates both text and image embeddings. The conversation highlights the fascinating results and the implications for language understanding and image generation.