Composing Multimodal Systems
The discussion highlights a novel approach to building multimodal systems by composing separate models that communicate through language. While traditional methods often rely on large neural networks, this method leverages the strengths of various models, enabling them to tackle complex applications like contextual image captioning and video understanding. The potential of using language as an intermediate representation opens up exciting possibilities for collaboration among diverse models.In this clip
From this podcast

The Gradient
Pete Florence: Dense Visual Representations, NeRFs, and LLMs for Robotics
Related Questions