Published Jul 3, 2023
Unifying Vision and Language Models with Mohit Bansal - 636
Mohit Bansal discusses the innovative unification of vision and language models to create more efficient, multimodal AI systems like VLT5 and UDOP, while addressing generative AI evaluation challenges like bias and factuality. He highlights significant advancements in processing diverse data types and emphasizes programmatic explainability to boost AI transparency and reliability.

Topics covered
Popular Clips
Episode Highlights
Related Episodes


ML Use Cases at Think Big Analytics with Mo Patel & Laura Frølich - #54
Answers 383 questions

Symbolic and Subsymbolic Natural Language Processing with Jonathan Mugan - #49
Answers 383 questions

Scaling Multi-Modal Generative AI with Luke Zettlemoyer - 650
Answers 383 questions

Trends in Computer Vision with Siddha Ganju - TWiML Talk #218
Answers 383 questions

Understanding AI’s Impact on Social Disparities with Vinodkumar Prabhakaran - 617
Answers 383 questions

More Language, Less Labeling with Kate Saenko - #580
Answers 383 questions

An Agentic Mixture of Experts for DevOps with Sunil Mallya - 708
Answers 383 questions

Learning Visiolinguistic Representations with ViLBERT w/ Stefan Lee - #358
Answers 383 questions

Understanding Cultural Trends with Computer Vision w/ Kavita Bala - #410
Answers 383 questions

The New DBfication of ML/AI with Arun Kumar - #553
Answers 383 questions

Building Maps and Spatial Awareness in Blind AI Agents with Dhruv Batra - 629
Answers 383 questions

AI Trends 2024: Computer Vision with Naila Murray - 665
Answers 383 questions

Building a Unified NLP Framework at LinkedIn with Huiji Gao - #481
Answers 383 questions













