Been Kim: Interpretable Machine Learning

Topics covered
Popular Clips
Episode Highlights
Conceptual Shift
emphasizes the shift from example-based to concept-based interpretability in AI, highlighting its efficiency in communication with experts. She explains that while example-based methods can miss nuances, concept-based approaches allow machines to speak the language of domain experts, such as medical professionals, enhancing efficiency and adaptation 1. This shift is exemplified by her work on TCAV, which uses concepts to explain AI decisions, proving effective even in early tests 2. notes the importance of selecting the right medium—whether examples, pixels, or concepts—based on the task at hand 1.
The decision boundaries tend to be complex. Distributions over these complex data points are also complex. So you need a nuance, which we call criticisms there, things that doesn't quite fit into that prototype, but human needs to know.
---
She also discusses dimensions of interpretability, such as cognitive chunks and compositionality, which are crucial for understanding AI models 3.
Interpretability Challenges
addresses the challenges in achieving interpretable machine learning models, particularly the issues of biases and testing rigor. She highlights the problem of confirmation bias, where humans tend to see what they expect in AI outputs, making it crucial to develop robust evaluation methods 4. Kim's work demonstrates that explanations from trained and random networks can appear indistinguishable, underscoring the need for careful evaluation 5.
We have to be careful with these explanations, how we evaluate it, how we set the goal, and how we claim the success.
---
She warns against the misuse of interpretability methods to falsely build trust in AI systems, stressing the importance of understanding the unknowns that machines might reveal 6.
Practical Applications
In real-world applications, underscores the necessity of human feedback in developing interpretable AI models. She recounts her experiences with domain experts during high-stakes situations, emphasizing that AI tools must be practical and user-friendly for experts who lack time for complex analyses 7. Kim's work on TCAV and moodboard search illustrates the potential of concept-based methods in creative fields, allowing users to define personal concepts like memories or emotions 8.
If my work didn't influence the way humans outside of a machine learning field, the whole world didn't change the way that the world works. That's not a success in my opinion.
---
She also explores extensions of TCAV, such as incorporating the magnitude of conceptual sensitivity, to enhance the method's applicability to non-visual concepts 9.
Related Episodes


Kyunghyun Cho: Neural Machine Translation, Language, and Doing Good Science
Answers 383 questions

Vera Liao: AI Explainability and Transparency
Answers 383 questions

Catherine Olsson and Nelson Elhage: Anthropic, Understanding Transformers
Answers 383 questions

Terry Winograd: AI, HCI, Language, and Cognition
Answers 383 questions

Ted Underwood: Machine Learning and the Literary Imagination
Answers 383 questions

Linus Lee: At the Boundary of Machine and Mind
Answers 383 questions

Melanie Mitchell: Abstraction and Analogy in AI
Answers 383 questions

Joon Park: Generative Agents and Human-Computer Interaction
Answers 383 questions

Hugo Larochelle: Deep Learning as Science
Answers 383 questions

Subbarao Kambhampati: Planning, Reasoning, and Interpretability in the Age of LLMs
Answers 383 questions

Zachary Lipton: Where Machine Learning Falls Short
Answers 383 questions

Helena Sarin on being an AI Artist
Answers 383 questions

Kate Park: Data Engines for Vision and Language
Answers 383 questions

François Chollet: Keras and Measures of Intelligence
Answers 383 questions

Alex Tamkin on Self-Supervised Learning and Large Language Models
Answers 383 questions
