Neel Nanda - Mechanistic Interpretability (Sparse Autoencoders)

Topics covered
Popular Clips
Questions from this episode
- Asked by 129 people
- Asked by 83 people
- Asked by 81 people
- Asked by 79 people
- Asked by 54 people
- Asked by 45 people
- Asked by 29 people
- Asked by 28 people
- Asked by 24 people
- Asked by 22 people
- Asked by 21 people
- Asked by 20 people
- Asked by 17 people
Episode Highlights
Feature Learning
Sparse Autoencoders (SAEs) play a crucial role in feature learning, with their performance improving as they are trained on more data. highlights that the frequency of a feature significantly impacts its learning, with more common features being more likely to be captured by the model 1. However, the process can be probabilistic, especially for niche features, leading to variability in feature learning across different model widths 1.
SAEs get better when trained on more data. As a general rule, sometimes they'll kind of saturate.
---
This variability suggests that practitioners should experiment with different SAE widths to capture the desired features effectively 2.
Interpretability
Understanding the interpretability of features in neural networks is a complex task. notes that larger Sparse Autoencoders (SAEs) often face training stability issues, with many features remaining inactive or "dead" 3. This suggests that while larger models may hold more concepts, they also require careful management to ensure meaningful feature activation.
I want to know what's going on in the model. What is the computation happening, what are the variables?
---
The goal is to understand the computation within the model, which involves identifying the variables and their interactions, rather than merely increasing the number of features 4.
Related Episodes


Neel Nanda - Mechanistic Interpretability
Answers 383 questions

#041 - Biologically Plausible Neural Networks - Dr. Simon Stringer
Answers 383 questions

The Elegant Math Behind Machine Learning - Anil Ananthaswamy
Answers 383 questions

#54 Gary Marcus and Luis Lamb - Neurosymbolic models
Answers 383 questions

Dr. Paul Lessard - Categorical/Structured Deep Learning
Answers 383 questions

#97 SREEJAN KUMAR - Human Inductive Biases in Machines from Language
Answers 383 questions

Explainability, Reasoning, Priors and GPT-3
Answers 383 questions

047 Interpretable Machine Learning - Christoph Molnar
Answers 383 questions

Dr. Sanjeev Namjoshi - Active Inference
Answers 383 questions

#55 Self-Supervised Vision Models (Dr. Ishan Misra - FAIR).
Answers 383 questions

Robert Lange on NN Pruning and Collective Intelligence
Answers 383 questions

Nora Belrose - AI Development, Safety, and Meaning
Answers 383 questions

#59 - Jeff Hawkins (Thousand Brains Theory)
Answers 383 questions
