Autoencoder Insights
SAEs improve with more data, but there's a question of whether they reach true saturation or if training methods limit their potential. Exploring the emergence of new features versus refinement of existing ones could yield fascinating insights. A recent evaluation compared specialized SAEs trained on specific tasks to broader models, revealing intriguing phenomena like feature oversplitting, which complicates analysis but highlights the nuances of training methods.In this clip
From this podcast

Machine Learning Street Talk (MLST)
Neel Nanda - Mechanistic Interpretability (Sparse Autoencoders)
Related Questions