Adversarial Grokking Explained
Grokking, traditionally linked to delayed generalization in specific settings, reveals a new dimension when examining adversarial noise. While clean train and test accuracies plateau simultaneously, robustness against adversarial attacks emerges significantly later. This phenomenon occurs without any adversarial training, highlighting the importance of prolonged training and the emergence of sparse solutions within neural networks.In this clip
From this podcast

Machine Learning Street Talk (MLST)
Want to Understand Neural Networks? Think Elastic Origami! - Prof. Randall Balestriero
Related Questions