Adversarial Grokking Explained

Grokking, traditionally linked to delayed generalization in specific settings, reveals a new dimension when examining adversarial noise. While clean train and test accuracies plateau simultaneously, robustness against adversarial attacks emerges significantly later. This phenomenon occurs without any adversarial training, highlighting the importance of prolonged training and the emergence of sparse solutions within neural networks.