Scrubbing Sensitive Information

The discussion delves into the challenges of ensuring AI models do not retain sensitive information, even after fine-tuning. Peter explains how traditional methods may leave residual knowledge within the model, while he proposes a more thorough approach to erase this information from both the final classification layer and intermediate representations. This robust method aims to enhance model safety against potential probing attacks.