Masking as Training
Hattie Zhou discusses the surprising findings of her study on masking in neural networks, revealing that masked subnetworks already contain valuable information about the task at initialization. This challenges the notion of pruning as removing useless weights and suggests that the remaining subnetworks work in tandem with the masked weights. The implications of this perspective on identifying lottery tickets at initialization and the training process are explored.In this clip
From this podcast

The Gradient
Hattie Zhou: Lottery Tickets and Algorithmic Reasoning in LLMs
Related Questions