Numerical Precision in Neural Networks
Daniel and Greg discuss the impact of numerical precision on theoretical properties in neural networks, highlighting the potential for divergences and subtle distortions in training trajectories. They emphasize the importance of observing the scale of activations and preactivations to avoid losing precision. While it is a solvable problem from an engineering perspective, ongoing innovations in floating point format are necessary as models continue to grow larger.In this clip
From this podcast

The Gradient
Greg Yang on Communicating Research, Tensor Programs, and µTransfer
Related Questions