Efficient Language Model Approximations

Sasha discusses the hypothesis behind approximating large model distributions over sequences rather than word-level distributions. Daniel explores the potential of combining pruning and knowledge distillation for more efficient models. Sasha suggests the potential benefits of distillation in conjunction with quantization for modern language models.