Evaluating Language Models

Mor Geva discusses the challenges of evaluating language models, highlighting the limitations of current benchmark tasks and the need for continuous updating and retraining processes.