Evaluating Language Models

Nathan delves into the complexities of evaluating large language models, highlighting the challenges of comparing models across companies and the potential for evaluation data contamination. He emphasizes the importance of looking beyond headline evaluation scores and suggests building better validation tools for open models.