Evaluating Language Models
Nathan delves into the complexities of evaluating large language models, highlighting the challenges of comparing models across companies and the potential for evaluation data contamination. He emphasizes the importance of looking beyond headline evaluation scores and suggests building better validation tools for open models.In this clip
From this podcast

Interconnects Audio
Big Tech's LLM evals are just marketing
Related Questions