Evaluating Language Models

Walid and Gábor discuss the challenges of evaluating language models, emphasizing the limitations of using perplexity as a sole metric. Gábor highlights the importance of optimizing for human perception rather than relying solely on proxies like perplexity or machine translation metrics.