Benchmarking AI Models

Junyang discusses the challenges of evaluating AI models through various benchmarks, highlighting the phenomenon of "big model smell." He emphasizes the importance of new, less exposed evaluations that can better capture a model's true performance. Alex adds that their goal is to gather insights from these evaluations to create a comprehensive understanding of model capabilities, particularly in areas like code editing.