Benchmark Evolution
Recent tests reveal that some models, like Mistral, are overfitting on benchmarks, while others, such as Claude and GPT, perform well on novel questions. The Arc dataset, created years ago, has its flaws, including redundancy and similarity among tasks. Plans are underway to release an updated version, allowing users to query an API to avoid accidental training on the dataset.In this clip
From this podcast

Dwarkesh Podcast
Francois Chollet - LLMs won’t lead to AGI - $1,000,000 Prize to find true solution
Related Questions