Benchmark Evolution

Recent tests reveal that some models, like Mistral, are overfitting on benchmarks, while others, such as Claude and GPT, perform well on novel questions. The Arc dataset, created years ago, has its flaws, including redundancy and similarity among tasks. Plans are underway to release an updated version, allowing users to query an API to avoid accidental training on the dataset.