Evaluation Challenges

The discussion revolves around the relevance of existing benchmarks to newly defined problems in AI. It raises the question of whether evaluation methods are primarily self-referential or if they can draw from established standards. Insights into the complexities of problem definition and its implications for effective evaluation are explored.