Evaluating Common Sense Intelligence

The discussion explores the effectiveness of different evaluation methods for common sense intelligence systems, highlighting the limitations of automated evaluation and the importance of generative evaluation. The potential scalability of human-based evaluation methods is also considered.