Benchmarking AI Models

The latest Humanities Last examination revealed a significant performance gap among AI models, with one achieving an impressive 26.6%, far surpassing competitors like OpenAI's models. The discussion highlights the importance of agentic tools that combine web browsing with reasoning capabilities, suggesting that access to such tools could be a game-changer for AI performance. Additionally, the hiring of PhD students to train models may also contribute to their success.