Evaluating AI Benchmarks

The discussion highlights the growing challenge of evaluating AI models using fewer seeds as benchmarks become increasingly complex. Rishabh points out that while this trend is driven by necessity rather than negligence, it raises questions about the robustness of comparisons across algorithms. The conversation also introduces a novel approach borrowed from computer science, suggesting that plotting score distributions could enhance evaluation methods in reinforcement learning.