Model Evaluation Insights

Two models, one from Berkeley and the other from Stanford, initially scored the same on benchmarks, but a deeper evaluation revealed intriguing biases in their assessments. By leveraging GPT4 to rank the models, a unique methodology emerged, exposing the preference biases inherent in both human and AI evaluations. The order of presentation proved crucial, leading to unexpected insights about model performance.