Benchmarking AI Models
Junyang discusses the challenges of evaluating AI models through various benchmarks, highlighting the phenomenon of "big model smell." He emphasizes the importance of new, less exposed evaluations that can better capture a model's true performance. Alex adds that their goal is to gather insights from these evaluations to create a comprehensive understanding of model capabilities, particularly in areas like code editing.In this clip
From this podcast

ThursdAI
📅 ThursdAI - Aug8 - Qwen2-MATH King, tiny OSS VLM beats GPT-4V, everyone slashes prices + 🍓 flavored OAI conspiracy
Related Questions