Shaping AI Benchmarks with Together AI Co-Founder Percy Liang

Topics covered
Popular Clips
Episode Highlights
Framework
, co-founder of Together AI and Stanford Associate Professor, developed HELM, a comprehensive framework for evaluating language models. HELM, which stands for Holistic Evaluation of Language Models, was created to address the broad and varied objectives of language models, unlike classical machine learning methods. The framework evaluates models holistically, considering both capabilities and risks, and was inspired by decentralized approaches like the Big Science project 1.
We created HELM, which stands for holistic evaluation of language models. Back in 2022, which is eons ago, the feeling was that language models were a different objective, unlike classical machine learning methods, because they were so broad.
---
HELM's development involved collaboration within the Stanford Center for Research on Foundation Models, organizing participants into areas like bias, reasoning, and robustness. This resulted in a robust infrastructure supporting exhaustive evaluations across multiple datasets 1.
Metrics
HELM's evaluation metrics extend beyond accuracy to include bias, calibration, robustness, and toxicity, among others. emphasized the importance of standardizing evaluations to ensure apples-to-apples comparisons across models and datasets. This transparency allows users to drill down into results, seeing actual predictions and prompts used, fostering reproducibility 2.
Evaluation, at least at the time, it was the most comprehensive evaluation where we took 30 different models that we had a hold of through either API or open weights, and we looked at seven different metrics from accuracy beyond accuracy, including bias, calibration, robustness, toxicity, and so on.
---
Liang also discussed the importance of considering factors like reliability, cost, and fine-tuning capabilities when selecting models, rather than solely relying on leaderboard performance 3.
Challenges
Benchmarking AI models comes with challenges, such as the risk of overfitting and the lack of transparency around datasets. noted that while some model providers filter benchmarks to prevent overt copying, the real issue lies in interpreting what benchmarks truly measure. He highlighted the slippery slope of overfitting to benchmarks, even unintentionally 2.
I feel like it's a pretty slippery slope before you're really overfitting to the benchmark without maybe even realizing it. Because, yeah, who doesn't want better benchmark numbers?
---
Liang also pointed out the difficulty in interpreting zero-shot performance due to the lack of transparency in datasets, which can lead to misleading results 4.
Related Episodes


Transforming Search with Perplexity AI’s CTO Denis Yarats
Answers 383 questions

Jerome Pesenti — Large Language Models, PyTorch, and Meta
Answers 383 questions

The Power of AI in Search with You.com's Richard Socher
Answers 383 questions

Transforming Data into Business Solutions with Salesforce AI CEO, Clara Shih
Answers 383 questions

Reinventing AI Agents with Imbue CEO Kanjun Qiu
Answers 383 questions

Revolutionizing AI Data Management with Jerry Liu, CEO of LlamaIndex
Answers 383 questions

Richard Socher — The Challenges of Making ML Work in the Real World
Answers 383 questions

Operationalizing Machine Learning: Interview with Shreya Shankar
Answers 383 questions

Emily M. Bender — Language Models and Linguistics
Answers 383 questions

How EleutherAI Trains and Releases LLMs: Interview with Stella Biderman
Answers 383 questions

Dave Rogenmoser & Saad Ansari on Growing & Maintaining Jasper AI
Answers 383 questions

Chip Huyen of Claypot AI— ML Research and Production Pipelines
Answers 383 questions

Peter Norvig – Singularity Is in the Eye of the Beholder
Answers 383 questions

Harnessing AI for legal practice with CoCounsel’s Jake Heller
Answers 383 questions













