Measuring LLM Hallucinations
Erwan discusses a novel approach to evaluating large language models by utilizing graph prompts instead of traditional binary questions. This method allows for a richer analysis of hallucinations, yielding insights that correlate well with existing metrics from the Hallucination Leaderboard. By focusing on just five small graphs, Erwan demonstrates how this technique can effectively measure and compare the performance of various LLMs.In this clip
From this podcast

Data Skeptic
Auditing LLMs and Twitter
Related Questions