768: Is Claude 3 Better than GPT-4? — with Jon Krohn (@JonKrohnLearns)

Topics covered
Popular Clips
Questions from this episode
Episode Highlights
Benchmarks
Jon Krohn evaluates the performance of Claude 3, a new model family from Anthropic, which includes the Haiku, Sonnet, and Opus models. He highlights that Claude 3 Opus, the most powerful of the three, potentially outperforms GPT-4 and Gemini 1.0 Ultra across various benchmarks like MMLU and GPQA 1. However, Jon notes the limitations of these benchmarks, as they may not fully capture the models' capabilities and could be subject to overfitting 1.
Benchmarks focus on specific tasks, and so they may not represent the full range of capabilities of an LLM that you're interested in.
---
Despite these concerns, he acknowledges that Claude 3 is at least in the same tier as its competitors.
Qualitative Insights
In a qualitative analysis, Jon shares his anecdotal experiences with Claude 3, particularly its ability to provide rare facts. He found that Claude 3 Opus excelled in generating unique information compared to GPT-4 and Gemini 1.0 Ultra 2. Jon emphasizes that while Claude 3 does not perform real-time internet searches, it offers excellent recall over a 200,000 token context window, making it a valuable tool for specific use cases 2.
Claude three actually did the best... I did get back some rare facts and quotes right off the bat.
---
He plans to continue testing Claude 3 on various tasks, expressing satisfaction with its performance so far.
Related Episodes

666: GPT-4 — with Jon Krohn (@JonKrohnLearns)
Answers 383 questions
670: LLaMA: GPT-3 performance, 10x smaller — with Jon Krohn (@JonKrohnLearns)
Answers 383 questions

808: In Case You Missed It in July 2024 — with Jon Krohn (@JonKrohnLearns)
Answers 383 questions
660: Five Ways to Use ChatGPT for Data Science — with Jon Krohn (@JonKrohnLearns)
Answers 383 questions
720: OpenAI’s DALL-E 3, Image Chat and Web Search — with Jon Krohn (@JonKrohnLearns)
Answers 383 questions
684: Get More Language Context out of your LLM — with Jon Krohn (@JonKrohnLearns)
Answers 383 questions
740: Q*: OpenAI's Rumored AGI Breakthrough — with @JonKrohnLearns
Answers 383 questions
818: In Case You Missed It in August 2024 — with Jon Krohn (@JonKrohnLearns)
Answers 383 questions
638: ChatGPT Holiday Greeting — with @JonKrohnLearns
Answers 383 questions




