712: Code Llama — with Jon Krohn (@JonKrohnLearns)

Topics covered
Popular Clips
Episode Highlights
Model Performance
Code Llama's suite of models, including the general, Python, and instruct families, offers a range of options for data scientists. highlights that these models have been benchmarked against major proprietary models like OpenAI's Codex and GPT-3.5, as well as open-source models such as Starcoder. The 34 billion parameter Code Llama model notably outperforms all open-source alternatives and competes closely with GPT-3.5.
Even the seven b code llama outperforms the seven db lama two on the majority of the coding benchmarks.
---
However, GPT-4 remains superior, outperforming even the largest Code Llama model by a significant margin 1.
Cautionary Notes
While the performance of Code Llama models is impressive, Jon advises caution when interpreting these results. Since Meta published the benchmarks, there might be inherent biases in the reported performance metrics. He suggests that the best way to evaluate Code Llama's effectiveness is through personal experimentation.
But of course all of these results should be taken with a grain of salt, since meta publish them themselves.
---
This hands-on approach allows users to tailor the models to their specific needs, potentially creating powerful tools without sharing proprietary data with third parties 1.
Related Episodes

670: LLaMA: GPT-3 performance, 10x smaller — with Jon Krohn (@JonKrohnLearns)
Answers 383 questions
772: In Case You Missed It in March 2024 — with Jon Krohn (@JonKrohnLearns)
Answers 383 questions

787: MLOps: The Job and The Key Tools — with Demetrios Brinkmann
Answers 383 questions
640: What I Learned in 2022 — with Jon Krohn (@JonKrohnLearns)
Answers 383 questions

754: A Code-Specialized LLM Will Realize AGI — with Jason Warner
Answers 383 questions
676: The Chinchilla Scaling Laws — with Jon Krohn (@JonKrohnLearns)
Answers 383 questions
666: GPT-4 — with Jon Krohn (@JonKrohnLearns)
Answers 383 questions
