Scaling Innovations
Joe discusses the complexities of developing the Llama 3 model, emphasizing the extensive effort required to write the accompanying paper. He highlights the innovative scaling techniques employed, including training on over 15 trillion tokens and utilizing synthetic data to enhance model performance. The conversation also touches on the significant infrastructure and teamwork needed to manage the challenges of training on thousands of GPUs.In this clip
From this podcast

Training Data
Meta’s Joe Spisak on Llama 3.1 405B and the Democratization of Frontier Models | Training Data
Related Questions
Which model of large language model (LLM) is discussed in the episode 806: Llama 3.1 405B: The First Open-Source Frontier LLM — with Jon Krohn (@JonKrohnLearns) and the clip Llama 3.1 Insights?
How are large language models (LLMs) trained, as discussed in the episode 670: LLaMA: GPT-3 performance, 10x smaller — with Jon Krohn (@JonKrohnLearns) and the clip Llama Model Insights?
How will large language models (LLMs) and AI change software engineering and the software development lifecycle (SDLC) as discussed in the episode Meta’s Joe Spisak on Llama 3.1 405B and the Democratization of Frontier Models | Training Data, and the clip Scaling Innovations?