Fine Tuning Transformers
Kirill and Jon discuss the intricacies of fine tuning transformer architectures, hinting at future episodes dedicated to this evolving topic. They reflect on the wealth of knowledge shared across two extensive episodes, covering both decoder-only and encoder-decoder transformers. Additionally, Kirill invites listeners to explore their exclusive course on large language models, offering deeper insights beyond the podcast format.In this clip
From this podcast

Super Data Science: ML & AI Podcast with Jon Krohn
759: Full Encoder-Decoder Transformers Fully Explained — with Kirill Eremenko
Related Questions
How are Large Language Models (LLMs) fine-tuned post-training in the episode Teaching Large Language Models to Reason with Reinforcement Learning with Alex Havrilla - 680 and the clip Exploration and Diversity?
I have a question about the episode Navigating Machine Learning Careers: Insights from Meta to Consulting // Ilya Reznik // #286 and the clip Fine Tuning Insights on how to train large language models (LLMs) accurately.
Teach me about neural networks as discussed in the episode 747: Technical Intro to Transformers and LLMs — with Kirill Eremenko and the clip Neural Network Dynamics