Learn more
Join Dexa

Model Scaling Insights

Dwarkesh and John discuss the mystery behind why bigger models with more parameters can be more sample efficient in training, suggesting that larger models act as ensembles of different computation circuits, increasing the chances of finding the right function.
  • In this clip

  • From this podcast

    Dwarkesh Podcast avatar

    Dwarkesh Podcast

    John Schulman (OpenAI Cofounder) - Reasoning, RLHF, & Plan for 2027 AGI

  • Related Questions

    • How does increasing model size affect performance in deep learning as discussed in the episode Ilya Sutskever: Deep Learning | Lex Fridman Podcast #94 and the clip Introduction to Deep Double Descent?

    • How does the size of a neural network affect its performance in deep learning, as discussed in the episode Ilya Sutskever: Deep Learning | Lex Fridman Podcast #94 and the clip Deep Double Descent?

    • How does the size of a neural network affect its performance in deep learning, as discussed in the episode Ilya Sutskever: Deep Learning | Lex Fridman Podcast #94 and the clip Introduction to Deep Double Descent?

Built by
Charlie AI
© 2024 DexaPressTermsPrivacySupport