Grokking
Version 1.0.0 · Updated 2026-07-28
CORE DEFINITION
In some experiments on small algorithmic datasets, neural networks first fit the training set while generalization remains near chance. Generalization improves substantially only after much more optimization. This delayed improvement is called grokking; it does not mean training accuracy stays low or that every learning process must pass through this stage.
SCAFFOLDING EFFECT
Reduce cognitive load
Distinguish memorization of training examples from generalization to unseen examples. Track their different time scales instead of treating a training fit as mastery or every plateau as an imminent breakthrough.
Anchor fast decisions
Fitting training data and generalizing to unseen data can occur on different time scales. Power and colleagues observed generalization long after overfitting in some networks. This observation alone establishes neither a single causal mechanism for all models nor a guarantee that plateaus end.
MINIMUM ACTION
In progress 0/2Practice this model in one real situation:
account_treeGenealogyexpand_more
menu_bookReferencesexpand_more
Source support: Explicit
- en.wikipedia.orghttps://en.wikipedia.org/wiki/Grokverified
- arxiv.orghttps://arxiv.org/abs/2201.02177verified
PRIVATE NOTES · Only visible to you
SAVED Q&A
ENTRY Q&A · Private saving available
Ask with a clear boundary
thinkingmodels answers from published entry context only.
Your question is sent to thinkingmodels. The answer uses public entry context only.
RELATED MODELS