Cognitive Scaffold

Preparing your thinking workspace

arrow_back_ios_new
MENTAL MODEL · M0025

Grokking

Grokking
SystemsHigh supportComplexity Science
Included
account_tree

Version 1.0.0 · Updated 2026-07-28

CORE DEFINITION

In some experiments on small algorithmic datasets, neural networks first fit the training set while generalization remains near chance. Generalization improves substantially only after much more optimization. This delayed improvement is called grokking; it does not mean training accuracy stays low or that every learning process must pass through this stage.

SCAFFOLDING EFFECT

psychology

Reduce cognitive load

Distinguish memorization of training examples from generalization to unseen examples. Track their different time scales instead of treating a training fit as mastery or every plateau as an imminent breakthrough.

anchor

Anchor fast decisions

Fitting training data and generalizing to unseen data can occur on different time scales. Power and colleagues observed generalization long after overfitting in some networks. This observation alone establishes neither a single causal mechanism for all models nor a guarantee that plateaus end.

MINIMUM ACTION

In progress 0/2

Practice this model in one real situation:

Check to track your progress (stored locally)
Learning progress0%
account_treeGenealogyexpand_more
menu_bookReferencesexpand_more

Source support: Explicit

  • link
    en.wikipedia.orghttps://en.wikipedia.org/wiki/GrokEN · Explicit
    verified
  • link
    arxiv.orghttps://arxiv.org/abs/2201.02177EN
    verified

RELATED MODELS