Learning Rate Decay
Updated 2026-08-09
INTRODUCTION
English translation pending.
CORE DEFINITION
Learning rate decay reduces the step size used in training over time. The core claim is that a large rate early allows rapid progress toward the region of the optimum, while a small rate later permits fine adjustment without overshooting, which prevents oscillation around the minimum. The qualification is that the schedule matters: starting too small wastes time on detail, and staying too large prevents convergence, so the transition point must be chosen deliberately.
SCAFFOLDING EFFECT
Reduce cognitive load
- Rhythm Control: In a new field, move broadly and tolerate roughness at first, then refine once you are an expert. - Phase Awareness: Early on, chase speed and coverage; later, chase precision and accept slower progress. - Switch Signal: Knowing when error rates plateau tells you when to reduce the rate rather than push harder.
Anchor fast decisions
A large learning rate in early training moves the parameters quickly toward the region of the optimum. Reducing it later allows fine convergence and prevents oscillation around the minimum. The analogy is direct: broad strokes for speed early, careful refinement for precision later.
MINIMUM ACTION
In progress 0/1Practice this model in one real situation:
account_treeGenealogyexpand_more
menu_bookReferencesexpand_more
Source support: Explicit
- en.wikipedia.orghttps://en.wikipedia.org/wiki/Learning_rateverified
PRIVATE NOTES · Only visible to you
SAVED Q&A
ENTRY Q&A · Private saving available
Ask with a clear boundary
thinkingmodels answers from published entry context only.
Your question is sent to thinkingmodels. The answer uses public entry context only.
RELATED MODELS