Exploding Gradient
Updated 2026-08-09
INTRODUCTION
English translation pending.
CORE DEFINITION
In training deep neural networks, the gradient of the loss with respect to early layers is computed by multiplying many terms along the chain of layers. When those terms are consistently greater than one, the product grows exponentially with depth, producing enormous updates and numeric overflow. The model then fails to converge, often showing NaN losses. Gradient clipping, careful initialization and normalization are the standard countermeasures.
SCAFFOLDING EFFECT
Reduce cognitive load
- Check the chain: ask whether small per-step exaggerations compound along a long chain of handoffs. - Clip at each node: normalize the signal before passing it on so it cannot grow without bound. - Distinguish the failure: separate amplification from attenuation, since the two require different remedies.
Anchor fast decisions
Backpropagation multiplies one derivative per layer, so the total gradient scales with the product of all of them. A product of many factors slightly above one grows exponentially with depth, and the parameter update becomes so large that it overshoots the minimum and destroys the learned state. Clipping bounds the norm of the update, which stops the amplification from compounding.
MINIMUM ACTION
In progress 0/1Practice this model in one real situation:
account_treeGenealogyexpand_more
menu_bookReferencesexpand_more
Source support: Explicit
- en.wikipedia.orghttps://en.wikipedia.org/wiki/Vanishing_gradient_problemverified
PRIVATE NOTES · Only visible to you
SAVED Q&A
ENTRY Q&A · Private saving available
Ask with a clear boundary
thinkingmodels answers from published entry context only.
Your question is sent to thinkingmodels. The answer uses public entry context only.
RELATED MODELS