The Lottery Ticket Hypothesis
Updated 2026-08-08
INTRODUCTION
English translation pending.
CORE DEFINITION
Proposed by Jonathan Frankle and Michael Carbin. They showed that a standard over-parameterised network contains subnetworks, called winning tickets, that can be trained in isolation to accuracy matching the full network, provided the surviving weights are reset to the values they held at initialisation. The implication is that dense training works largely as a search over a pool of candidate subnetworks. It also explains why pruning and then rewinding, rather than reinitialising, is what preserves performance. The claim is empirical and architecture-dependent, not a guarantee that every task hides a tiny sufficient network.
SCAFFOLDING EFFECT
Reduce cognitive load
- Redundancy reading: Treat excess capacity as a search pool rather than as waste to cut on sight. - Pruning timing: Wait until the winning subnetwork is identified before removing the rest. - Portfolio triage: Keep the few connections that carry the result and cut the dormant ones.
Anchor fast decisions
Over-parameterisation supplies a large space of possible subnetworks. Gradient descent explores that space and, through the initialisation it happens to draw, some sparse subnetwork becomes the one carrying the learned function. Because the crucial ingredient is the initial weights rather than the final ones, rewinding to the start recovers the ticket while reinitialising destroys it. Redundancy therefore pays for exploration, and efficiency is what remains after the search.
MINIMUM ACTION
In progress 0/1Practice this model in one real situation:
account_treeGenealogyexpand_more
menu_bookReferencesexpand_more
Source support: Explicit
- en.wikipedia.orghttps://en.wikipedia.org/wiki/Lottery_ticket_hypothesisverified
PRIVATE NOTES · Only visible to you
SAVED Q&A
ENTRY Q&A · Private saving available
Ask with a clear boundary
thinkingmodels answers from published entry context only.
Your question is sent to thinkingmodels. The answer uses public entry context only.
RELATED MODELS