Heaps' Law
Updated 2026-08-05
INTRODUCTION
English translation pending.
CORE DEFINITION
Formulated by Harold Stanley Heaps in his 1978 book on information retrieval, the law states that vocabulary size grows with text length according to a power function whose exponent typically lies between 0.4 and 0.6 for natural language. It is closely related to Herdan's law and follows from the Zipfian distribution of word frequencies. The key constraint is that the relation is statistical and corpus-dependent, since the exponent varies with language, genre and text type, so it predicts averages rather than the count in any single document.
SCAFFOLDING EFFECT
Reduce cognitive load
- Plan vocabulary: Estimate how large a lexicon a corpus of a given size will require. - Budget learning: Expect new concepts per page to fall as you read on, and plan effort accordingly. - Stop sensibly: Use the flattening curve to decide when more reading yields too few new terms.
Anchor fast decisions
Because word frequencies follow a Zipfian distribution, a few words recur constantly while a long tail of words appears rarely. As a text lengthens, the chance that the next word has never been seen before declines in proportion to a power of the length and never reaches zero. Vocabulary therefore grows sublinearly, producing a curve that keeps rising but with diminishing returns.
MINIMUM ACTION
In progress 0/1Practice this model in one real situation:
account_treeGenealogyexpand_more
menu_bookReferencesexpand_more
Source support: Explicit
- en.wikipedia.orghttps://en.wikipedia.org/wiki/Heaps's_lawverified
PRIVATE NOTES · Only visible to you
SAVED Q&A
ENTRY Q&A · Private saving available
Ask with a clear boundary
thinkingmodels answers from published entry context only.
Your question is sent to thinkingmodels. The answer uses public entry context only.
RELATED MODELS