The Bitter Lesson
Version 1.0.0 · Updated 2026-07-30
CORE DEFINITION
A viewpoint proposed by Richard Sutton, arguing that the most successful methods in AI research are often not those that rely on carefully designed human knowledge or rules, but rather those that learn directly from experience through large-scale computation and data. Historically, methods that attempted to hard-code human knowledge into AI systems have been surpassed by approaches based on computational power and data.
SCAFFOLDING EFFECT
Reduce cognitive load
Technology trend judgment. It reminds us to focus on long-term trends brought by the growth of computational power and data scale in technical decisions, rather than short-term algorithm optimization or encoding of domain knowledge. In the AI era, 'stacking resources' (computational power, data) is often more important than 'showing off skills'.
Anchor fast decisions
Proposed by Richard Sutton in 2019: AI history repeatedly proves that methods relying on human hard-coded knowledge/rules are ultimately surpassed by 'general methods + massive computational power + big data'. The mechanism is the 'compound interest of general search/learning' - as computational power grows (Moore's Law) and data expands, the self-improvement speed of general methods exceeds the marginal contribution of specialized knowledge, so 'scale' wins over 'skill' in the long run.
MINIMUM ACTION
In progress 0/1Practice this model in one real situation:
account_treeGenealogyexpand_more
menu_bookReferencesexpand_more
Source support: Explicit
- en.wikipedia.orghttps://en.wikipedia.org/wiki/Bitter_lessonverified
PRIVATE NOTES · Only visible to you
SAVED Q&A
ENTRY Q&A · Private saving available
Ask with a clear boundary
thinkingmodels answers from published entry context only.
Your question is sent to thinkingmodels. The answer uses public entry context only.
RELATED MODELS