Bellman Equation
Updated 2026-08-05
INTRODUCTION
English translation pending.
CORE DEFINITION
Named after Richard Bellman, the equation expresses the value of a current state as the sum of two terms: the immediate reward received now and the discounted expected value of the state reached by the next optimal action. Because the future term is itself defined by the same relation, the value of any state depends on the entire chain of decisions that follows, which makes the equation the mathematical form of long-term thinking. Solving it yields the policy that maximizes total discounted reward.
SCAFFOLDING EFFECT
Reduce cognitive load
- Price the destination: value a choice by the worth of the state it delivers, not only its payoff now. - Solve backward: compute from the desired end state toward the present decision. - Choose the discount rate: pick explicitly how heavily you weight the future relative to today.
Anchor fast decisions
The equation breaks a multi-step choice into a single step plus the rest, and defines the value of the current state through the best achievable value of the next one. Discounting converts future rewards into present terms, so each stage's contribution shrinks with distance. A policy is optimal exactly when every action maximizes the sum of immediate reward and the discounted continuation value, which is why the whole chain, not the single step, determines what the current choice is worth.
MINIMUM ACTION
In progress 0/1Practice this model in one real situation:
account_treeGenealogyexpand_more
menu_bookReferencesexpand_more
Source support: Explicit
- zh.wikipedia.orghttps://zh.wikipedia.org/wiki/%E8%B2%9D%E7%88%BE%E6%9B%BC%E6%96%B9%E7%A8%8Bverified
PRIVATE NOTES · Only visible to you
SAVED Q&A
ENTRY Q&A · Private saving available
Ask with a clear boundary
thinkingmodels answers from published entry context only.
Your question is sent to thinkingmodels. The answer uses public entry context only.
RELATED MODELS