Policy Iteration
Updated 2026-08-13
INTRODUCTION
English translation pending.
CORE DEFINITION
Strategy iteration treats a strategy as a continuously revised hypothesis rather than a fixed plan. Its core proposition is that an initial strategy is necessarily imperfect because it is formed under incomplete information, so the decisive capability is a reliable loop of executing, measuring, and adjusting rather than the brilliance of the first plan. The key qualification is direction: iteration improves a strategy only when each cycle is measured against a stable objective, otherwise repeated adjustment becomes undirected churn. The idea connects to build-measure-learn practice in lean startup work and to the observe-orient-decide-act loop.
SCAFFOLDING EFFECT
Reduce cognitive load
- Cycle design: define the loop length and the metric that closes each cycle. - Small bets: launch the smallest version that yields usable feedback and a measurable signal. - Revision rule: decide in advance what evidence would trigger a change of course.
Anchor fast decisions
Plans formed without feedback encode assumptions that may be wrong, and those errors compound over time. Short cycles convert assumptions into testable predictions and return evidence quickly, so corrections arrive while they are still cheap. Each revision narrows the gap between the plan and reality, which is what makes a modest starting strategy competitive.
MINIMUM ACTION
In progress 0/1Practice this model in one real situation:
account_treeGenealogyexpand_more
menu_bookReferencesexpand_more
Source support: Explicit
- en.wikipedia.orghttps://en.wikipedia.org/wiki/Markov_decision_processverified
PRIVATE NOTES · Only visible to you
SAVED Q&A
ENTRY Q&A · Private saving available
Ask with a clear boundary
thinkingmodels answers from published entry context only.
Your question is sent to thinkingmodels. The answer uses public entry context only.
RELATED MODELS