Counterfactual Explanation
Updated 2026-08-11
INTRODUCTION
English translation pending.
CORE DEFINITION
A counterfactual explanation describes a model's decision by stating the minimal change to the input that would have produced a different output, not why the loan was rejected but what would have had to be true for approval. Its value is that it matches the form in which people naturally ask why, and it is actionable in a way a list of feature weights is not, because it names the specific lever and the distance to the decision boundary.
SCAFFOLDING EFFECT
Reduce cognitive load
- Name the lever: identify which feature controls the decision boundary. - Compute the distance: find the smallest change that flips the output. - Check it is real: confirm the counterfactual lies within the plausible data distribution.
Anchor fast decisions
People understand a decision by contrasting it with a nearby alternative, which is exactly the structure a counterfactual supplies. Feature weights describe the model's interior, while a counterfactual describes the boundary in terms of a change the subject could actually make, which is what converts an explanation into a possible action.
MINIMUM ACTION
In progress 0/1Practice this model in one real situation:
account_treeGenealogyexpand_more
menu_bookReferencesexpand_more
Source support: Explicit
- doi.orghttps://doi.org/10.48550/arXiv.1711.00399verified
PRIVATE NOTES · Only visible to you
SAVED Q&A
ENTRY Q&A · Private saving available
Ask with a clear boundary
thinkingmodels answers from published entry context only.
Your question is sent to thinkingmodels. The answer uses public entry context only.
RELATED MODELS