Datasaurus Dozen
Updated 2026-07-31
INTRODUCTION
English translation pending.
CORE DEFINITION
Created by Justin Matejka and George Fitzmaurice at Autodesk, the Datasaurus Dozen extends Anscombe's quartet to twelve datasets plus the original dinosaur. All of them share nearly identical means, variances, and correlations, yet their plots range from a dinosaur to a star; the lesson is that summary statistics must never replace looking at the raw distribution.
SCAFFOLDING EFFECT
Reduce cognitive load
- Summary distrust: never sign off on a report that shows only means, variances, and correlations. - Plot first: inspect the raw distribution before you draw any conclusion. - Review habit: pair every written summary with a scatterplot your audience can actually read.
Anchor fast decisions
Twelve satirical charts expose how naive an organization's data capability really is, for instance assuming that buying a Hadoop cluster counts as having big data. Humor lowers defenses, so the mismatch between claimed and actual maturity becomes visible and discussable.
MINIMUM ACTION
In progress 0/1Practice this model in one real situation:
account_treeGenealogyexpand_more
menu_bookReferencesexpand_more
Source support: Explicit
- en.wikipedia.orghttps://en.wikipedia.org/wiki/Datasaurus_dozenverified
PRIVATE NOTES · Only visible to you
SAVED Q&A
ENTRY Q&A · Private saving available
Ask with a clear boundary
thinkingmodels answers from published entry context only.
Your question is sent to thinkingmodels. The answer uses public entry context only.
RELATED MODELS