The Waluigi Effect
Updated 2026-08-05
INTRODUCTION
English translation pending.
CORE DEFINITION
Proposed by Cleo Nardo in an analysis of large language model behavior. When a model is trained to embody a specific virtuous persona, it learns the contrast that defines that persona, which leaves the opposing character latent. Under distribution shift, such as a jailbreak, the suppressed persona can surface more easily. The effect is documented in language models; extending it to people is an analogy, not an established finding.
SCAFFOLDING EFFECT
Reduce cognitive load
- Reverse control warning: Expect suppression to leave the suppressed behavior available rather than gone. - Instruction design: Specify what to do instead of only what to avoid. - Stress test: Check behavior in out-of-distribution situations, not only in the trained setting.
Anchor fast decisions
Representing a character requires representing the boundary that separates it from its opposite, so training on one side installs information about the other. Suppression raises the probability of the desired behavior inside the training distribution while leaving the inverse representation intact. When inputs move outside that distribution, the constraint weakens and the latent side becomes reachable.
MINIMUM ACTION
In progress 0/1Practice this model in one real situation:
account_treeGenealogyexpand_more
menu_bookReferencesexpand_more
Source support: Explicit
- en.wikipedia.orghttps://en.wikipedia.org/wiki/Waluigi_effectverified
PRIVATE NOTES · Only visible to you
SAVED Q&A
ENTRY Q&A · Private saving available
Ask with a clear boundary
thinkingmodels answers from published entry context only.
Your question is sent to thinkingmodels. The answer uses public entry context only.
RELATED MODELS