Alignment Tax
Updated 2026-08-08
INTRODUCTION
English translation pending.
CORE DEFINITION
A term used in AI safety research, notably by Paul Christiano, for the costs incurred when making a system behave as intended. Techniques such as reinforcement learning from human feedback, red teaming, and runtime monitoring consume resources and often reduce raw capability on some tasks. The core claim is that alignment is not free and must be budgeted for. The qualification is that the tax is not fixed, since better methods lower it while the cost of failing to pay it grows with capability.
SCAFFOLDING EFFECT
Reduce cognitive load
- Use Cost Naming: Make the extra compute and time spent on safety an explicit budget line. - Use Trade-off Measure: Track capability lost against risk reduced for each alignment technique. - Use Budget Scaling: Raise the alignment share as system capability and stakes increase.
Anchor fast decisions
Aligning a model requires additional mechanisms that shape its outputs, and each mechanism consumes compute, engineering time, or behavioral flexibility. Because stronger systems have more ways to cause harm, the expected cost of an alignment failure rises faster than the cost of the safeguards, so under-investing early looks cheap only until capability grows. The tax is therefore best understood as insurance whose premium scales with the value at risk.
MINIMUM ACTION
In progress 0/1Practice this model in one real situation:
account_treeGenealogyexpand_more
menu_bookReferencesexpand_more
Source support: Explicit
- en.wikipedia.orghttps://en.wikipedia.org/wiki/AI_alignmentverified
PRIVATE NOTES · Only visible to you
SAVED Q&A
ENTRY Q&A · Private saving available
Ask with a clear boundary
thinkingmodels answers from published entry context only.
Your question is sent to thinkingmodels. The answer uses public entry context only.
RELATED MODELS