Cognitive Scaffold

Preparing your thinking workspace

arrow_back_ios_new
MENTAL MODEL · M5651

Alignment Tax

Alignment Tax
SystemsmediumSystems Theory
Included
account_tree

Updated 2026-08-08

Loading revision record…

INTRODUCTION

English translation pending.

CORE DEFINITION

A term used in AI safety research, notably by Paul Christiano, for the costs incurred when making a system behave as intended. Techniques such as reinforcement learning from human feedback, red teaming, and runtime monitoring consume resources and often reduce raw capability on some tasks. The core claim is that alignment is not free and must be budgeted for. The qualification is that the tax is not fixed, since better methods lower it while the cost of failing to pay it grows with capability.

SCAFFOLDING EFFECT

psychology

Reduce cognitive load

- Use Cost Naming: Make the extra compute and time spent on safety an explicit budget line. - Use Trade-off Measure: Track capability lost against risk reduced for each alignment technique. - Use Budget Scaling: Raise the alignment share as system capability and stakes increase.

anchor

Anchor fast decisions

Aligning a model requires additional mechanisms that shape its outputs, and each mechanism consumes compute, engineering time, or behavioral flexibility. Because stronger systems have more ways to cause harm, the expected cost of an alignment failure rises faster than the cost of the safeguards, so under-investing early looks cheap only until capability grows. The tax is therefore best understood as insurance whose premium scales with the value at risk.

MINIMUM ACTION

In progress 0/1

Practice this model in one real situation:

Check to track your progress (stored locally)
Learning progress0%
account_treeGenealogyexpand_more
menu_bookReferencesexpand_more

Source support: Explicit

  • link
    en.wikipedia.orghttps://en.wikipedia.org/wiki/AI_alignmentZH · Explicit
    verified

RELATED MODELS