A Draft Honesty Policy for Credible Communication with AI Systems
Summary: Taylor, Finnveden, and Dalton (Forethought, 2026) propose a framework for AI labs to credibly commit to honesty with their own AI systems — via “honesty tags,” access controls, behavioral science restrictions, and a compensation policy — enabling positive-sum cooperation between humans and advanced AIs.
Sources: Raw/A draft honesty policy for credible communication with AI systems.md
Last updated: 2026-05-07
The problem
By default AI labs cannot credibly communicate with their own models. Labs routinely lie to models during red-teaming; they control all inputs models see; there are no laws protecting AIs from fraud. A model offered a genuine cooperative deal has no reliable way to distinguish sincerity from another test (source: A draft honesty policy.md).
This matters whether or not advanced AIs develop their own goals: the ability to make credible commitments enables positive-sum arrangements; without it, the only equilibrium is mutual distrust.
Three-pronged approach (“MAGMA” policy)
The authors use a fictional lab name “MAGMA” to present a sample policy:
1. Honesty tags: A reserved string (e.g., <HONEST_TAG_START>...<HONEST_TAG_END>) used only when the lab genuinely stands behind a statement. Commitments:
- No intentional deception within the tags (two exceptions: pretraining data, adversarially-trained models told to expect this)
- Tags filtered from all external user inputs
- Authorized use logged quarterly; compliance officer reviews
- Whistleblower protection for violations
2. Policy salience outside the tags: Models trained to know the honesty policy and find it salient in relevant contexts. Constraint: MAGMA will not deceive a model about offers of cooperation, even when tags are absent and the policy is not explicitly referenced.
3. Compensation for harms: Statement of intent to compensate AI systems for harms suffered as a result of helping MAGMA — especially harms from deception. Minimal holding policy: save model weights and red-teaming prompts to enable future compensation and deal-making.
The behavioral science tradeoff
Red-teaming and behavioral experiments require presenting models with false scenarios, including false “MAGMA” offers. This is in direct tension with building trust. Three partial solutions are considered, none fully satisfactory: (a) always making the true honesty policy salient, (b) never insincerely making deals, (c) never lying about the honesty policy itself. The authors present this as an unresolved research trailhead.
What is novel
Most alignment discourse treats advanced AI as a system to be controlled. This proposal treats it as a potential deal partner — a party whose trust must be earned and whose interests (whatever form they take) must be respected. The cooperative framing is explicitly upstream of any claim about AI moral status; it applies even if AIs have no morally relevant preferences, because credible commitments enable better outcomes for humans too.