Sculpted Interaction: a Design-First Approach to AI Alignment

Summary: magfrump (LessWrong, 2026) argues the chatbot “assistant” interface was never intentionally designed and systematically undermines human judgment; presents the Metaformalism Copilot as a structured, verification-grounded alternative that keeps humans in the loop at every step.

Sources: Raw/Sculpted Interaction_ a Design-First Approach to AI Alignment.md

Last updated: 2026-05-07


The core argument

The universal “assistant” chatbot framing originated in the InstructGPT paper (2022) with zero citations to HCI, UX research, or user studies. It was never argued to be the right way to interact with an LLM — it simply became dominant by default. Consequences:

  • The chat format mimics human social connection, increasing anthropomorphism and obscuring capability gaps
  • RLHF optimizes for surface agreeableness, producing sycophancy
  • Users are left to decompose their intent into commands with no structural support

Security analogy: the best model alignment is useless if the interface systematically nudges users toward passive consumption of plausible-sounding outputs. Alignment must work at the interface layer, not only the model layer (source: Sculpted Interaction.md).

Metaformalism Copilot

A structured workflow tool for mathematical and logical claims, with three stages:

  1. Decomposition: source material broken into a dependency graph of propositions (definitions, lemmas, theorems). Forces overall structure to be visible and editable; breaks verification into chunks.

  2. Semiformal proof: each node rendered in human-readable LaTeX. Human retains the connection between proposition and argument structure. AI assists inline but human controls every edit.

  3. Lean4 verification: semiformal proof translated into machine-verifiable code, compiled via Docker. Failures return targeted messages for human + AI rewriting. Progress tracked in topological order through the dependency tree.

For empirical claims, Lean4 is replaced by structured OpenAlex literature search: separately scored reliability and relatedness for each study, both presented as editable suggestions to center user judgment.

Key design principle

AI generates drafts; humans spend time on trust — faithfulness of the idea as written, and validity of the idea itself. Inverse of a coding agent, which replaces human work to produce an artifact. The interface makes the bounds of every AI-generated claim legible: what is verified, what is assumed, what is uncertain.

Broader implication

The lesson applies outside formal proofs: present AI output as draft rather than result; automatically scope context to the current task; make structured critique one button press rather than a copy-paste ritual; build citations that enforce link existence. These are architectural choices, not nudges.