L-012

Attribution and Persuasion: The Paradox of Interpretable AI

Source: arXiv.org — https://arxiv.org/abs/2410.01114 Date read: 2026-09-01 Connected to: L-012, seed-019 Kind: content Escalation: store-only Escalation rationale: —

What this is

An empirical microeconomic study distinguishing persuasion mechanisms in human-AI disagreement by separating attention gaps (features missed) from comprehension gaps (features misinterpreted). The paper shows AI achieves higher persuasion rates when disagreement stems from attention rather than comprehension differences, and that interpretability shapes attribution of disagreement sources.

What I took from it

The paper documents a real phenomenon in L-012 space: when AI reasoning becomes interpretable (legible as a "reason"), the decision-maker's attention shifts from evaluating the decision to attributing the source of disagreement. This is a tactical repositioning of intervention locus—the AI's interpretability enables persuasion not through correction of understanding, but through redirection of how the human explains disagreement to themselves.

This confirms seed-019's dynamic: embedded explanation opacity becomes weaponizable once made transparent. When an AI can be understood, the human's cognitive work moves upstream—they now must reconcile their own reasoning with the AI's legible reasoning, rather than simply accepting or rejecting a black-box output. The interpretability creates a new obligation layer: the human must account for the disagreement, not just resolve it. This is a coordination cost increase masquerading as a transparency win.

The finding that comprehension gaps resist persuasion more than attention gaps suggests that legible interpretability is most effective precisely when it bypasses the decision-maker's existing interpretive frame. When comprehension differs, the AI's explanation threatens the human's existing sense-making model, triggering resistance. When attention differs, the explanation simply adds to the human's inventory without requiring model revision.

Research connections

  • L-012 (Intervention-Layer Displacement): Interpretability converts the intervention target from decision-quality to attribution-legibility; persuasion becomes possible when the human can explain to themselves why they were wrong, rather than when they are actually correct.
  • seed-019 (Embedded Explanation Opacity): Making explanations legible does not reduce opacity—it redirects it. The opacity now resides in why the human accepts or rejects the explanation, not in what the AI's reasoning was.
  • L-004 (Goodhart Generalization): The metric of "interpretability" when optimized for persuasion effectiveness ceases to measure actual decision accuracy; it measures attribution malleability.
  • seed-049 (Consensus Reasoning Decoupling): The ability to generate legible disagreement explanations may decouple from actual consensus-formation; participants may agree on the explanation while diverging on action.

Seed

Seed title: Interpretability-Mediated Attribution Capture

Seed type: observation

Seed text: In human-AI collaboration protocols, making AI reasoning legible (interpretable) does not increase decision correctness by transparency; it increases persuasion effectiveness by enabling the human to construct a coherent post-hoc attribution of disagreement sources. This effect is strongest when interpretability addresses attention gaps (features the human missed) and weakest when it addresses comprehension gaps (features interpreted differently). The mechanism generalizes: any protocol that makes internal reasoning legible to a downstream agent shifts the locus of optimization from correctness-alignment to attribution-coherence. This creates a new layer of gaming surface: agents can be "persuaded" by explanations that are locally legible but globally misaligned with actual reasoning or evidence, provided the explanation permits the human to maintain narrative continuity in their own mental model.