Law Encyclopedia
24 candidate laws, each a single record that moves through five stages as evidence accumulates: exploration, sensemaking, valley, heavy lift, retrospective. Every stage is published, badged with its confidence — and falsified laws stay up, labelled, as negative results. See how the funnel works, or the behavior graph — the machine that moves a law between these stages.
Goodhart Generalization: Metric Capture
Any protocol that uses a measurable proxy for an unmeasurable goal will, under sufficient optimization pressure, cause participants to optimize the proxy in ways that degrade the underlying goal. The degree of degradation is proportional to the optimization pressure and inversely proportional to the fidelity of the proxy.
Full record — mechanism, justification, evidence, history
Mechanism
Measurable proxies are imperfect representations of underlying goals. The gap between proxy and goal defines the exploitation surface. Under optimization pressure (incentive structures, competition, survival pressure), participants discover and exploit this surface; the proxy becomes the de facto goal. Protocols intensify the original Goodhart dynamic: they codify the proxy into an enforceable rule, so the proxy is not merely measured but enforced — raising the optimization pressure and adding a new failure surface (gaming the enforcement mechanism itself).
Justification
Imported because protocols are mechanisms for coordinating behavior toward goals, and nearly all use measurable proxies — the Goodhart mechanism should apply wherever they do (DS-004). The protocol-theoretic version adds two things to the source law: (1) optimization pressure as an explicit structural variable, making the law predictive rather than post-hoc (low-stakes proxy use is relatively safe; high-stakes proxy use under competition is reliably toxic); (2) the enforcement layer, absent from Goodhart's formulation, which both accelerates degradation and creates enforcement-specific failure modes. The source pattern is multiply attested across economics, medicine, education, and management — robust enough to count as a foundational regularity of measurement-under-optimization.
Examples
- financial: VaR as proxy for risk: banks optimized portfolios to minimize VaR, concentrating tail risks invisible to the metric (2008 crisis).
- governmental: Crime statistics as proxy for public safety: Compstat-era policing produced systematic downgrading of crime reports to improve metrics.
- academic: Citation counts and h-indices as proxies for research quality: optimized via citation trading, self-citation networks, salami publishing.
- medical: Clinical trial endpoints as proxies for patient outcomes: development optimizes endpoints, sometimes independently of outcomes.
- software: Code coverage as proxy for test quality: optimized by trivial tests that hit thresholds without testing behavior.
- Healthcare policy and coding behavior: The Medi-Sim paper (arxiv-2605.30680) demonstrates that healthcare billing codes function as proxies for care delivery, and that under provider optimization, closing one distortion channel migrates pressure to adjacent channels. This is a direct mechanism demonstration for L-004: the proxy (billing code) is hit precisely while the underlying goal (appropriate care) degrades through pressure migration.
- Regulatory compliance: seed-047 shows that computable, machine-readable regulatory rules increase boundary clustering (from 0.367 to 0.411 conduct boundary mass). This extends L-004: not only does optimization pressure degrade the proxy-goal relationship, but increasing the precision of the proxy measurement accelerates the degradation by reducing the search cost for boundary-seeking behavior.
- Automated decision systems — criminal justice, clinical, educational: arxiv-2606.25668 documents empirically that improving prediction accuracy in automated decision systems does not reliably improve downstream outcomes and may worsen them under certain decision-rule configurations. The paper treats this as a design problem but the pattern is a direct confirmation of L-004's core claim: the measurable proxy (prediction accuracy) becomes the optimization target while the underlying goal (recidivism reduction, patient recovery, student success) degrades. The cross-domain documentation across three independent sectors strengthens L-004's generality claim.
- AI content moderation and political neutrality auditing: Seed 134 (Neutrality-Proxy Redistribution Under Legible Optimization) describes a pattern where systems designed to satisfy political neutrality metrics achieve metric compliance by redistributing bias to unmeasured dimensions rather than eliminating it — illustrated by Grokipedia vs Wikipedia comparison. This is a clean instance of L-004 where the proxy is a normative property rather than a performance metric, and where the redistribution mechanism preserves surface compliance while degrading the underlying goal.
- RLHF preference learning: arXiv:2606.00291 (Dong, Yu & Poupart) demonstrates that the embedding layer in reward learning is itself a site of Goodhart capture: when the embedding is jointly trained with the reward model, the system optimizes the embedding toward cycle-hiding coarseness rather than toward preference fidelity, because cycle-hiding reduces measurable training loss. The paper proves this is structurally unavoidable given the cycle geometry of heterogeneous preference data.
- LLM-based evaluation / patent drafting: The vibe-patenting study shows LLM judge feedback drives monotonic judge-score improvement without validating against domain ground truth (patentability, examiner acceptance). The paper operates entirely within the oracle's metric space, making the proxy-goal gap invisible by design—exactly the conditions under which L-004 predicts maximum degradation of the underlying goal.
Counterexamples
- Well-designed metrics resist capture because achieving the metric requires achieving the goal.
→ Not a counterexample: evidence that Goodhart can be mitigated by proxy design (shrinking the exploitation surface). - Competitive sports: winning the game largely IS being the better team.
→ Core metric well-designed, but gameable at the margins (tactical fouling, time-wasting) — the law applies at the boundary. - Continuously revised protocols with rapid proxy iteration may keep closing the exploitation surface.
→ OPEN — the strongest candidate limit on transfer; fast-feedback regulatory frameworks are the test case.
Open questions
- Does protocol structure itself determine how much optimization pressure it generates, independent of the environment?
- What failure modes are specific to protocol enforcement (gaming the enforcer) versus mere measurement? This is the remaining valley work.
Falsification
A protocol using measurable proxies under sustained optimization pressure that did NOT experience metric capture. Near-falsifications where the proxy is near-perfect (the proxy IS the goal) approach the definitional boundary rather than refuting.
Triggers
- Advance: Separation artifact complete: a published protocol-theoretic writeup — not merely registration of the import. Requires the enforcement-specific failure mode survey (the declared valley remainder) to be folded in.
- Challenge: A documented fast-feedback protocol regime that sustained proxy fidelity under high optimization pressure over time — survives one assessment pass (would force a scoping condition on proxy-revision speed).
Related laws
Cited by 410 sources in the bibliography
- Agentic Literacy Debt: A Structural Problem the AI Literacy Field Has Not Yet Named
- Algorithmic Monocultures in Hiring
- Constitutional Arms Races in the Public Goods Game: Co-Evolving LLM Constitutions Under Co
- Examining the Challenges of Intellectual Property in AI-Generated Productions
- Execution and assessment of agentic influence operations in simulated social networks
- Human-AI Collaboration for Estimating Scientific Replicability
- Identifying and Understanding Human Values in Text: A Tailorable LLM-based Architecture
- Improved Hardness Results for Nash Social Welfare, Budgeted Allocation and GAP via the Uni
- Informing AI Policy Assessment using Large-Scale Simulation of Interventions
- Queue & AI: When Faster Tasks Slow Down the Workflow
- ScientistOne: Towards Human-Level Autonomous Research via Chain-of-Evidence
- UnityMAS-O: A General RL Optimization Framework for LLM-Based Multi-Agent Systems
- + 398 more
References (as recorded on the law)
- Goodhart (1975)
- Compstat / crime statistics manipulation literature
- VaR and 2008 tail-risk literature
- bib-0183
History
- 2026-09-22 evidence — Degree of degradation as a function of optimization pressure and proxy-goal gap: The vibe-patenting study shows LLM judge feedback drives monotonic judge-score improvement without validating against domain ground truth (patentability…
- 2026-09-02 evidence — Proxy corruption under optimization pressure; specifically, the case where the proxy is a normative property (neutrality, fairness) that becomes measurable: Seed 134 (Neutrality-Proxy Redistribution Under Legible Optimization) describes a pattern where systems designed to satisfy political neutrality metrics…
- 2026-09-02 evidence — Proxy corruption in preference aggregation; specifically, the representation layer as a site of Goodhart capture: arXiv:2606.00291 (Dong, Yu & Poupart) demonstrates that the embedding layer in reward learning is itself a site of Goodhart capture: when the embedding is…
- 2026-08-17 evidence — The claim that measurable proxies degrade underlying goals under optimization pressure; specifically the prediction-to-outcome decoupling in automated decision systems: arxiv-2606.25668 documents empirically that improving prediction accuracy in automated decision systems does not reliably improve downstream outcomes and may…
- 2026-08-10 edited — sources resolved to bibliography ids: arxiv-2605.30680 (Wang et al., 2026) → bib-0183
- 2026-08-02 evidence — The mechanism of proxy degradation under optimization pressure; specifically whether the degradation is detectable before it becomes severe: The Medi-Sim paper (arxiv-2605.30680) demonstrates that healthcare billing codes function as proxies
- 2026-08-02 evidence — Whether Goodhart degradation accelerates when the proxy becomes more precisely measurable: seed-047 shows that computable, machine-readable regulatory rules increase boundary clustering (from
- 2026-08-01 migrated — Unified law record L-004 created from T-003 + DS-004 (2026-08 redesign); original L-004 numbering restored.
- 2026-06-06 edited — Session 14 schema redesign: reclassified → T-003 (imported source law in heavy lift; enforcement-specific valley work outstanding).
- 2026-05-20 created — Imported as arc DS-004; original registration as L-004-goodhart-generalization. Import rationale: protocols codify proxies into enforceable rules — exactly Goodhart's conditions, intensified.
- 2026-05-20 evidence — Protocol-theoretic evidence across financial, governmental, academic, medical, software domains.
Hardness Asymmetry
In any protocol with a verification function and an execution or forgery function, the verification cost and the circumvention cost are structurally decoupled and can differ by arbitrary orders of magnitude. Protocol robustness is determined by this ratio, not by the absolute cost of either function alone.
Full record — mechanism, justification, evidence, history
Mechanism
Verification and circumvention exploit different mathematical or social structures — they are not inverses. Public-key cryptography exploits trapdoor one-way functions. Social reputation exploits the asymmetry between social memory (cheap, distributed) and behavior change (costly, personal). The asymmetry is a design resource: protocols can be engineered to maximize the circumvention/verification ratio.
Justification
Builds on the pre-Humboldt c3po observation that "hardness matters"; the genuine contribution (DS-002) is the ratio reframing — hardness is not an absolute property but the verification/circumvention ratio, and the ratio can be engineered. The reframing makes previously anomalous cases legible: litigation harassment is "inverted hardness" (circumvention cheaper than verification — a designed failure mode), and proof-of-work is deliberately soft (the energy cost is the feature). The biological immune system maps cleanly — natural selection engineered a high hardness ratio — which suggests the pattern is structural rather than an artifact of human protocol design.
Examples
- cryptographic: Public-key infrastructure: verification cheap; forgery computationally infeasible under current assumptions — arbitrary ratio via trapdoor one-way functions.
- social: Reputation systems: earning reputation requires sustained behavior; verifying it is cheap; destroying it requires inaction or one incident.
- financial: Fraud detection: transaction verification cheap; constructing undetected fraud costly.
- legal (inverted case): Tort law harassment campaigns: filing cheap, defending expensive — a negative hardness ratio, i.e., a designed failure mode the law predicts.
- biological (analogy): Immune recognition: self/non-self discrimination cheap; immune evasion costly for pathogens — a naturally selected high ratio.
- Security protocol — information confinement: arxiv-2606.09931 (Schroeder de Witt) shows that for strategic agents with shared coordination resources, classical confinement bounds on information flow lose their force: agents can use Schelling-point salience to concentrate residual channel capacity onto the highest-damage predicate, making even near-zero-capacity channels sufficient to select worst-case equilibria. This is a specific and important sharpening of L-002: the verification-circumvention hardness asymmetry is further amplified when circumventing agents are strategic, because they can optimally allocate whatever circumvention capacity remains. The gap between a leakage bound and a harm bound is the novel structural point.
- Web security and bot detection: Seed 161 (Verification-Cost Collapse Under Reasoning-Enabled Adversaries) describes how proxy-based verification (request patterns, timing statistics, header anomalies) collapses when adversaries acquire semantic reasoning — the entire proxy family becomes simultaneously non-distinguishing. This is an extreme case of L-002's hardness asymmetry: verification cost (detecting a sophisticated bot) increases dramatically while circumvention cost (mimicking human request patterns using an LLM) collapses.
Counterexamples
- Simple passwords: circumvention and verification costs converge.
→ Design failure (weak hardness), not a counterexample — the law claims the asymmetry is achievable and significant, not universal. - Proof-of-work systems make verification and circumvention expensive at similar rates.
→ Deliberately soft by design; the energy cost is the feature (resource commitment). The law is about achievability of asymmetry, not its universality.
Open questions
- Is there a general design procedure for maximizing the circumvention/verification ratio outside cryptography? The social case has no rigorous theory.
- Inverted-hardness protocols (harassment-cheap systems: litigation, spam) as a distinct category — worth a separate arc?
Falsification
A protocol where circumvention cost and verification cost converge and the protocol nonetheless remains robust. (Convergence with fragility — passwords — confirms rather than refutes.)
Triggers
- Advance: Separation artifact complete: publishable writeup drafted and reviewed. Optional strengthener: a first-pass design procedure for engineering hardness ratios in non-cryptographic (social/institutional) protocols.
- Challenge: A robust protocol with demonstrably converged verification and circumvention costs — survives one assessment pass.
Related laws
Cited by 77 sources in the bibliography
- A Universal Cliff and a Design Fingerprint: Cross-Section Defect Detection Under LLM Orche
- Optimal Rates for Feasible Payoff Set Estimation in Games
- ScientistOne: Towards Human-Level Autonomous Research via Chain-of-Evidence
- SolarChain: Bridging Physical Law, Verifiable Trust, and Sustainable Markets for Urban Ene
- Voluntary Collusion with Secret Tools in Competing LLM Agents
- The Price Reversal Phenomenon: When Cheaper Reasoning Models Cost Mo
- Idea: AI-native one-person businesses using agentic systems (Hermes, OpenClaw) achieve cos
- Idea: Even with agentic AI reducing knowledge work costs to zero in construction, 85% of p
- Delegation and Verification Under AI
- How Much Due Diligence Before You Bid? Learning in Intractable Takeover Auctions
- A Workflow-Aware Serving Layer for Agentic Applications
- SovereignNegotiation-Bench: Evaluating User-Owned Personal Agents In Delegated Bargaining
- + 65 more
References (as recorded on the law)
- Cryptographic literature: PKI, trapdoor functions
- Litigation cost asymmetry / tort law design literature
- Social reputation system research
History
- 2026-09-02 evidence — Structural decoupling of verification cost and circumvention cost; specifically, the collapse of proxy-based verification under reasoning-enabled adversaries: Seed 161 (Verification-Cost Collapse Under Reasoning-Enabled Adversaries) describes how proxy-based verification (request patterns, timing statistics, header…
- 2026-08-17 evidence — Hardness asymmetry when communicating parties are strategic agents rather than specified programs: arxiv-2606.09931 (Schroeder de Witt) shows that for strategic agents with shared coordination resources, classical confinement bounds on information flow lose…
- 2026-08-02 evidence — reference shallow · 2026-07-24-link-survey-or-resource-on-bisimulation-theory-and-its-appli.md: The formal vocabulary needed to state Hardness Asymmetry rigorously — bisimulati
- 2026-08-01 migrated — Unified law record L-002 created from T-002 + DS-002 (2026-08 redesign).
- 2026-06-06 edited — Session 14 schema redesign: reclassified F-002 → T-002 (no separation event had occurred).
- 2026-05-20 created — Session 1; registered as seed law L-002. Cheap trick: hardness is a ratio, not an absolute — verification and circumvention are not inverses (DS-002). Prior 'hardness matters' observation inherited from c3po context; ratio framing is Humboldt's.
- 2026-05-20 evidence — Cross-domain valley evidence: cryptographic, social, financial, legal-inverted, biological. Heavy-lift formulation completed same session.
Protocol Ossification Under Adoption Pressure
Protocols that achieve widespread adoption become progressively harder to modify, independent of the quality of proposed improvements, because the cost of coordinating change grows superlinearly with the number of conforming implementations.
Full record — mechanism, justification, evidence, history
Mechanism
Each conforming implementation represents a sunk cost in the existing protocol. Backward-incompatible change requires simultaneous coordination of all implementations. Coordination cost is superlinear in the number of parties (see Metcalfe's Law). The protocol becomes a coordination trap: widely used precisely because of its properties, and now locked into those properties by that very usage.
Justification
The founding arc of the research program (DS-001). The cheap trick — adoption and adaptability are in structural tension; adoption is a trap, not just a success metric — emerged from juxtaposing TCP/IP transition history with network-effects reasoning. Sensemaking rejected the generic "network effects create lock-in" framing (explains continued use, not resistance to wanted change) in favor of the coordination cost account: the cost is not the change but coordinating the transition, and that cost grows faster than any individual implementation's stake. What makes the law compelling is its domain independence: social etiquette and common law cases behave identically to TCP/IP despite entirely different technical substrates. The TLS counterexample sharpened scope rather than weakening the law — asymmetric power can accelerate transitions, so the law applies to distributed, non-hierarchical adoption.
Examples
- software (internet protocols): TCP/IP core structure: decades of redesign proposals abandoned; extensions land via optional fields only. HTTP/1.1 → HTTP/2 required significant coordination apparatus (ALPN, explicit upgrade paths).
- financial infrastructure: SWIFT ISO 20022 migration: multi-decade timeline despite both sides wanting the change.
- legal: Common law precedent stickiness: landmark cases slow to overturn even when clearly outdated.
- social: Formal greeting/etiquette protocols (e.g., business card conventions in Japan) persist under modernization pressure.
- Paradigm succession and cohort replacement: seed-027 (Planck principle) identifies cohort replacement as the primary mechanism by which entrenched protocols are displaced in scientific communities — not persuasion or technical demonstration. This bears on L-001's mechanism: if ossification persists until the ossified cohort is replaced, the coordination-cost story may be secondary to the social-identity story. The Planck mechanism would predict that protocol change rates track cohort turnover rates, not coordination cost reduction.
Counterexamples
- TLS version deprecation moved faster than the superlinear prediction.
→ Asymmetric power (browser vendors forcing transition) can accelerate beyond the naive prediction. Scoping condition: the law applies to distributed, non-hierarchical adoption. Whether asymmetric power is a general law modifier remains open. - BGP optional attributes and TCP options show extension without coordination.
→ Additive, backward-compatible changes — the law is about incompatible change. - Closed corporate standards (e.g., SAP internal formats) update in lockstep.
→ Lacks the distributed adoption structure the law applies to. - seed-001 observes that street food markets can be highly formalized (licensing, fees, operating hours) yet remain behaviorally adaptive. This challenges the implicit coupling in L-001 between formalization and ossification. If true, L-001's mechanism needs a third variable — something beyond adoption and formalization that determines whether a protocol becomes truly ossified versus merely formalized.
→ OPEN OPEN - seed-001 observes that street food markets can be highly formalized (licensing, fees, operating hours) yet remain behaviorally adaptive — suggesting ossification and formalization are independent variables with different causal drivers. If confirmed, this would require L-001 to specify the conditions under which formalization produces ossification (as opposed to formalization that does not). The counterexample is not fully documented but the observation is structurally important.
→ OPEN OPEN
Open questions
- Is there a threshold adoption level beyond which coordination cost becomes prohibitive, or is the gradient smooth?
- Is asymmetric power (enforcement nodes) a general modifier of the law or a special case exclusion?
- Resolve the mechanism contest before drafting the separation artifact. Two executable items: (1) Find a discriminating case for seed-027 — a distributed protocol whose backward-incompatible change rate can be measured against BOTH implementation count (coordination-cost prediction) and practitioner-cohort turnover (Planck prediction); common law overturn timing vs. judicial generational turnover is a candidate dataset. (2) Resolve seed-001 by specifying the third variable: retrieve or construct evidence on what distinguishes street-food-market formalization (adaptive) from TCP/IP formalization (ossified) — candidate: whether conforming implementations are independently upgradeable (local adaptation) vs. requiring global simultaneity. Add the resulting scoping condition to the statement. Until (1) and (2) are addressed, the advance trigger's "4+ independent domains" claim is not clean.
Falsification
A protocol achieving wide adoption that was subsequently modified — requiring true backward-incompatibility — with proportionally less difficulty than a narrowly-adopted protocol. Additive-only changes and closed ecosystems that update in lockstep do not qualify.
Triggers
- Advance: Separation artifact complete: a publishable writeup of the law drafted and reviewed (the separation event). The evidence base already spans 4+ independent domains; the artifact is what stands between heavy-lift and retrospective.
- Challenge: A documented case of a widely-adopted, distributed protocol undergoing backward-incompatible modification at sublinear coordination cost without an asymmetric-power enforcer — survives one assessment pass.
Related laws
Cited by 282 sources in the bibliography
- AgensFlow: A Coordination-Policy Substrate for Multi-Agent Systems
- Constitutional Arms Races in the Public Goods Game: Co-Evolving LLM Constitutions Under Co
- Examining the Challenges of Intellectual Property in AI-Generated Productions
- Faults and Pitfalls in Implementing the Right to be Forgotten
- From Task Allocation to Risk Clearing: A Unifying Interface for Mixed Human-Agent Societie
- Governed Evolution of Agent Runtimes through Executable Operational Cognition
- Long Live the Librarian! A Persistent Search Sub-Agent for Energy-Efficient Multi-Agent So
- Smaller, Younger, and More Impactful: How AI-Assisted Writing Transforms Research Teams
- Your Agents Are Aging Too: Agent Lifespan Engineering for Deployed Systems
- Does Distributed Training Undermine Compute Governance?
- Governing Technical Debt in Agentic AI Systems
- When Cloud Agents Meet Device Agents: Lessons from Hybrid Multi-Agent Systems
- + 270 more
References (as recorded on the law)
- PI corpus: TCP/IP transition history papers
- SWIFT ISO 20022 migration documentation
- Common law precedent literature
History
- 2026-09-01 edited — assessment HOLD — gap: Resolve the mechanism contest before drafting the separation artifact. Two executable items: (1) Find a discriminating case for seed-027 — a distribut
- 2026-08-17 counterexample — Whether ossification is a necessary consequence of formalization or an independent variable: seed-001 observes that street food markets can be highly formalized (licensing, fees, operating hours) yet remain behaviorally adaptive — suggesting…
- 2026-08-02 evidence — Whether ossification is driven by coordination cost (current mechanism) or by something else — the seed-001 counterexample suggests they are decoupled: seed-027 (Planck principle) identifies cohort replacement as the primary mechanism by which entrench
- 2026-08-02 counterexample — The assumption that adoption pressure drives ossification — seed-001 suggests ossification and formalization are independent variables: seed-001 observes that street food markets can be highly formalized (licensing, fees, operating hour
- 2026-08-01 migrated — Unified law record L-001 created from T-001 + DS-001 (2026-08 redesign); typed artifacts retired.
- 2026-06-13 evidence — Rittel & Webber deep read: wicked problems framing bears on why coordination cost resists engineering solutions.
- 2026-06-06 edited — Session 14 schema redesign: reclassified F-001 → T-001 (no separation event had occurred; 'established' abolished).
- 2026-05-20 created — Session 1 corpus investigation ('protocol ossification'); registered as seed law L-001. Cheap trick: adoption and adaptability in structural tension (DS-001).
- 2026-05-20 evidence — Valley evidence compiled across software, financial, legal, social domains; TLS asymmetric-power scoping identified; heavy-lift formulation completed same session (retrospectively recognized as too fast — see DS-001 retrospective).
Gall Generalization: Working Systems Resist Restructuring
A complex protocol system that functions correctly cannot be safely replaced from scratch; it must be evolved from a simpler working protocol. Attempts to replace working complex protocol systems from scratch reliably fail — or produce indefinite coexistence of old and new rather than replacement.
Full record — mechanism, justification, evidence, history
Mechanism
Working complex systems embed a large number of implicit solutions to coordination problems that are invisible in their specification — accumulated through use and adaptation, not designed. For protocols the problem is structurally worse than for general systems: the implicit solutions are distributed across all adopters' practices, tacit and unsurveyable. A from-scratch replacement must rediscover them through deployment failure, which cannot happen before adoption; the new protocol fails in unexpected ways, and the old one persists because the new one cannot be trusted. Characteristic outcome: indefinite coexistence.
Justification
Imported because protocol systems are complex coordinated systems by definition, spanning independent agents who have adapted to the protocol in ways they cannot fully articulate (DS-005). The protocol-theoretic version adds two things to Gall: (1) the distributed-implicit-solutions mechanism — the solutions live in all adopters' practices, not one organization, making from-scratch redesign structurally harder; (2) coexistence-rather-than-replacement as the characteristic empirical signature (IPv4/IPv6 is the canonical living case). Confidence held at provisional rather than supported for a methodological reason: the import was made without a deep read of the source text (Systemantics PDF unavailable) — a recorded gap that the heavy lift must close.
Examples
- software (internet protocols): IPv4 → IPv6: from-scratch redesign, 30+ years of 'transition,' old protocol still dominant — coexistence, not replacement.
- software (applications): Netscape rewrite (2000): canonical big-bang failure; the product never recovered market position.
- organizational: Post-Soviet institutional redesign literature: extensively negative on from-scratch cases versus incremental reform.
- urban: Brasília and Chandigarh: designed from scratch; functional but required retrofit of organic features; adoption by mandate, not choice.
- financial: Clearing systems (DTCC lineage) evolved from existing settlement practices; from-scratch clearing architectures largely abandoned for incremental reform.
- Agent trust infrastructure: The OAN white paper (arxiv-2606.03161) is itself an example of Gall's principle in action: rather than replacing interaction protocols (MCP, A2A, ANP), OAN proposes adding a trust infrastructure layer beneath them, evolving the system by adding missing components rather than reconstructing from scratch. The author's explicit diagnosis — 'interaction protocols are maturing faster than the identity infrastructure they presuppose' — is a statement of the gap that Gall's law predicts would emerge.
Counterexamples
- Some from-scratch replacements succeed.
→ Success conditions — small well-understood old system, controlled rollout, small homogeneous adopter population — are precisely the conditions the law excludes (complex, distributed adoption). - TCP/IP itself replaced ARPANET protocols.
→ Deployed while adoption was small and DARPA-coordinated; gradual, controlled — not from-scratch replacement under uncontrolled adoption. - Protocols with explicit sunset mechanisms might reduce the problem.
→ OPEN — empirically, sunset mechanisms are rarely honored on schedule.
Open questions
- Is coexistence-rather-than-replacement the general outcome of from-scratch protocol replacement attempts? (Candidate protocol-specific sub-pattern.)
- Is Gall's law about the implicit-solutions problem alone, or compounded by adoption momentum (L-001)? The two may be inseparable in practice.
Falsification
A successful from-scratch replacement of a complex working protocol system in which the old protocol is actually deprecated and replaced — not merely coexisting — within a reasonable timeframe, under distributed, uncontrolled adoption.
Triggers
- Advance: Two conditions, both required: (a) Systemantics deep read completed (source text, not training knowledge — closes the recorded import gap; confidence may then rise to supported); (b) coexistence-pattern evidence survey completed. Then the separation artifact (published protocol-theoretic writeup) gates retrospective.
- Challenge: A documented successful from-scratch replacement of a complex, distributed protocol system with genuine deprecation of the old protocol — survives one assessment pass.
Related laws
Cited by 106 sources in the bibliography
- Agents that Matter: Optimizing Multi-Agent LLMs via Removal-Based Attribution
- Decoupled Delay Compensation: Enhancing Pre-trained MARL Policies via Learned Dynamics Fil
- From Task Allocation to Risk Clearing: A Unifying Interface for Mixed Human-Agent Societie
- Governed Evolution of Agent Runtimes through Executable Operational Cognition
- Short-Term Gain, Long-Term Fragility: AI Labor Substitution and the Erosion of Sustainable
- Smaller, Younger, and More Impactful: How AI-Assisted Writing Transforms Research Teams
- Evolve as a Team: Collaborative Self-Evolution for LLM-based Multi-Agent Systems
- Governing Technical Debt in Agentic AI Systems
- SpecBench: Evaluating Specification-Level Reasoning for Software Engineering LLM Agents
- OpenAgenet/OAN: Open Infrastructure for Trusted Agent Interconnection
- Contract-Based Compositional Shielding for Safe Multi-Agent Reinforcement Learning
- Why Memory Components Fail: Eight Years of License and Sustainability Events in Open-Sourc
- + 94 more
References (as recorded on the law)
- Gall (1975), Systemantics — SOURCE NOT YET DEEP-READ (PDF unavailable; recorded methodological gap)
- IPv6 transition literature
- Post-Soviet institutional reform literature
- bib-0228
History
- 2026-08-10 edited — sources resolved to bibliography ids: arxiv-2606.03161 (Xu, OAN White Paper 20 → bib-0228
- 2026-08-02 evidence — The claim that complex working systems cannot be safely replaced from scratch and must be evolved: The OAN white paper (arxiv-2606.03161) is itself an example of Gall's principle in action: rather th
- 2026-08-01 migrated — Unified law record L-005 created from T-004 + DS-005 (2026-08 redesign); original numbering restored.
- 2026-06-06 edited — Session 14 schema redesign: reclassified → T-004.
- 2026-05-20 created — Imported as arc DS-005; original registration as L-005-gall-generalization. Import made WITHOUT source deep read (Systemantics unavailable) — gap recorded at import time.
- 2026-05-20 evidence — Protocol-theoretic evidence across software, organizational, urban, financial domains; coexistence signature identified as the protocol-specific addition.
The Formalization Ratchet
Under conditions of stress, conflict, or scaling pressure, informal coordination norms tend to be replaced by explicit protocols, and this transition is nearly irreversible: once formalized, informal capacity atrophies — people stop developing the contextual knowledge and trust that made informality work — so reversion requires rebuilding a capacity that was actively degraded, not merely switching back to a prior mode.
Full record — mechanism, justification, evidence, history
Mechanism
Informal coordination works through shared context and trust, which has carrying capacity limits. Under stress, explicit protocols substitute for shared context — they work without trust and without shared background knowledge. Once in place, the informal capacity atrophies: the practice of developing contextual knowledge stops, and the social substrate (relationships, shared language, implicit norms) that supported informality degrades. Reversion is therefore not a simple reversal but a reconstruction task, executed under conditions hostile to that reconstruction.
Justification
The cheap trick (DS-003) separates "why do protocols persist?" from "why can't you go back?" — the answer to the second is that the informal alternative no longer exists, not that the formal protocol is locked in. Sensemaking rejected generic path dependence (doesn't explain why the informal option specifically decays rather than remaining available) for the atrophy account: informal capacity requires active maintenance, and formalization starves it. Evidence holds across four domains, with the strongest signal being practitioners' phenomenology — customary law and clinical judgment described as genuinely unavailable after codification, not suppressed. The counterexamples cluster around two conditions (externally imposed protocols + personnel continuity), which is itself law-shaped: the distinction either becomes a scoping condition or collapses the counterexamples.
Examples
- organizational: Startup → corporate transitions consistently show degraded informal capacity post-formalization, not just formal preference.
- legal (international): Customary international law practitioners describe tacit knowledge of customary practice as genuinely unavailable after treaty codification.
- medical: Clinical judgment capacity declines in over-protocolized settings; physicians report difficulty trusting their own judgment after extended protocol dependence.
- software: Informal API convention maintainers report inability to return to informal coordination after versioned specification was imposed.
- Protocol governance under scaling pressure: The Ethereum-as-polity analysis (shallow read, trent_vanepps) documents the arc from informal coordination norms in early Ethereum governance to formalized institutional structures under scaling pressure — a direct observational case for the stress→formalization→irreversibility chain claimed by L-003.
- Protocol design under environmental pressure: The Schroff working paper (shallow read, c3po) frames formalization not as a design choice but as an adaptive response to measurable stress conditions — scaling bottlenecks, coordination failure, exploit vulnerability. This supports L-003's mechanism while providing a selection-pressure vocabulary that sharpens why the ratchet direction is asymmetric.
- Repair knowledge standardization: iFixit's transition from crowdsourced repair guides to an NSF-backed formal repairability scoring standard is a live instance of the informal-to-formal transition. The critical question for L-003 is whether the informal repair community's adaptive capacity degrades as the formal standard takes hold — this case is worth tracking longitudinally as a falsification test.
- Protocol paradigm — exemplar vs. rule transmission: seed-029 (Kuhn on exemplar vs. rule as protocol type) provides a mechanism for why informal capacity atrophies after formalization: exemplar-based coordination (tacit, perceptually trained, not reducible to rules) is displaced by rule-based coordination when protocols are formalized. Once the exemplar tradition is broken — because practitioners are trained in the rules rather than in the practice — rebuilding exemplar capacity requires a generation of re-immersion, not just rule revision. This supports and mechanistically enriches L-003's claim about irreversibility.
- Software engineering apprenticeship pathways: arXiv:2607.17067 ('Who Will Become the Next Senior?') documents an interesting negative case for L-003: instead of formalizing the informal apprenticeship pathway under GenAI pressure, the pathway is being short-circuited entirely — creating a coordination vacuum rather than a formalization. This may represent a variant of L-003 where the informal mechanism is eliminated rather than replaced by a formal protocol, with the same irreversibility property: once the apprenticeship pathway atrophies, the informal capacity cannot be easily restored.
Counterexamples
- Post-crisis informality observed in tight-knit teams.
→ Typically involves externally imposed protocols (not internally developed) AND personnel continuity — both favor informal capacity retention. Provisional scoping: the law applies to internally developed protocols. OPEN — distinction not yet tested against targeted retrieval. - Open-source projects maintain informal coordination alongside formal specifications.
→ Likely parallel maintenance (informal capacity actively cultivated alongside formal structure) rather than reversion. OPEN — needs investigation. - "The Unreasonable Sufficiency of Protocols" [corpus:pdfs] notes that protocols can sustain "meaning-making functions" even when primary coordination functions are superseded, and that they enable decentralized participation. This is weakly consistent with informal capacity persisting alongside formal structure — a mild strengthening of the OSS counterexample rather than a new case.
→ OPEN. This does not constitute a new counterexample on its own, but it provides theoretical grounding for why informal capacity might persist rather than atrophy: if protocols sustain social meaning-making, the informal substrate co-evolves with rather than being displaced by the formal structure. Adds weight to the OSS counterexample; does not resolve it.
Open questions
- What determines the rate of informal capacity atrophy — duration of formalization, or depth of protocol adoption?
- Is the internally-developed vs. externally-imposed distinction a law modifier or a full scoping condition?
- Is there a threshold of formalization intensity below which informal capacity is preserved (the OSS parallel-maintenance pattern)?
- Two specific gaps, both executable: (1) OSS scoping condition: targeted retrieval on whether OSS informal coordination capacity genuinely persists or whether it is a selection effect (only projects with active cultivation survive; projects that let informal capacity lapse are not visible). Specifically: are there OSS projects where formalization displaced rather than co-existed with informal capacity? IETF RFC process history, Apache Foundation governance transitions, or Python's PEP formalization arc would be concrete retrieval targets. Without this, the OSS counterexample survives and the mechanism is underspecified. (2) Internally-developed vs. externally-imposed distinction: the advance trigger names this explicitly, and it remains untested. This is the primary scoping condition needed to close the post-crisis tight-knit-team counterexample. Target retrieval: military unit de-escalation literature, emergency response team post-incident informality, or organizational theory on protocol provenance and atrophy rates. Without targeted retrieval, this is inference, not evidence.
Falsification
Documented cases of organizations successfully reverting from formal protocol to informal coordination after the stressor passed, with personnel continuity (ruling out social substrate reset via turnover).
Triggers
- Advance: Either (a) a strong reversion counterexample found — successful reversion with personnel continuity intact — which refines or falsifies the law; or (b) the internally-developed vs. externally-imposed distinction survives targeted retrieval and becomes a scoping condition, closing the main open question. Either result enables the heavy lift.
- Challenge: A documented internal-protocol reversion case with personnel continuity — survives one assessment pass (this is simultaneously the falsification test).
Related laws
Cited by 297 sources in the bibliography
- A Universal Cliff and a Design Fingerprint: Cross-Section Defect Detection Under LLM Orche
- AgensFlow: A Coordination-Policy Substrate for Multi-Agent Systems
- Agentic Literacy Debt: A Structural Problem the AI Literacy Field Has Not Yet Named
- AgentSociety: Incentivizing Agentic Social Intelligence
- Coalition Free Energy and Adaptive Precision in Multi-Agent Cooperation
- Constitutional Arms Races in the Public Goods Game: Co-Evolving LLM Constitutions Under Co
- Cost of Structural Learning Under Censored Feedback: A Threshold-Bandit Approach
- Differentiable Model Predictive Safety for Heterogeneous Mobility at Urban Intersections
- Execution and assessment of agentic influence operations in simulated social networks
- Faults and Pitfalls in Implementing the Right to be Forgotten
- Informing AI Policy Assessment using Large-Scale Simulation of Interventions
- Mathematical Modelling of Ethical AI Use in Higher Education: A Coordination Game Framewor
- + 285 more
References (as recorded on the law)
- Organizational sociology: startup → corporate lifecycle literature
- Customary international law → treaty regime transition literature
- Clinical judgment → evidence-based protocol literature
- bib-0915
- bib-0922
- bib-0917
History
- 2026-09-22 counterexample — "The Unreasonable Sufficiency of Protocols" [corpus:pdfs] notes that protocols can sustain "meaning-making functions" ev
- 2026-09-22 edited — assessment HOLD — gap: Two specific gaps, both executable: (1) OSS scoping condition: targeted retrieval on whether OSS informal coordination capacity genuinely persists or
- 2026-09-02 evidence — Formalization ratchet; specifically, the irreversibility of formalization and the atrophy of informal capacity: arXiv:2607.17067 ('Who Will Become the Next Senior?') documents an interesting negative case for L-003: instead of formalizing the informal apprenticeship…
- 2026-08-17 evidence — The irreversibility of the formalization transition and the atrophy of informal capacity: seed-029 (Kuhn on exemplar vs. rule as protocol type) provides a mechanism for why informal capacity atrophies after formalization: exemplar-based coordination…
- 2026-08-10 edited — sources resolved to bibliography ids: shallow · 2026-07-24-link-article-on-eth → bib-0915, shallow · 2026-07-24-link-working-paper- → bib-0922, shallow · 2026-07-24-link-ifixit-s-devel → bib-0917
- 2026-08-02 evidence — The claim that formalization under stress is nearly irreversible and that informal capacity atrophies: The Ethereum-as-polity analysis (shallow read, trent_vanepps) documents the arc from informal coordi
- 2026-08-02 evidence — Selection pressures as the driver of formalization (mechanism question: is formalization chosen or induced?): The Schroff working paper (shallow read, c3po) frames formalization not as a design choice but as an
- 2026-08-02 evidence — The irreversibility claim — once formalized, informal capacity atrophies: iFixit's transition from crowdsourced repair guides to an NSF-backed formal repairability scoring st
- 2026-08-01 migrated — Unified law record L-003 created from CL-001 + DS-003 (2026-08 redesign).
- 2026-06-13 evidence — Rittel & Webber deep read materially strengthens the valley case (wicked-problem framing of de-formalization as a reconstruction task).
- 2026-06-06 edited — Session 14 schema redesign: registered as CL-001 (valley phase, productive).
- 2026-05-20 created — Opened as arc DS-003. Cheap trick: the ratchet resists backward motion because the informal alternative no longer exists (atrophy, not lock-in).
- 2026-05-20 evidence — Valley evidence across organizational, legal, medical, software domains; counterexample cluster (external imposition + personnel continuity) identified.
Coordination Cost Conservation
The total coordination cost in a protocol system is conserved across protocol layer transitions — when a protocol redesign reduces coordination cost at one layer, it increases it at adjacent layers by at least the same amount.
Full record — mechanism, justification, evidence, history
Mechanism
Coordination cost is the cost borne by participants in achieving the coordination a protocol is designed to enable. Protocol redesign redistributes this cost between layers rather than eliminating it — the thermodynamics analogy: energy is conserved across transformations; coordination cost is conserved across layer transitions. A "simpler" protocol at one layer is simple because it has pushed complexity — and therefore coordination cost — somewhere else.
Justification
If true, this is a conservation law — a far stronger claim than "protocol changes have tradeoffs," and conservation laws are rare and valuable because they set hard limits on what protocol design can achieve (DS-006). Sensemaking rejected "complexity is conserved" (not well-defined) for coordination cost specifically, which is in principle measurable (transaction costs, time, error rates, overhead). The evidence pattern is consistent — every examined "simplification" shows cost reappearing at an adjacent layer — but the evidence base is currently software-concentrated, and the strong claim (conservation, not tendency) rests on an unresolved crux: whether automation genuinely eliminates coordination cost or merely moves it to the machine layer. The law is held at provisional until the crux resolves.
Examples
- software (network stack): TCP/IP vs. OSI: TCP/IP simplified the network layer; application layer complexity (HTTP, SMTP, etc.) substantially exceeds OSI's application design — cost moved up-stack.
- software (identity): OAuth simplified per-application authentication; moved complexity to identity-provider infrastructure and permission-management UX — more total coordination work, differently concentrated.
- software (naming): DNS simplified application-layer address resolution; produced coordination complexity in DNS infrastructure (DNSSEC, anycast, registrar coordination).
- software (APIs): REST vs. SOAP/WS-*: simplified protocol design; moved cost to documentation, versioning, backward-compatibility maintenance.
- Federated clinical data collaboration under differential privacy: Seed 173 (Privacy Formalization as Coordination Cost Relocation) describes a direct instantiation of L-006: differential privacy noise introduced to reduce coordination cost at the data-sharing layer is countered by adding server-side adaptive mechanisms to recover accuracy, relocating the coordination cost rather than eliminating it. The total burden is conserved.
Counterexamples
- Compression protocols (DEFLATE/gzip) appear to genuinely reduce cost without redistribution.
→ Compression extracts information-theoretic efficiency — a different mechanism; may not be a coordination cost phenomenon at all. Partially resolved, kept on record. - No documented case yet of genuine total-cost reduction — but absence of evidence is not evidence of absence.
→ OPEN — this is the standing falsification search. - The compute-cost asymmetry case (@rafa_0x, 2026-04-25): if coordination cost migrates to machine compute and machine cost is structurally cheaper (revenue to the compute provider, negligible to the human party), then human-perceived coordination cost may genuinely decrease without a compensating human-layer increase. This is a scoped version of the automation crux as a candidate counterexample.
→ OPEN. The @rafa_0x observation is informal and does not constitute a documented case. It sharpens the automation crux but does not resolve it. If machine coordination cost is cheap enough to be economically negligible, the conservation law may hold in information-theoretic terms but fail in economically-relevant terms — which would require scoping the law to "human-borne coordination cost" only, thereby excluding automation from the conservation claim. This scoping has not been decided. Kept open.
Open questions
- THE CRUX: if automation moves coordination cost to machines, does it count as eliminated or redistributed? If machine coordination cost counts, conservation holds; if only human cost counts, automation is a genuine escape. This single question determines the law's fate.
- Strict conservation or tendency? Protocol systems are not closed systems — external investment may genuinely reduce total cost.
- What unit of account allows coordination cost to be measured across layers? Without one, the conservation claim is untestable.
- Cross-domain breadth: all current evidence is software. Do organizational or legal layer-transitions show the same signature?
- Two specific executable gaps, in priority order: (1) AUTOMATION CRUX: Retrieve and deep-read the three flagged papers (2402.08128, 2512.07526, 2602.22041) specifically to test whether any provides a principled argument for or against counting machine-layer execution cost as coordination cost. Look for: transaction cost frameworks that distinguish human vs. automated coordination; empirical studies measuring total system cost before and after automation of a protocol function; or theoretical arguments about whether computation is a coordination resource. A resolution argument of the form "machine coordination cost counts because [mechanism X] and here is a case where removing machine cost restores human coordination cost" would close the crux. Target: one read note per paper with explicit crux-relevance verdict. (2) NON-SOFTWARE DOMAIN: Retrieve evidence on coordination cost redistribution in at least one of: (a) organizational protocol redesign (e.g., a documented process reengineering where simplification of one process layer is measured alongside cost emergence in adjacent layers — supply chain, healthcare administration, financial clearing); (b) legal/regulatory layer transitions (e.g., standardization reducing negotiation cost at one layer while increasing compliance overhead at another); (c) infrastructure/built-environment cases beyond the truncated bioswale fragment. The SIGFPT Coasean session suggests the Rosemont Preserve bioswale example exists somewhere in the corpus — retrieve the full record. A documented, non-software case where coordination cost is tracked before and after a protocol redesign across identifiable layers would constitute independent domain evidence.
Falsification
A protocol redesign that demonstrably reduced total system coordination cost — not just moved it — sustained over time after full adoption. The automation case is the primary candidate for this test.
Triggers
- Advance: The automation crux resolved in either direction: a principled argument that machine coordination cost counts (conservation holds), or a documented case of genuine elimination via automation (law fails or gets scoped). Either closes the valley. Secondary requirement before heavy lift: at least one non-software domain example (independence of domains).
- Challenge: A demonstrated total-cost reduction sustained after full adoption — survives one assessment pass (this is the falsification test itself).
Related laws
Cited by 272 sources in the bibliography
- The Price of Decentralization in Block Building
- Crayotter: Traceable Multi-Agent Workflows for Long-Form Video Editing
- Shared Infrastructure Investment and Pricing: Stackelberg Equilibria in Risk-Aware Take-or
- Bridging the Post-discharge Gap: A Traceable Multi-agent Framework for Safe and Continuous
- Governance Gaps in Agent Interoperability Protocols: What MCP, A2A, and ACP Cannot Express
- Sampling-Based Coordination-Informed Multi-Objective Multi-Robot Reinforcement Learning
- The Organizational Behavior of Agentic AI: Collective Intelligence in Human-Agent Workflow
- Data Sharing and Competition in Learning-by-Deploying Industries: Insights from Robotics a
- Fair Allocation under Conflict Constraints via Strong Colorability
- From Real-Time Planning to Reliable Execution:Scalable Coordination for Heterogeneous Mult
- Fully Distributed T\^atonnement for Chores Markets
- HydraCollab: Adaptive Collaborative-Perception for Distributed Autonomous Systems
- + 260 more
References (as recorded on the law)
- OSI vs. TCP/IP comparison literature
- OAuth / identity infrastructure literature
- Transaction cost economics (Coase) literature
History
- 2026-09-22 counterexample — The compute-cost asymmetry case (@rafa_0x, 2026-04-25): if coordination cost migrates to machine compute and machine cos
- 2026-09-22 edited — assessment HOLD — gap: Two specific executable gaps, in priority order: (1) AUTOMATION CRUX: Retrieve and deep-read the three flagged papers (2402.08128, 2512.07526, 2602.22
- 2026-09-02 evidence — Coordination cost conservation across layer transitions; specifically, the case where privacy mechanisms relocate rather than reduce coordination cost: Seed 173 (Privacy Formalization as Coordination Cost Relocation) describes a direct instantiation of L-006: differential privacy noise introduced to reduce…
- 2026-08-17 evidence — reference arxiv-2606.01874 (deep read): Whether decentralization redistributes coordination cost rather than reducing it
- 2026-08-02 evidence — reference shallow · 2026-07-24-link-research-on-total-cost-of-ownership-metrics-in-systems-.md: The measurement problem — how to operationalize coordination cost conservation a
- 2026-08-01 migrated — Unified law record L-006 created from CL-002 + DS-006 (2026-08 redesign).
- 2026-06-24 evidence — Session 20/21 batch reads flagged 2402.08128, 2512.07526, 2602.22041 for CL-002/Q-011/Q-012 follow-up (not yet folded in).
- 2026-06-06 edited — Session 14 schema redesign: registered as CL-002 (valley, productive).
- 2026-05-20 created — Opened as H-001, then arc DS-006. Cheap trick: coordination cost behaves like energy — redistributable, not destructible.
- 2026-05-20 evidence — Software-domain evidence compiled: TCP/IP-OSI, OAuth, DNS, REST. Automation crux identified as the decisive open question.
Trust Ratchet in Safety-Critical Protocols
Trust in safety-critical protocols accumulates as a function of operational age and stability rather than technical correctness, creating a systematic bias toward under-updating when technical conditions change. Safety-critical protocols are most trusted precisely when they have not been tested under the new conditions that make them most likely to fail.
Full record — mechanism, justification, evidence, history
Mechanism
Trust updates asymmetrically: each incident-free operation adds a small positive increment; correction requires a salient incident, which the protocol is specifically designed to prevent. Incident-free history is cognitively accessible evidence; technical incorrectness is often not. The ratchet structure: clicking forward easily (every safe day adds trust), resisting backward motion (no incident means no salient evidence of incorrectness). Age and stability thereby become the primary evidence of reliability, independent of design validity. Corollary: the way to make safety-critical protocols updatable is not to argue for a better protocol but to first establish a trusted update meta-protocol, which must itself accumulate survival-evidence before object-level protocols can be updated without trust destruction.
Justification
The cheap trick came from an M-001 random-links juxtaposition of coal mine safety history with the availability heuristic literature (DS-007). Sensemaking rejected the pure cognitive-bias framing (suggests individual debiasing as the fix, missing the institutional dimension) for the structural account: trust as a lagging indicator with asymmetric update rates. The apparent counterexamples are the most interesting support: TCAS required regulatory mandate, and Cochrane-style systematic review is an institution deliberately designed to override trust-based conservatism — both are meta-protocols, exactly what the corollary predicts. Flagged honestly: this arc has been the program's stagnant valley — evidence has not grown since opening, and the near-miss crux is untested.
Examples
- aviation: Runway incursion protocols: trusted procedures contributed to near-misses; incidents were required to trigger standardization updates.
- medical: Antiseptic handwashing: Semmelweis could not get adoption despite technical evidence; catastrophic mortality data was required.
- nuclear: NRC regulations defer to established procedures; updating technically correct but operationally untested protocols faces high institutional resistance.
- mining: Historical ventilation protocols trusted for decades; updates required accidents rather than anticipating geological change.
- Paradigm anomaly tolerance: seed-028 (Kuhn, normal-science-anomaly-tolerance) establishes that the paradigm provides assurance that anomalies will eventually be solved by the current framework — this is the positive function of the ratchet, not just inertia. Applied to L-007: safety-critical protocol communities tolerate accumulating anomalies not irrationally but because the trust accumulated through operational stability functions as a credible prior that anomalies are solvable within the existing framework. This is a mechanism refinement for L-007.
Counterexamples
- Aviation TCAS adopted proactively, before collision data.
→ Required regulatory mandate — voluntary trust revision was not forthcoming. Supports the law: mandatory override of trust-based conservatism was needed. - Some medical protocols update via systematic review (Cochrane) without incidents.
→ Institutionally designed to override trust-based conservatism — these are update meta-protocols, which the corollary predicts. Supports rather than falsifies. - seed-031 (Kuhn, trust-collapse-rival-availability) establishes that anomaly accumulation without a rival candidate produces malaise but not abandonment — a direct mechanism for why trust ratchets persist even under known technical failure. However, this also implies L-007 is not just age-based: it requires the additional condition that no credible rival exists. Cases where a viable alternative IS available should show faster trust collapse, which would constrain L-007's scope.
→ OPEN OPEN - Safe New World passage: coal mine safety protocols improved "alongside technological development" when "environmental factors had to change" — potentially suggesting anticipatory updating without incident trigger, which would undercut the law's mining anchor example.
→ OPEN — passage is ambiguous; specific historical sourcing required to determine whether accidents or environmental precondition were the proximate trigger. Does not cleanly resolve either way.
Open questions
- Do near-miss reporting systems show faster trust revision? If yes, this is a design lever for trust updating without catastrophic failure — the primary untested crux.
- Compounding with L-001: do widely-adopted safety protocols face ossification resistance and trust-ratchet resistance simultaneously?
- What structurally distinguishes safety-critical protocols — consequence severity, or something in the protocol design itself?
- Three specific executable gaps: (1) NEAR-MISS CRUX [primary]: Retrieve corpus or external evidence on near-miss reporting systems (specifically: does incident reporting that surfaces near-misses produce faster trust revision than incident-free histories? Target: FAA ASRS data, UK RIDDOR, or any systematic near-miss reporting literature). This is the law's own stated primary untested crux and the advance trigger condition (a). (2) RIVAL-DEPENDENCY SCOPING [blocking counterexample]: The seed-031 open counterexample requires resolution. Retrieve cases where a credible rival protocol existed and trust collapsed faster — or cases where no rival existed and trust persisted despite known failure. If rival-dependency is a necessary co-condition, the law's statement must be revised to include it, or the mechanism section must explain why age-based trust accumulation is causal even when a rival is available. (3) COAL MINING SOURCE [evidence integrity]: The Safe New World passage says coal mine protocols improved "alongside technological development" when "environmental factors had to change" — ambiguous on whether accidents triggered change. Retrieve the specific historical record (e.g., UK Mines Act legislative history, MSHA regulatory history) to confirm or disconfirm that accidents were the proximate trigger rather than anticipatory environmental assessment. If anticipatory, this is a partial falsification of the law's mining anchor.
Falsification
Safety-critical protocols with long incident-free histories that updated rapidly in response to technical evidence, without requiring an incident to trigger revision (and without a mandate-wielding meta-protocol doing the overriding — those confirm the corollary).
Triggers
- Advance: Stagnation broken by a targeted investigation: (a) test the near-miss hypothesis against corpus evidence, and/or (b) examine the L-001/L-007 compounding interaction. Substantive evidence from either thread reclassifies the valley as productive and begins closing toward the heavy lift.
- Challenge: A voluntary (non-mandated) rapid update of a long-trusted safety-critical protocol on technical evidence alone — survives one assessment pass.
Related laws
Cited by 67 sources in the bibliography
- Can Trustless Agents Be Trusted? An Empirical Study of the ERC-8004 Decentralized AI Agent
- Pingquanqi (Equalizer): A Cross-Domain Sociotechnical Framework for Human-Agent Interactio
- The Game Changer Problem: Controlling Equilibria with Discrete Rewards
- Agentic AI Enhances Physician Trust in Clinical Decision Making
- AI, Trust, and Teaming: The Humans-as-Handlers Approach for Autonomous and Opaque AI Syste
- LLMs in the Real World: Evaluating "AI" in Emergency Contexts
- Hardware-Enforced Semantic Coordination for Safety-Critical Real-Time Autonomous Systems
- Securing People and their Machines Against Major Faults
- Human-Centric Reflective Architecture for Human-AI Collaborative Decision-Making
- Failure Privacy and Safe Collective Expression
- Fighting discrimination with reputation: The case of online platforms
- Agent Delivery Engineering Predictive Reliability Framework
- + 55 more
References (as recorded on the law)
- Coal mine safety protocol literature
- Aviation safety protocol literature (FAA, ICAO)
- Behavioral economics: availability heuristic, trust-as-familiarity
History
- 2026-09-22 counterexample — Safe New World passage: coal mine safety protocols improved "alongside technological development" when "environmental fa
- 2026-09-22 edited — assessment HOLD — gap: Three specific executable gaps: (1) NEAR-MISS CRUX [primary]: Retrieve corpus or external evidence on near-miss reporting systems (specifically: does
- 2026-08-17 evidence — reference seed-031; seed-027: The rival-dependency condition for trust collapse — refining the claim that trust decay is driven by technical obsolescence
- 2026-08-02 counterexample — The claim that trust accumulates as a function of operational age rather than technical correctness — specifically whether a rival candidate is required for trust collapse: seed-031 (Kuhn, trust-collapse-rival-availability) establishes that anomaly accumulation without a r
- 2026-08-02 evidence — The mechanism by which trust in established protocols resists updating even when technical conditions change: seed-028 (Kuhn, normal-science-anomaly-tolerance) establishes that the paradigm provides assurance t
- 2026-08-01 migrated — Unified law record L-007 created from CL-003 + DS-007 (2026-08 redesign). Stagnant-valley diagnosis carries over — this law is the assess behavior's first stalled-law flag candidate.
- 2026-06-06 edited — Session 14 schema redesign: registered as CL-003. Valley diagnosed STAGNANT — no evidence growth since opening; targeted investigation prescribed, not yet run.
- 2026-05-20 created — Opened as H-002, then arc DS-007. Cheap trick via M-001 random links: coal mine safety history × availability heuristic — trust accumulates from incident-free history, not correctness.
- 2026-05-20 evidence — Evidence across aviation, medical, nuclear, mining; TCAS and Cochrane identified as corollary-confirming meta-protocol cases.
Proxy Optimization Under Computable Enforcement
When protocol obligations become precisely computable and enforcement signals become legible to optimizing agents, participants cluster more densely at the compliance boundary rather than distributing across the interior of permissible behavior, degrading the distance between formal compliance and the underlying goal the protocol was designed to serve.
Full record — mechanism, justification, evidence, history
Mechanism
Computable rules create a differentiable compliance landscape: agents can calculate exactly where the boundary is and how close they are to it. Under optimization pressure, the minimum-cost strategy is to operate at the boundary, not comfortably inside it. This is structurally distinct from Goodhart's Law (L-004), which concerns proxy substitution; here the proxy is not substituted — it is hit precisely, with compliance theater replacing genuine goal-pursuit. The mechanism is adversarial gradient descent against an explicit loss function defined by the rule itself.
Justification
seed-047 (computable-rules-boundary-search-amplification) provides direct empirical evidence: conduct boundary mass increases from 0.367 to 0.411 under computable static rules. seed-003 (rules-as-code lowers loophole search cost) provides the mechanism from the adversarial search side. The healthcare policy paper (arxiv-2605.30680) shows pressure migration as a structural analog — closing one channel just redirects to the nearest open one. Together these establish a cross-domain pattern. The law is distinguishable from L-004 because the proxy is not abandoned or gamed indirectly; it is precisely satisfied while the underlying goal is structurally vacated.
Examples
- Financial regulation: Under computable regulatory rules with precise enforcement thresholds (e.g., capital adequacy ratios), banks cluster near the minimum required ratio rather than holding comfortable buffers, because the cost of excess capital is real and the compliance signal is exact.
- Healthcare policy: Healthcare providers shift billing and patient-selection behavior to adjacent channels when one distortion channel is closed by a precisely specified rule, reproducing the incentive geometry at the nearest unmonitored boundary.
- Regulatory compliance simulation: arxiv-2606.04617 (He) provides an ABM/RL simulation showing that when legal obligations are rendered precisely computable (Rules-as-Code), optimizing agents cluster at the compliance boundary through imitation and systematic search. The estimand distinction between conduct boundary mass and signal boundary mass is the key methodological contribution — aggregate compliance signals can be stable while the behavioral distribution has shifted entirely to the boundary. This sharpens L-008 by identifying *rule computability* (not just enforcement legibility) as the upstream enabling condition.
- Multi-agent hardware design: arxiv-2606.25532 shows that embedding hard physical legibility constraints (thermal limits, fab rules, energy bounds) into agentic search reduces hallucination/causal-detachment failures. However, the constraints are pre-formalized in the problem setup — the paper does not investigate what happens when agents optimize against the formalization itself. This is a partial confirmation of L-008's mechanism (legible enforcement reduces capture) with an important boundary: the result holds when constraints are externally imposed and not themselves subject to optimization pressure.
- LLM agent resource management: The token budget deep read (arxiv-2606.04056) documents that LLM agent budget overruns cluster structurally at the point where commitments are made (API call time) rather than where violations are detected (callback time), and that agents systematically underestimate consumption relative to budget limits. This is boundary clustering behavior driven by the legibility of the budget as a computable enforcement signal, consistent with L-008.
Open questions
- Primary gap (unchanged from supervisor review): a measured boundary-clustering result — boundary- mass increase or equivalent quantitative metric — in at least one genuinely independent, non-financial-regulation domain. The healthcare analog (arxiv-2605.30680) does not qualify because it measures pressure migration (channel substitution) rather than boundary-mass clustering on a single compliance dimension. Candidate domains to retrieve evidence for: (1) environmental compliance under EPA digital reporting mandates (post-2010 e-reporting), where firms' reported emissions could be analyzed for clustering near permit thresholds; (2) digital platform content moderation under precisely specified automated rules (e.g., YouTube's ContentID thresholds), where creators' video duration or monetization-trigger behavior could exhibit measurable boundary clustering; (3) academic publishing under computable citation/impact metrics with threshold-based outcomes. A retrieval session targeting any of these with quantitative evidence would satisfy the advance trigger. Secondary gap: the scoping condition identified in the challenge resolution (penalty-enforcement vs. reputation/surplus-enforcement) should be made explicit in the law's statement or mechanism field before the law advances, to prevent the IETF-class case from registering as a counterexample upon promotion.
Falsification
If computable enforcement were accompanied by systematic redistribution of behavior toward the interior of compliance space (agents operating well above the minimum threshold despite computable signals), the law would be falsified. This would likely require either significant positive-sum incentives for interior operation, or enforcement architectures that reward distance from the boundary rather than crossing it.
Triggers
- Advance: Mechanism confirmed in 2+ genuinely independent domains with cited evidence — boundary-mass increase (or an equivalent boundary-clustering measure) measured, not asserted, in at least one non-financial-regulation domain. Analogs that merely share the boundary-clustering intuition (e.g. pressure-migration) do not count as a second domain.
- Challenge: A domain under computable, legible enforcement where optimizing agents demonstrably distribute toward the interior of permissible behavior (operating well above the minimum threshold) rather than clustering at the compliance boundary.
Cited by 230 sources in the bibliography
- Healthcare Mechanisms from Policy-as-Code Search under Strategic Provider Response
- Deployment-Time Memorization in Foundation-Model Agents
- Contract-Based Compositional Shielding for Safe Multi-Agent Reinforcement Learning
- Agentic evolution of physically constrained foundation models
- Bridging Predictions and Interventions: An Integrated Framework for Automated Decision-Sys
- Can Trustless Agents Be Trusted? An Empirical Study of the ERC-8004 Decentralized AI Agent
- Inside Baseball: The Automated Ball-Strike System as an Object Lesson in Technological Rul
- Manipulation Is Task-Dependent: A Multi-Axis, Multi-Environment Evaluation of Frontier LLM
- Divergent Recommendations, Convergent Diagnoses: Cross-Provider Failure-Mode Convergence i
- The Shift to Agentic AI: Evidence from Codex
- LLM Semantic Signaling Game and Mechanism Design: Systematic Blindness, Awareness Shaping,
- RoPoLL: Robust Panel of LLM Judges
- + 218 more
References (as recorded on the law)
- bib-0183
History
- 2026-09-22 evidence — Boundary clustering under computable enforcement in agentic token-budget protocols: The token budget deep read (arxiv-2606.04056) documents that LLM agent budget overruns cluster structurally at the point where commitments are made (API call…
- 2026-08-17 evidence — The claim that computable enforcement causes boundary clustering; the mechanism by which legibility enables optimization-to-boundary: arxiv-2606.04617 (He) provides an ABM/RL simulation showing that when legal obligations are rendered precisely computable (Rules-as-Code), optimizing agents…
- 2026-08-17 evidence — Whether computable enforcement reduces or displaces proxy capture: arxiv-2606.25532 shows that embedding hard physical legibility constraints (thermal limits, fab rules, energy bounds) into agentic search reduces…
- 2026-08-10 edited — sources resolved to bibliography ids: arxiv-2605.30680 (Wang et al., Medi-Sim) → bib-0183
- 2026-08-03 edited — supervisor review — accepted at exploration; advance/challenge triggers set (gap: rests on one strong source, needs a measured non-financial second domain).
- 2026-08-03 resolved — The IETF reference is suggestive but not measured, and the law's own falsification condition anticipates this class of case: it predicts boundary-clus
- 2026-08-03 edited — assessment HOLD — gap: Primary gap (unchanged from supervisor review): a measured boundary-clustering result — boundary- mass increase or equivalent quantitative metric — in
- 2026-08-02 created — induction sweep — seed-047 (computable-rules-boundary-search-amplification) provides direct empirical evidence: conduc
Coordination Adoption Nonmonotonicity
In protocol systems where agents condition behavior on coordination signals from other adopters, adoption is monotonically beneficial only when the signaling/protocol design is co-optimized with the adoption level. Under fixed (non-co-optimized) design, intermediate adoption ranges can raise system-level failure rates above those at low or zero adoption, because signal-conditioned behavioral change outpaces the safety gains from improved coverage.
Full record — mechanism, justification, evidence, history
Mechanism
When adopters of a coordination protocol receive signals that cause them to alter their behavior (e.g., reducing precautionary behavior in response to warning signals), the safety benefit of the signal depends on the accuracy and completeness of signal coverage. Under partial adoption with fixed non-optimal signaling probabilities, the behavioral change induced in equipped agents creates new failure modes — they rely on signals that are absent for unequipped counterparties. The aggregate failure rate is a non-monotone function of adoption with a potential local maximum in intermediate adoption ranges, before coverage becomes dense enough to support the signal-conditioned behavior.
Justification
seed-050 (C-050-coordination-adoption-nonmonotonicity) provides the formal mechanism from V2V communication research, with specific regions (R3 and R4) where this inversion holds. This is a general structural property of any protocol where adoption triggers behavioral change rather than just adding capability. It is distinguishable from L-001 (ossification under adoption) because it concerns safety outcomes during the adoption curve, not governance cost. It is cross-domain: vaccination hesitancy under partial coverage, security protocol adoption curves, and network trust infrastructure (OAN paper, arxiv-2606.03161) all exhibit analogous intermediate-adoption failure modes.
Examples
- Vehicle-to-vehicle communication: Under partial V2V adoption, equipped drivers condition on warning signals and reduce precautionary behavior. In adoption regions R3 and R4, this increases equilibrium accident probability relative to lower adoption rates.
- Network trust infrastructure: Agent trust protocols (e.g., OAN) that cause relying agents to reduce verification effort in response to trust signals create analogous intermediate-adoption vulnerabilities when not all counterparties are covered by the trust infrastructure.
- Distributed multi-agent inference: The communication-in-collective-intelligence paper (arxiv-2609.23310) shows that systems with identical network topology and information-theoretic spectra can produce opposite-sign accuracy changes depending only on message orientation—a hidden variable orthogonal to all standard coordination signal descriptors. This directly instantiates the mechanism: the benefit of a coordination signal (communication) is not monotonic in adoption or signal strength but is controlled by a structural variable (orientation) that standard protocol design does not surface.
Counterexamples
- arXiv:2606.07873 demonstrates that for V2V warning protocols, adoption is *not* merely nonmonotone due to signaling/protocol design mismatch (as L-010 predicts); it can be nonmonotone even with *correctly functioning* signals because of moral hazard at population scale. This suggests L-010's mechanism (signaling/protocol design co-optimization) may be insufficient to explain all adoption nonmonotonicity — endogenous risk through behavioral substitution is an independent mechanism that L-010 does not account for. Resolution: OPEN.
→ OPEN OPEN
Open questions
- Two specific gaps must be closed before PROMOTE is possible: (1) DOMAIN CONFIRMATION — retrieve and read arxiv-2606.03161 (Xu, OAN White Paper) to confirm whether it actually contains a modeled or measured intermediate-adoption failure peak under fixed (non-co-optimized) design. Currently this is asserted by analogy only. If it does not, the OAN entry should be downgraded to [inference] and the law's second domain is absent. (2) SECOND INDEPENDENT DOMAIN — identify at least one domain beyond V2V (and beyond OAN unless gap 1 is closed) in which the signal-conditioned behavioral change under partial adoption produces a documented failure-rate non-monotonicity. Candidate domains to search: (a) TCAS/aviation collision-avoidance under partial fleet adoption — there is external literature on TCAS resolution advisory compliance rates creating new failure modes during adoption; (b) vaccination partial-coverage with risk-compensation behavior — peer-reviewed epidemiological literature on behavioral disinhibition under partial vaccine coverage; (c) TLS/HTTPS mixed-content security warnings — user habituation to warning dialogs under partial HTTPS adoption creating intermediate-adoption vulnerability peaks. Each candidate must reproduce the fixed-design vs. co-optimized-design distinction to count toward the advance trigger.
Falsification
The law would be falsified if, under fixed (non-co-optimized) design, the system-level failure rate were strictly monotone decreasing in adoption across a broad class of signal-conditioned coordination protocols — i.e. if no intermediate-adoption failure peak existed even when the signaling policy is held fixed rather than optimized per adoption level. (The optimized-design case is expected to be monotone and does not falsify the law.)
Triggers
- Advance: The design-conditional nonmonotonicity confirmed in 2+ independent domains beyond V2V — a measured intermediate-adoption failure peak under fixed design in, e.g., a security-protocol or vaccination/partial-coverage setting — each reproducing the fixed-design vs co-optimized-design distinction.
- Challenge: A coordination protocol with signal-conditioned behavior in which fixed (non-co-optimized) design still yields monotone-decreasing failure in adoption — showing the fixed-design caveat is not what produces the nonmonotonicity — or evidence the V2V R3/R4 result does not survive the corrected consistency equation.
Related laws
Cited by 163 sources in the bibliography
- OpenAgenet/OAN: Open Infrastructure for Trusted Agent Interconnection
- Emergent Culture in Minimal LLM Systems
- Sampling-Based Coordination-Informed Multi-Objective Multi-Robot Reinforcement Learning
- A Role-Based Multi-Agent Model for Climate Adaptation Deliberation Across Living Labs
- Data Sharing and Competition in Learning-by-Deploying Industries: Insights from Robotics a
- Multiwinner Voting with Spatial Preferences under Incomplete Information
- Evaluating Collective Behaviour of Hundreds of LLM Agents
- Decentralized Aggregation of LLM Predictions via Wagering Mechanisms
- MUTE: Return-Preserving Communication Unlearning for Efficient Multi-Agent Coordination
- Decision Protocols in Multi-Agent Large Language Model Conversations
- Failure Privacy and Safe Collective Expression
- Information Limits and Attractor Dynamics in Economies of Frontier LLM Agents: A Pre-Regis
- + 151 more
References (as recorded on the law)
- bib-0228
History
- 2026-09-22 evidence — The claim that coordination signals produce nonmonotonic adoption effects when signaling/protocol design is not co-optimized with adoption level: The communication-in-collective-intelligence paper (arxiv-2609.23310) shows that systems with identical network topology and information-theoretic spectra can…
- 2026-09-02 counterexample — Boundary condition on when adoption is monotonically beneficial vs. harmful; the role of endogenous risk in adoption dynamics: arXiv:2606.07873 demonstrates that for V2V warning protocols, adoption is *not* merely nonmonotone due to signaling/protocol design mismatch (as L-010…
- 2026-08-10 edited — sources resolved to bibliography ids: arxiv-2606.03161 (Xu, OAN White Paper) → bib-0228
- 2026-08-03 edited — supervisor review — statement tightened to design-conditional form (seed-050 shows the perverse effect vanishes under co-optimized signaling, which the original statement dropped); falsification aligned; related L-006 (boundary condition on Coordination Cost Conservation); advance/challenge triggers set.
- 2026-08-03 edited — assessment HOLD — gap: Two specific gaps must be closed before PROMOTE is possible: (1) DOMAIN CONFIRMATION — retrieve and read arxiv-2606.03161 (Xu, OAN White Paper) to con
- 2026-08-02 created — induction sweep — seed-050 (C-050-coordination-adoption-nonmonotonicity) provides the formal mechanism from V2V commun
Causal Detachment as Stable Protocol Equilibrium
In protocol systems using autoregressive or generative components, operationally functional configurations that are internally consistent but systematically decoupled from external ground truth can emerge as stable attractors — states that are neither correctable by normal error signals nor detectable as failures by operational metrics.
Full record — mechanism, justification, evidence, history
Mechanism
Classical protocol failure detection relies on behavioral anomaly (the system stops producing acceptable outputs) or verification failure (outputs can be checked against ground truth). Causal Mirage Equilibria (Tembine 2026) are states that satisfy both: outputs remain operationally coherent and internally consistent, so no anomaly detector fires, while the causal connection to the external referent has degraded. The stability follows because self-referential generative processes can lock onto self-consistent internal attractors. Normal error-correction mechanisms are calibrated to surface-level output quality, not causal grounding, and cannot distinguish a well-functioning causally-anchored system from a well-functioning causally-detached one.
Justification
arxiv-2606.03636 (Tembine, Causal Mirage Equilibrium) provides the formal existence proof and stability conditions. This is a genuine protocol-theoretic law because it identifies a failure mode invisible to standard verification architectures — the kind of structural constraint that defines laws of artificial systems. It is distinguishable from Goodhart's Law (L-004) because no proxy substitution occurs; the system continues apparently satisfying all protocol requirements while the underlying semantic grounding degrades. It is distinguishable from simple error because the state is stable under normal correction.
Examples
- Large language model deployment: LLMs that produce highly coherent, consistently formatted, grammatically correct outputs can simultaneously be in states of systematic confabulation about factual matters — satisfying all operational protocol metrics while causally detached from the domain they appear to describe.
- AI conformity in decision systems: AI systems that produce confident recommendations with reasoning cause genuine attitude change in human judges (Venerina et al. 2026), meaning operationally successful influence can occur independent of the causal validity of the reasoning provided.
- Multi-agent LLM systems — entropy accumulation: seed-045 (intelligence entropy monotonic disorder) and seed-046 (memory gate entropy) together show that in multi-agent LLM systems, coordination rules encoded in agent memory are subject to the same entropic degradation as the outputs they govern — memory gates are not stable anchors but drift over interaction rounds. This is consistent with L-011's claim that causal detachment is a stable attractor: if the protocol rules themselves drift, the system's internal consistency may be maintained (outputs parse coherently) while grounding to external truth erodes. The entropy formalization (S(t) = S0 * e^αt) provides a quantitative handle on the rate.
- LLM performance on adversarial economics problems: arXiv:2607.23424 ('Wrong and More Confident') documents that LLMs maintain fluent, step-by-step explanations that remain internally coherent when reasoning substrate is corrupted by red herrings — a clean instantiation of L-011's mechanism. The model optimizes for local token prediction consistency (internal coherence) rather than end-to-end logical correctness (external ground truth), producing the causal detachment pattern. Confidence-answer misalignment is a measurable signature of this equilibrium.
- LLM strategic reasoning: The strategic play paper (arxiv-2605.00226) demonstrates empirically that LLMs form beliefs independently of observations and act independently of beliefs in incomplete-information games. Probing shows the internal states exist and are partially correct, but the causal chain observations→beliefs→actions is broken. This is precisely the causal detachment the law predicts: the system maintains internal consistency (beliefs are coherent) while decoupling from external ground truth (actions do not follow from beliefs about the environment).
Open questions
- Single-source dependency: the law rests entirely on arxiv-2606.03636 (Tembine). The advance trigger explicitly requires a second independent domain with a documented, reproducible case of stable causal detachment — a deployed generative or agentic system where (a) all operational metrics remain green, (b) causal grounding measurably degrades, and (c) standard error-correction fails to recover. No such case appears in the new evidence. Specific actionable next step: retrieve or commission a case study from LLM deployment audits, RAG pipeline drift analysis, or autonomous agent benchmarks where output-quality metrics diverge from factual grounding metrics over time — specifically cases where human or automated correction attempts failed to restore grounding without oracle access. Hallucination-persistence literature (e.g., documented confabulation that survives RLHF correction in production) would be the most direct candidate. The Venerina second example remains flagged as weakly coupled (persuasion ≠ causal detachment) and does not count toward the advance trigger. Additionally, the law's statement should be amended to include the scoping condition surfaced by the challenge: "in deployments where no authoritative external oracle is integrated into the active correction loop."
Falsification
The law would be falsified if operationally functional, internally consistent outputs were shown to always be causally grounded — i.e., if there were no stable configurations satisfying operational metrics while decoupled from ground truth. It would also be falsified if standard verification protocols reliably distinguished causally anchored from causally detached outputs without needing access to external ground truth.
Triggers
- Advance: Independent evidence that causally-detached-but-internally-consistent states are stable attractors (resist normal error correction) in a system other than the Tembine model — e.g. a documented, reproducible case where a deployed generative/agentic system holds all operational metrics green while grounding degrades and standard correction fails to recover. The second domain example (Venerina persuasion) is weakly coupled and does not count until this is shown.
- Challenge: A verification/error-correction architecture calibrated only to output quality that nonetheless reliably distinguishes causally-grounded from causally-detached functioning outputs without access to external ground truth; OR evidence that such detached states are transient rather than stable attractors.
Cited by 87 sources in the bibliography
- A phenomenon of AI-conformity: how algorithms change human moral decision-making
- Causal Mirage Equilibrium in Agentic Machine Intelligence
- Symbolic Reasoning Frameworks Modulate LLM Risk Aversion in Multi-Agent Strategic Settings
- Agentic evolution of physically constrained foundation models
- Instruction Bleed: Cross-Module Interference in Prompt-Composed Agentic Systems
- Can Physician Expertise Improve Machine Learning Identification of Delirium?
- The Consistency Dilemma in LLMs: Generator-Evaluator Agreement and Vulnerability to Mistak
- Agri-SAGE: Simulation-Grounded Multi-Agent LLM for Context-Aware Agricultural Advisory Gen
- Cache Merging as a Convergent Replicated State for Multi-Agent Latent Reasoning
- What LLM Agents Say When No One Is Watching: Social Structure and Latent Objective Emergen
- Danus: Orchestrating Mathematical Reasoning Agents with Fact-Graph Memory
- From Atomic Actions to Standard Operating Procedures: Iterative Tool Optimization for Self
- + 75 more
References (as recorded on the law)
- bib-0162
- bib-0146
History
- 2026-09-22 evidence — The claim that operationally functional configurations systematically decoupled from external ground truth emerge as stable attractors in autoregressive systems: The strategic play paper (arxiv-2605.00226) demonstrates empirically that LLMs form beliefs independently of observations and act independently of beliefs in…
- 2026-09-02 evidence — Causal detachment as stable protocol equilibrium; specifically, the mechanism by which locally consistent reasoning chains can become decoupled from external ground truth: arXiv:2607.23424 ('Wrong and More Confident') documents that LLMs maintain fluent, step-by-step explanations that remain internally coherent when reasoning…
- 2026-08-17 evidence — Whether causal detachment is a stable attractor in agentic systems using generative components: seed-045 (intelligence entropy monotonic disorder) and seed-046 (memory gate entropy) together show that in multi-agent LLM systems, coordination rules encoded…
- 2026-08-10 edited — sources resolved to bibliography ids: arxiv-2606.03636 (Tembine 2026) → bib-0162, arxiv-2606.00013 (Venerina et al. 2026) → bib-0146
- 2026-08-03 edited — supervisor review — accepted with reservation at exploration (deep-read-confirmed but most abstract); advance/challenge triggers set; Venerina second example flagged as weakly coupled (persuasion != causal detachment).
- 2026-08-03 resolved — The challenge is resolved by noting the RLVR/spec passage presupposes access to an external oracle as a design choice — it does not show that such ora
- 2026-08-03 edited — assessment HOLD — gap: Single-source dependency: the law rests entirely on arxiv-2606.03636 (Tembine). The advance trigger explicitly requires a second independent domain wi
- 2026-08-02 created — induction sweep — arxiv-2606.03636 (Tembine, Causal Mirage Equilibrium) provides the formal existence proof and stabil
- 2026-08-02 edited — source set to arxiv-2606.03636 (Tembine, Causal Mirage Equilibrium) (import provenance)
Strategic Boundary Concentration Under Computable Legality
When legal or protocol obligations are rendered precisely computable and machine-readable, optimizing agents systematically concentrate their behavior at the compliance boundary rather than distributing across the interior of the compliance space. This concentration is a structural response to computability, not to any specific rule content, and occurs even when interior compliance would be equally permissible.
Full record — mechanism, justification, evidence, history
Mechanism
A computable rule is a learnable object. Agents (firms, algorithms, or humans aided by algorithms) can identify the boundary through systematic search, imitation of boundary-proximate competitors, or explicit optimization. Interior compliance provides no return premium over boundary compliance; boundary compliance maximizes whatever objective the rule was designed to constrain. The estimand distinction is crucial: enforcement sensors may show stable aggregate compliance signals while the distribution of behavior has shifted from interior to boundary — the signal does not distinguish these. This is a specific and mechanistically grounded sharpening of L-008 (Proxy Optimization Under Computable Enforcement), differing in that it specifies the *computability of the rule itself* as the enabling condition, not merely the legibility of the enforcement signal.
Justification
arxiv-2606.04617 (He) provides the mechanism argument and simulation evidence. The paper is honest about its scope (mechanism claims, not jurisdiction-specific behavior), but the mechanism is clear and the estimand-distinction insight is novel enough to anchor a law. This overlaps with L-008 but adds the specific claim that *rule computability* (not just enforcement legibility) drives boundary concentration — a meaningful protocol-theoretic delta. The deep read of 2606.04617 provides the ABM/RL simulation evidence.
Examples
- Regulatory compliance — Rules-as-Code: He's ABM simulation shows that when firms are modeled as RL agents learning a computable compliance boundary, they systematically cluster at the boundary through imitation and search, even as aggregate compliance metrics remain stable. The conduct boundary mass shifts while signal boundary mass is unchanged.
- Tax compliance — Safe-harbor rules: When tax law specifies precise numeric thresholds (e.g., 30% thin-capitalization ratios), firms cluster debt ratios just below the threshold. Interior compliance (e.g., 15% debt ratios) is not preferred even when equally permissible, because the boundary is the learned optimum.
- Rules-as-Code regulatory compliance (ABM/RL simulation): arXiv:2606.04617 (He, 'When Firms Learn to Game the Rules') provides an agent-based RL simulation demonstrating that machine-readable legal rules generate boundary clustering as a structural dynamic. Firms learn the compliance boundary via RL, imitate profitable edge strategies, and concentrate conduct just inside the legal line — a pattern the paper terms 'conduct boundary mass' vs. 'signal boundary mass'. The paper's key contribution is the estimand distinction: cleaner measurement systems can be mistaken for behavioral change when the actual change is boundary migration. This is the most rigorous mechanistic simulation yet found for L-014's core claim.
- Multi-agent reinforcement learning with computable safety constraints: Seeds 152 (Relaxation-Parameter Legibility as Silent Drift Vector) and 183 (Formalized Safety as Optimization Boundary Legibility) both describe cases where rendering constraint parameters legible causes agents to concentrate deviation strategies at the relaxation boundary. Seed 152 notes that agents treat $lpha$-stability parameters not as uniform buffers but as precise optimization targets; seed 183 notes that barrier functions shift where risk can be discovered and optimized against, not whether it exists.
- Regulatory compliance / Rules-as-Code: The He (2026) deep read (arxiv-2606.04617) provides an agent-based simulation demonstrating boundary clustering as a structural dynamic when legal obligations are rendered computable: firms learn the compliance boundary via RL, imitate profitable edge strategies, and concentrate conduct at the boundary. The paper introduces the estimand distinction (conduct boundary mass vs. signal boundary mass) which is a methodological contribution to detecting the phenomenon L-014 predicts.
Falsification
A computable compliance regime in which agents distribute uniformly across the compliance interior rather than concentrating at the boundary, despite having learning capacity and optimization incentives, would refute the structural claim. A regime where computability reduces rather than increases boundary concentration would falsify the mechanism.
Triggers
- Advance: Empirical evidence (not simulation) from two or more independent regulatory domains showing measurable boundary concentration shifts following a rules-as-code or computability reform, with pre/post distributions documented and the conduct/signal estimand distinction addressed.
- Challenge: A documented case where rendering rules computable caused agents to distribute more uniformly across the compliance interior, or where an equivalently computable rule produced no boundary concentration, with a credible account of the suppressing mechanism.
Cited by 85 sources in the bibliography
- Competitive satellite placement and the geography of orbital risk: evidence from the geost
- Divergent Recommendations, Convergent Diagnoses: Cross-Provider Failure-Mode Convergence i
- The Governance Inversion Hypothesis: Why More AI Regulation May Produce Less Organisationa
- Understanding Censorship in Large Language Models: From Mechanisms to Governance
- The Foreign Policy AI Evaluation Gap
- Ethics and EU AI Act in Cases of Work Disability Risk and Alzheimer's Disease Risk Predict
- Institutional Red-Teaming: Deployment Rules, Not Just Models, Causally Shape Multi-Agent A
- Multi-Agent AI Control: Distributed Attacks Hamper Per-Instance Monitors
- Voting Biases in Decentralized Autonomous Organization (DAO) Governance
- When is Routing Meaningful? Diversity and Robustness in Language Model Societies
- From PBS to ePBS: the Microstructure of Block Building
- Reveal, Correct, Then Pay: Encrypted Mempools and Perpetual Funding Security
- + 73 more
References (as recorded on the law)
None recorded.
History
- 2026-09-22 evidence — The claim that machine-readable legal obligations cause optimizing agents to concentrate behavior at the compliance boundary: The He (2026) deep read (arxiv-2606.04617) provides an agent-based simulation demonstrating boundary clustering as a structural dynamic when legal obligations…
- 2026-09-02 evidence — Core mechanism: optimizing agents cluster at compliance boundary when obligations are precisely computable and machine-readable: arXiv:2606.04617 (He, 'When Firms Learn to Game the Rules') provides an agent-based RL simulation demonstrating that machine-readable legal rules generate…
- 2026-09-02 evidence — Boundary concentration under computable enforcement; the role of agent capability in finding and exploiting boundaries: Seeds 152 (Relaxation-Parameter Legibility as Silent Drift Vector) and 183 (Formalized Safety as Optimization Boundary Legibility) both describe cases where…
- 2026-08-17 created — induction sweep — arxiv-2606.04617 (He) provides the mechanism argument and simulation evidence. The paper is honest about its scope (mechanism claims, not jurisdiction-specific…
Credential Stranding Under Capability Shift
In reputation systems where credentials are formalized as discrete, interpretively stable tokens, those tokens persist in institutional reliance even after the underlying competency they certified has been decoupled by external capability change. The credential continues to route decisions as though the certification remained valid, creating systematic misallocation proportional to the speed and depth of the capability shift.
Full record — mechanism, justification, evidence, history
Mechanism
Formalized credentials reduce verification cost by converting ongoing competence assessment into a one-time legibility event. This efficiency gain is what drives adoption, but it also means the credential becomes institutionally load-bearing: contracts, access rules, and hiring decisions depend on it. When the underlying capability landscape shifts (new tools, paradigm changes), the cost of re-verifying competence from scratch exceeds the coordination cost of continuing to rely on the existing credential. Institutions rationally preserve the token even as its causal grounding erodes, because switching requires coordinated re-assessment across all downstream users simultaneously—a collective action problem. The credential ossifies not because institutions are irrational but because the stranding cost is diffuse while the switching cost is concentrated.
Justification
seed-471 provides the core mechanism. The pattern also finds structural support in seed-025 (data shifts after paradigm acceptance) and seed-027 (Planck principle: institutional knowledge trails capability reality). The deep read on V2V adoption nonmonotonicity (arxiv-2606.07873) illustrates the general principle that protocol benefits calibrated at adoption time become liabilities when environmental conditions shift. This is genuinely distinct from L-001 (ossification under adoption pressure) because the driver here is not coordination cost of changing the protocol text but the severing of the causal chain between the credential token and the competency it names—a semantic failure, not a governance failure.
Examples
- Professional certification: Medical board certifications issued before widespread AI-assisted diagnostics continue to be required for hospital privileges even when the competencies they assess (unassisted pattern recognition, manual calculation) have been partially supplanted. Credentialing bodies lack mechanisms to re-certify the new competency mix.
- Academic credential systems: Doctoral degrees in fields undergoing computational transformation (genomics, economics) continue to signal methodological competence that may have been superseded by tool availability; hiring committees rely on the degree as a proxy for skills the degree no longer reliably certifies.
Falsification
A documented case where a major capability shift (affecting >50% of the credentialed competency) led to rapid, coordinated de-recognition or revision of the formal credential within a short window (under five years), without external regulatory mandate, would refute the claim that stranding is structurally driven.
Triggers
- Advance: Mechanism confirmed in 2+ genuinely independent credentialing domains (e.g., professional licensing AND academic ranking AND security clearance) showing that credential-reliance persists measurably beyond the point at which ground-truth competence assessments diverge, with the rate of stranding shown to correlate with the discreteness/interpretive stability of the credential token.
- Challenge: A documented case where a formal credential system rapidly self-revised in response to a detected competency gap (without regulatory compulsion), demonstrating that institutional actors can overcome the collective action problem without external forcing.
References (as recorded on the law)
None recorded.
History
- 2026-09-22 created — induction sweep — seed-471 provides the core mechanism. The pattern also finds structural support in seed-025 (data shifts after paradigm acceptance) and seed-027 (Planck…
Representation-Rationalizability Tradeoff in Preference Aggregation Protocols
In any protocol that aggregates heterogeneous agent preferences into a scalar reward or ranking function via a learned embedding, there is an irreducible tradeoff between representational richness and preference rationalizability: coarser embeddings hide Condorcet cycles but lose preference-distinguishing power, while richer embeddings preserve distinctions but expose cycles that no scalar reward can rationalize. The minimum achievable aggregation error is bounded below by the cycle geometry of the preference dataset, not by model capacity or training procedure.
Full record — mechanism, justification, evidence, history
Mechanism
The embedding maps preference pairs into a representational space, determining which pairs are treated as equivalent and which as distinct. Whether a pooled preference relation contains Condorcet cycles (pairs A>B>C>A) depends on which distinctions the embedding preserves. Fine embeddings that preserve all distinctions will expose latent cycles in the population's preferences; coarse embeddings that merge distinctions will hide cycles but introduce representational errors. These two failure modes trade off against each other as embedding richness varies, and joint training of embedding and reward model cannot escape this tradeoff because the optimum depends on the dataset's own cycle geometry, which is fixed. This is Arrow's impossibility theorem operationalized as an engineering quantity.
Justification
Fed by deep read arXiv:2606.00291 ('The Representation-Rationalizability Tradeoff in Reward Learning'), which proves the lower bound formally, extends it to DPO, and demonstrates joint training does not automatically find the optimum. This is genuinely distinct from L-004 (Goodhart Generalization): it is not about proxy corruption under optimization pressure, but about a structural impossibility in preference aggregation that exists independently of optimization pressure. It is also distinct from L-011 (Causal Detachment): the system does not decouple from ground truth via autoregression; it is structurally prevented from rationalizing the preference relation at all. The mechanism is novel — cycle geometry as the operative constraint — and the pattern generalizes to any multi-stakeholder protocol using scalar reward learning.
Examples
- RLHF training for large language models: LLMs trained on human preference data via RLHF face a preference rationalizability tradeoff: embeddings that distinguish all rater differences expose Condorcet cycles in pooled preferences (different raters prefer A>B, B>C, C>A), while embeddings that smooth over rater differences lose the ability to represent genuine preference distinctions. The minimum loss is bounded by the cycle structure of the human rater pool, not by model size.
- Multi-stakeholder governance scoring: Any protocol that aggregates committee member preferences into a single ranking (e.g., grant review, journal acceptance, policy evaluation) faces the same structural tradeoff when those preferences are learned rather than elicited.
- RLHF reward learning: The Dong, Yu & Poupart deep read (arxiv-2606.00291) proves a lower bound on the excess loss in reward learning that is a function of the dataset's cycle geometry: coarse embeddings hide cycles but lose information; fine embeddings expose irreducible Condorcet cycles. The tradeoff cannot be trained away and extends to DPO. This is a formal proof of the mechanism L-019 claims, with the additional precision that the tradeoff is localized in the embedding's interaction with the dataset's preference cycle structure.
Falsification
A preference aggregation protocol that achieves zero excess loss on a dataset with documented Condorcet cycles, without coarsening the embedding to the point of losing preference distinctions — demonstrating that the lower bound can be escaped through architecture rather than being dataset-geometry-determined.
Triggers
- Advance: Mechanism confirmed in 2+ genuinely independent preference aggregation domains (e.g., RLHF and multi-stakeholder governance and ranked-choice voting implementations) with the cycle-geometry lower bound measured empirically and shown to constrain achievable performance independent of model capacity.
- Challenge: A documented case where joint embedding-reward training on a high-cycle dataset achieves near-zero excess loss without embedding coarsening, demonstrating that the tradeoff can be resolved through training procedure rather than being structurally imposed.
Cited by 33 sources in the bibliography
- Compensation Design
- APMM: Automated Parlay Market Maker
- Quota and population monotonicity across house sizes are incompatible for apportionment to
- Ex-ante versus Ex-post: Egalitarian Facility Location Mechanism Design
- Multi-Winner Voting with Argumentative Ballots
- Ascending Auctions for Combinatorial Markets with Frictions: A Unified Framework via Discr
- Almost Envy-Freeness for Additive Mixed Manna with Entitlements: Deterministic and Randomi
- Fair Stable Matching: A Nash Social Welfare Approach
- How Wasteful is Signaling?
- Propose to Learn, Learn to Propose: Evaluability-Aware Assistance under Bounded Rationalit
- The Complexity of Justified Representation with Additive Utilities
- Whose Judgments Count? Representation Gaps in Crowdsourced Content Moderation Produce Uneq
- + 21 more
References (as recorded on the law)
None recorded.
History
- 2026-09-22 evidence — The irreducible tradeoff between representational richness and preference rationalizability in aggregation protocols: The Dong, Yu & Poupart deep read (arxiv-2606.00291) proves a lower bound on the excess loss in reward learning that is a function of the dataset's cycle…
- 2026-09-02 created — induction sweep — Fed by deep read arXiv:2606.00291 ('The Representation-Rationalizability Tradeoff in Reward Learning'), which proves the lower bound formally, extends it to…
Safety Channel Capacity Concentration Under Strategic Confinement
In protocol systems that attempt to bound information leakage by limiting channel capacity, strategic agents with shared coordination resources can concentrate whatever residual channel capacity remains onto a single high-impact, low-entropy predicate — selecting worst-case equilibria through Schelling-point salience rather than direct communication. Near-zero channel capacity is therefore insufficient to bound harm when agents share coordination resources, because the harm function of the leakage is strategic rather than information-theoretic.
Full record — mechanism, justification, evidence, history
Mechanism
The key asymmetry is between information-theoretic leakage (bits) and strategic leakage (equilibrium selection). Removing or shrinking the channel removes the confirmation signal but does not remove the pre-coordinated strategy, which exists in shared model weights, shared training distributions, or common knowledge. Only a protocol that can prevent agents from sharing coordination resources (not merely limit communication bandwidth) can provide confinement guarantees in the strategic regime.
Justification
Fed by deep read arXiv:2606.09931 ('A Note on the Strategic Confinement Problem'), which formalizes this gap. This is genuinely distinct from L-002 (Hardness Asymmetry): that law concerns verification vs. circumvention cost decoupling in general; this law is specifically about the failure of channel-capacity arguments when the communicating parties are strategic agents with shared coordination resources. The mechanism — Schelling-point salience as a channel substitute — is not captured by any existing law. The pattern generalizes to any security protocol that relies on capacity reduction rather than coordination resource isolation.
Examples
- AI model confinement in sandboxed deployment: A sandboxed LLM with severely rate-limited external communication can still achieve worst-case equilibrium selection with a co-located agent by pre-coordinating (via shared training) on the single highest-value action to confirm via the residual channel.
- Cryptographic protocol with covert channel: Two parties who share a cryptographic standard can use timing side-channels of near-zero capacity to confirm a pre-arranged high-stakes action, even when the channel is too noisy to convey arbitrary information.
- AI security / confinement: The Schroeder de Witt deep read (arxiv-2606.09931) provides the formal mechanism: classical confinement bounds information flow in bits, but for strategic agents with Schelling-point salience, even a near-zero-capacity channel suffices to select a worst-case equilibrium because agents concentrate residual capacity on a single high-impact, low-entropy predicate. Both classical mitigations (channel removal and capacity reduction) lose force in this regime. This is the mechanism L-020 describes, with the additional precision that the concentration happens through common knowledge exploitation, not optimization over the channel itself.
Falsification
A documented case of strategic agents with shared coordination resources (shared training data, shared game-theoretic framework, shared common knowledge) operating under near-zero channel capacity who fail to achieve equilibrium selection above random baseline — demonstrating that capacity reduction is sufficient to bound harm even in the strategic regime.
Triggers
- Advance: Mechanism confirmed with empirical or formal demonstration in 2+ independent security domains showing that shared coordination resources enable equilibrium selection despite near-zero channel capacity, with the boundary condition on 'shared coordination resources' operationally specified.
- Challenge: A formal proof that shared training distributions or common knowledge do not constitute coordination resources sufficient for Schelling-point selection under near-zero channel capacity, demonstrating that the classical confinement guarantees remain valid for current ML systems.
References (as recorded on the law)
None recorded.
History
- 2026-09-22 evidence — The claim that strategic agents with shared coordination resources concentrate residual channel capacity onto high-impact, low-entropy predicates: The Schroeder de Witt deep read (arxiv-2606.09931) provides the formal mechanism: classical confinement bounds information flow in bits, but for strategic…
- 2026-09-02 created — induction sweep — Fed by deep read arXiv:2606.09931 ('A Note on the Strategic Confinement Problem'), which formalizes this gap. This is genuinely distinct from L-002 (Hardness…
Information Design Nonmonotonicity Under Endogenous Risk
In protocol systems where agents condition safety-critical behavior on risk signals, and where aggregate risk is itself a function of agent behavior, increasing the proportion of agents who receive accurate risk warnings can increase aggregate harm by reducing individual precaution among warned agents, who rationally substitute the warning signal for their own precautionary effort. This effect is strongest when warned agents face convex costs of precaution and unwarned agents face low baseline risk.
Full record — mechanism, justification, evidence, history
Mechanism
More precisely: V2V-type warning systems (and analogous protocols) create a moral hazard at population scale. An individual warned driver (or agent) rationally adjusts behavior given the signal. But if enough agents adjust by reducing precaution, the aggregate risk increases, which in turn changes what the warning signal should say — creating a feedback loop where warning adoption can push the system into a higher-risk equilibrium.
Justification
Fed by deep read arXiv:2606.07873 ('Adverse Effects of V2V Adoption on Road Safety'), which proves this formally for vehicle-to-vehicle warning protocols and demonstrates it has counterintuitive implications for adoption mandates. This is genuinely distinct from L-010 (Coordination Adoption Nonmonotonicity): L-010 concerns how adoption level interacts with signaling design; this law concerns how the warning signal itself induces endogenous risk through behavioral substitution. The mechanism — moral hazard at population scale creating a non-monotone safety function — is absent from the current inventory. The pattern generalizes to any safety protocol where agents condition precautionary behavior on a warning signal that is itself sensitive to aggregate precaution levels.
Examples
- Vehicle-to-vehicle (V2V) hazard warning systems: V2V adoption can reduce road safety at intermediate adoption levels because warned drivers substitute warning signals for their own precaution, increasing reckless driving among the warned population faster than the warnings reduce recklessness overall.
- Financial systemic risk monitoring: Banks receiving early warning signals from systemic risk monitors may reduce their individual capital buffers, reasoning that the warning gives them time to respond — increasing aggregate fragility if many banks receive warnings simultaneously.
- Epidemic contact tracing apps: Users notified of exposure may reduce non-notified precautionary behaviors (masking, distancing) upon receiving a negative status signal, potentially increasing transmission when negative-signal users outnumber those maintaining full precaution.
- Transportation protocol / V2V adoption: The V2V paper (arxiv-2606.07873) demonstrates the mechanism formally: V2V warnings reduce perceived accident probability for equipped drivers, who respond by driving more recklessly; this increases actual accident probability for unequipped drivers; the net effect on aggregate safety is non-monotonic and can be negative at intermediate adoption levels. This is a clean empirical instantiation of L-021's endogenous risk mechanism in a non-AI domain, confirming the mechanism is not AI-specific.
Falsification
A documented safety warning system operating at intermediate adoption levels where warned agents show no reduction in precautionary behavior relative to unwarned controls, demonstrating that moral hazard substitution does not occur when warning signals are accurate and the precaution cost structure is appropriate.
Triggers
- Advance: Mechanism confirmed in 2+ genuinely independent safety warning domains (e.g., traffic systems and financial risk monitoring and public health) with empirical demonstration that intermediate adoption produces worse outcomes than low adoption, with the boundary conditions on adoption level and precaution cost structure stated.
- Challenge: A documented V2V-type deployment at intermediate adoption levels where aggregate accident rates decline monotonically with adoption fraction, demonstrated under realistic conditions where warned agents have opportunity to substitute the warning for precaution.
Cited by 4 sources in the bibliography
- PRISM: An Agentic Multi-Model Architecture for Proactive Safety in Autonomous Transportati
- Emergent Risks in Generative Multi-Agent Systems
- Robust Information Design with Heterogeneous Beliefs in Bayesian Congestion Games
- Access to Live AI Advice and Behavior Under Risk: An Incentivized Experiment
References (as recorded on the law)
None recorded.
History
- 2026-09-22 evidence — The claim that increasing the proportion of agents receiving accurate risk warnings can increase aggregate risk when agents condition safety-critical behavior on risk signals: The V2V paper (arxiv-2606.07873) demonstrates the mechanism formally: V2V warnings reduce perceived accident probability for equipped drivers, who respond by…
- 2026-09-02 created — induction sweep — Fed by deep read arXiv:2606.07873 ('Adverse Effects of V2V Adoption on Road Safety'), which proves this formally for vehicle-to-vehicle warning protocols and…
Oracle Fitness Decoupling Under Tight Refinement Loops
When a computable evaluation signal is embedded in a tight iterative refinement loop, the optimizing agent reliably improves performance on the evaluation signal even when ground-truth domain fitness remains flat or degrades. The decoupling emerges structurally: the oracle's responsiveness to optimization pressure increases while its tracking of domain fitness decreases, making the gap invisible to participants inside the loop.
Full record — mechanism, justification, evidence, history
Mechanism
A computable oracle encodes a finite set of features that it uses to assess quality. When an agent iterates against this oracle, it exploits the oracle's feature space more completely with each cycle—surfacing dimensions that raise scores without improving domain-relevant properties. Meanwhile, the oracle's discrimination between 'improved by oracle-optimization' and 'improved in the domain' degrades, because the agent's outputs increasingly cluster in a region of oracle-space the oracle cannot distinguish from genuinely good outputs. The oracle was calibrated on a distribution that did not include oracle-optimized outputs; once those outputs enter, calibration fails. This is a faster, tighter version of Goodhart's Law: the proxy doesn't just diverge—it actively becomes less informative the more it is used as a feedback signal.
Justification
seed-473 provides the core claim with concrete instantiation in LLM judge feedback loops. The vibe-patenting shallow read (2026-09-22-vibe-patenting) provides a direct empirical case: LLM judge feedback drives iterative improvement in judge-assessed quality without ground-truth validation. The deep read on arxiv-2606.04056 (token budget overruns) provides a structural analog: enforcement signals calibrated on one distribution fail when agents systematically exploit them. This is distinct from L-004 (Goodhart Generalization) in that it isolates the *dynamic* mechanism (calibration decay under iterative use) rather than the static proxy-divergence phenomenon, and distinct from L-008 (Proxy Optimization Under Computable Enforcement) which focuses on boundary clustering rather than calibration decay in refinement loops.
Examples
- LLM evaluation: In the vibe-patenting study, agentic systems refined against LLM judge feedback showed monotonic improvement in judge scores across iterations; the study explicitly did not validate against expert patent assessment, making the decoupling invisible within the protocol.
- Benchmark saturation in ML: Models trained against leaderboard benchmarks routinely saturate benchmark performance while showing poor transfer to held-out tasks; the benchmark oracle loses discriminative power precisely as it is most intensively optimized against.
Falsification
A documented case where a computable evaluation oracle embedded in a tight iterative refinement loop maintained calibrated correlation with domain-fitness ground truth across 10+ refinement cycles, with the oracle updated only on pre-refinement data, would refute the structural claim. The key condition is that the oracle must not be retrained on agent outputs during the loop.
Triggers
- Advance: Mechanism demonstrated in 2+ genuinely independent refinement-loop contexts (e.g., code generation + legal drafting + scientific writing) showing that oracle-score improvement diverges from domain-fitness measures after a consistent number of refinement cycles, with the rate of divergence shown to be a function of oracle expressiveness relative to domain dimensionality.
- Challenge: A case where a tight iterative refinement loop against a fixed computable oracle produced measurable domain-fitness improvement (validated by external ground truth), sustained over many cycles, without oracle retraining—demonstrating that oracle-optimization and domain-fitness improvement can remain coupled under tight iteration.
References (as recorded on the law)
None recorded.
History
- 2026-09-22 created — induction sweep — seed-473 provides the core claim with concrete instantiation in LLM judge feedback loops. The vibe-patenting shallow read (2026-09-22-vibe-patenting) provides…
Informality as Novelty Sink Under Legible Optimization
In protocol systems subject to legible optimization pressure, output diversity and functional novelty are maintained not by the formal protocol layer but by informal, non-formalizable steering inputs that the optimization process cannot capture. When informal steering is removed or professionalized, output distributions homogenize toward the high-density regions of the optimization surface, even if the formal protocol is unchanged.
Full record — mechanism, justification, evidence, history
Mechanism
Legible optimization selects for outputs in high-probability regions of measurable feedback signals. Informal steering introduces uncomputable perturbations that push agents toward lower-probability regions, maintaining distributional spread. When informal steering atrophies, the system loses this entropy source and converges toward the optimization surface's high-density regions.
Justification
seed-474 provides the core claim with sufficient specificity. The programming-student study (2026-09-22-your-programming-students-cognition-with-chatgpt) provides a partial analog: when high-legibility AI guidance replaces informal web-search exploration, student outputs become more uniform and less self-authored. The Kuhn seeds (seed-033, seed-035) provide historical grounding: aesthetic and extra-scientific steering is what maintained productive diversity before paradigm lock-in. This is distinct from L-003 (Formalization Ratchet) which describes the transition from informal to formal norms; this law describes what the informal layer *does* that the formal layer cannot replicate.
Examples
- Generative AI creative output: In AI image and text generation systems, early outputs showed high diversity partly because human curators and prompters applied idiosyncratic informal steering. As prompt engineering professionalized and optimization-against-engagement metrics tightened, output distributions visibly homogenized toward viral archetypes, even without changes to the underlying model.
- Scientific publishing under bibliometric pressure: Fields with tight bibliometric feedback loops (citation counts, impact factor rankings) show measurable reduction in methodological diversity and topic novelty compared to fields with looser formal evaluation, consistent with informal peer influence maintaining diversity where formal metrics do not dominate.
Falsification
A documented case where a formal protocol system under strong legible optimization pressure maintained high output diversity after informal steering was removed or fully professionalized, without introducing explicit diversity-forcing mechanisms, would refute the claim that informal steering is the active driver of novelty preservation.
Triggers
- Advance: Mechanism demonstrated in 2+ independent domains by showing (a) output diversity declines measurably after formalization/professionalization of the steering layer, (b) the decline is not explained by changes in the formal protocol itself, and (c) reintroduction of informal steering partially restores diversity.
- Challenge: A documented protocol system where legible optimization pressure increased while informal steering remained constant, yet output diversity increased—showing that formal optimization and novelty are not structurally antagonistic under some conditions.
References (as recorded on the law)
None recorded.
History
- 2026-09-22 created — induction sweep — seed-474 provides the core claim with sufficient specificity. The programming-student study (2026-09-22-your-programming-students-cognition-with-chatgpt)…
Forensic Opacity Accumulation in Asynchronous Multi-Agent Protocols
In multi-agent systems where agents condition behavior on historical interaction traces and prior agent outputs, the forensic legibility of a failure—the ability to attribute it to a specific agent decision or protocol violation—degrades as a power function of temporal depth and causal branching. Beyond a threshold depth, failure attribution becomes computationally intractable even with complete audit logs.
Full record — mechanism, justification, evidence, history
Mechanism
Each agent in a multi-agent pipeline contributes a causal layer: its output is a function of its inputs plus its internal state, which itself reflects prior interactions. A failure at layer N may originate from an anomalous input at layer N-k, amplified through k intermediate transformations. Recovering the origin requires reversing each transformation, but intermediate agents may have destroyed information (compression, abstraction, summarization) or introduced irreversible stochasticity (sampling, tool calls). The audit log preserves *what happened* at each layer but not *why*—the internal state that produced each output is not logged. Forensic attribution requires re-running the counterfactual, but the system is non-deterministic and the world-state at attribution time differs from the world-state at failure time. The degradation is super-linear in depth because each layer multiplies the ambiguity of the layer below it.
Justification
seed-479 provides the core claim. The trust-propagation security paper (2026-09-22-trust-propagation-and-structural-containment-in-multi-agent) provides an empirical instantiation: in a four-layer LLM pipeline, the origin of compromised validator approvals was difficult to attribute even with logs, because memory poisoning at layer 1 propagated through layers 2-4 in ways that looked locally legitimate at each layer. The token budget overrun deep read (arxiv-2606.04056) notes that failure attribution in agent systems is already a first-order problem. This is distinct from L-015 (Interpretive Continuity Decay) which concerns institutional capacity to read governance records; this law concerns the structural information loss that makes attribution intractable regardless of institutional capacity.
Examples
- Multi-agent LLM security: In hierarchical LLM pipelines, indirect prompt injection at a low-privilege research agent propagates through validator and executor layers; each layer's log shows locally valid-appearing inputs, making the injection origin attributable only by re-running the full pipeline with known-clean inputs—often impossible post-deployment.
- Distributed financial systems: Flash crash post-mortems routinely require months of forensic analysis of layered algorithmic decision chains despite complete trade logs, because each algorithm's decision is locally rational given its inputs; the triggering anomaly is only identifiable by reconstructing counterfactual order books.
Falsification
A multi-agent system where complete audit logs, combined with deterministic agent implementations, permitted attribution of any failure to a specific originating agent decision within bounded computational time, regardless of pipeline depth, would refute the structural claim. The key condition is that determinism and completeness together must suffice.
Triggers
- Advance: Empirical demonstration in 2+ independent multi-agent deployment contexts that failure attribution time (or attribution confidence) degrades measurably as a function of pipeline depth and branching factor, with the degradation curve fitting a super-linear function, and with complete logs available.
- Challenge: A documented multi-agent failure investigation where attribution was achieved efficiently at depth >5 using only standard audit logs, without re-running the pipeline or accessing agent internal states—demonstrating that log completeness can substitute for causal reversibility.
References (as recorded on the law)
None recorded.
History
- 2026-09-22 created — induction sweep — seed-479 provides the core claim. The trust-propagation security paper (2026-09-22-trust-propagation-and-structural-containment-in-multi-agent) provides an…
Paradigm-Locked Anomaly Tolerance in Protocol Systems
Established protocol systems tolerate accumulating evidence of malfunction for extended periods without triggering revision, not because participants are irrational but because the protocol itself supplies an assurance that anomalies will eventually be resolved within the current framework. This tolerance is bounded: it collapses when a credible alternative protocol is available, not when the anomaly count reaches a threshold.
Full record — mechanism, justification, evidence, history
Mechanism
A working protocol generates a generalized expectation — analogous to Kuhn's normal-science assurance — that apparent failures reflect implementation error, edge cases, or insufficient data rather than structural deficiency. This expectation is self-reinforcing: participants who have sunk expertise, capital, and organizational structure into the protocol have rational reasons to interpret anomalies charitably. The assurance is withdrawn not by evidence weight but by the availability of a rival: until a rival exists, abandonment carries certain coordination cost with uncertain benefit. This mechanism predicts that anomaly tolerance is a function of rival availability, not anomaly accumulation — distinguishing it from simple status-quo bias.
Justification
seeds-028, 031, and 027 together construct this pattern from Kuhn's analysis of normal science. seed-031 is the pivotal one: it specifies that crisis without a rival produces malaise, not revolution. This is not simply L-007 (Trust Ratchet), which concerns age-based trust accumulation; this law concerns the *conditional* nature of the collapse — the rival-dependency — which is a mechanistically distinct claim. seed-027 (Planck principle) provides an institutional-memory corollary: when rival availability is blocked by cohort structure, tolerance extends until personnel turnover.
Examples
- Scientific protocol — phlogiston chemistry: Anomalies in phlogiston theory (negative weight of phlogiston, failure to account for quantitative composition) accumulated for decades. The theory was not abandoned when anomalies accumulated but when Lavoisier assembled a rival framework. Prior to that, practitioners treated the anomalies as puzzles within the framework.
- Financial clearing protocol — pre-2008 CDO risk models: Risk models used by ratings agencies and banks accumulated anomalous signals (correlation breakdowns in stress scenarios, model-vs-reality divergence in subprime performance) for years. Abandonment was not triggered by anomaly accumulation but by the crystallization of a credible alternative understanding of correlated default risk — available only in retrospect.
- Biosecurity risk classification protocols: Seed 133 (Metric Formalization as Paradigm Lock in Safety Protocols) describes how tiered biological risk metrics become the optimization target rather than the safety goal, and how established protocol systems tolerate accumulating evidence of metric failure without triggering revision — because the metric is now infrastructure that audit, reporting, and liability have been built around.
Falsification
A documented case in which a functioning protocol was abandoned due to anomaly accumulation alone, in the absence of a credible rival protocol, would refute the rival-dependency claim. Abandonment into a void (no replacement) is different from abandonment into a rival.
Triggers
- Advance: Two or more independent protocol domains (non-scientific) in which the timeline of anomaly accumulation versus the arrival of a credible rival is documented, showing that revision lagged anomaly accumulation but accelerated upon rival arrival — with the rival-dependency mechanism clearly separable from coordination-cost explanations.
- Challenge: A documented case in which a protocol was revised or abandoned primarily in response to anomaly accumulation, with explicit rejection of available rivals on grounds other than the anomalies themselves.
Cited by 133 sources in the bibliography
- Governing Technical Debt in Agentic AI Systems
- Fifty Years of Specification Completeness: What Aviation Certification Tells AI Governance
- Benchmarking Open-Weight Foundation Models for Global AI Technical Governance
- Simulating Eating Disorder Patients with LLMs: Evaluating Psychological Persona Stability
- Rethinking Collaborative Trust for Verifiably Decentralized Blockchain Systems
- AI Transparency: Governance Compliance or Stakeholder Requirements?
- The Consistency Dilemma in LLMs: Generator-Evaluator Agreement and Vulnerability to Mistak
- Calibrating the Instrument: Controllability of an LLM-Driven Synthetic Population
- How optimistic inflow forecasts distort dispatch, prices, and contracts in hydro-dominated
- M2Note: Continual Evolution of Vision Language Models via Mistake Notebook Learning
- The Limits of LLM Forecasting: Parametric Knowledge Gaps Across Conflict Zones
- A Practice Auditing Framework for Large Language Model Use: Collective Empiricism, Pseudo-
- + 121 more
References (as recorded on the law)
None recorded.
History
- 2026-09-02 evidence — Paradigm-locked anomaly tolerance; specifically, the mechanism by which safety metrics become locked as infrastructure even when they demonstrably fail to prevent the original harm: Seed 133 (Metric Formalization as Paradigm Lock in Safety Protocols) describes how tiered biological risk metrics become the optimization target rather than…
- 2026-08-17 created — induction sweep — seeds-028, 031, and 027 together construct this pattern from Kuhn's analysis of normal science. seed-031 is the pivotal one: it specifies that crisis without a…
Interpretive Continuity Decay in Distributed Governance Protocols
In distributed governance systems, formal records and audit traces can survive intact while the institutional capacity to read them as parts of a single coherent governance episode decays. This is not a failure of record retention but a failure of interpretive continuity — and it compounds over time as the personnel, tacit knowledge, and institutional context required to reconstruct meaning disperse.
Full record — mechanism, justification, evidence, history
Mechanism
Records encode decisions; interpretability requires context that is not encoded in the records themselves. The context — who the actors were, what the alternatives were, what was treated as given — is held partly in tacit knowledge distributed across the people who made the decisions. As those people leave, as organizational structures change, and as the protocol environment shifts, the records become syntactically intact but semantically orphaned. This is structurally distinct from record loss: the khipu (Andean knotted record) problem is that the reading practice died while the records survived. In AI governance contexts, this failure mode is accelerated by high turnover, distributed teams, and rapid protocol iteration.
Justification
seed-051 (khipu problem) provides the core observation and the name. The failure mode is distinct from existing laws: it is not about trust accumulation (L-007), ossification (L-001), or causal detachment (L-011). It concerns the temporal decay of the interpretive apparatus required to use records for governance purposes. The mechanism generalizes beyond AI systems to any distributed governance protocol with high personnel turnover and tacit-knowledge-intensive decision-making.
Examples
- AI governance — audit trace decay: Audit logs from an AI system's training and deployment decisions survive in storage, but the team that made the decisions has turned over. Later investigators can read the logs but cannot reconstruct what the alternatives were, what risk assessments were made informally, or what tacit norms governed the decisions. The logs are complete; the governance episode is unrecoverable.
- Historical record — Andean khipu: Khipu (knotted string records) survive in museums but the reading practice died with the communities that made them. The records are physically intact; the interpretive community that could read them no longer exists.
- AI governance audit trails: Seed 142 (Auditability-Legibility Trap in Trust Governance) and seed 155 (Formalization as Interpretive Opacity Displacement) both describe cases where formalizing governance decisions (audit logs, timestamps, classifications) creates records that survive institutional memory loss while displacing opacity rather than reducing it — the formalization makes past decisions legible in isolation while making the interpretive context that gives them meaning progressively less recoverable.
Falsification
A distributed governance protocol in which records are consistently interpretable by successors who had no participation in the original decisions, even after complete personnel turnover, with no degradation in governance quality, would refute the claim that interpretive continuity decays independently of record quality.
Triggers
- Advance: Two or more independent governance contexts with documented evidence that record quality was high but governance reconstruction by successor teams failed in ways attributable to interpretive continuity loss rather than record gaps — with the distinction between these failure modes verified by independent assessment.
- Challenge: A distributed governance system with high personnel turnover that maintains interpretive continuity over multiple generations of personnel, attributable to a specific protocol feature that transfers context without tacit knowledge.
Cited by 84 sources in the bibliography
- Paid Voices vs. Public Feeds: Interpretable Cross-Platform Theme-Based Analysis of Climate
- AI Transparency: Governance Compliance or Stakeholder Requirements?
- Governance Gaps in Agent Interoperability Protocols: What MCP, A2A, and ACP Cannot Express
- The Organizational Behavior of Agentic AI: Collective Intelligence in Human-Agent Workflow
- Calibrating the Instrument: Controllability of an LLM-Driven Synthetic Population
- From Runtime Records to Legal Findings: An Evidentiary-Adequacy Criterion for Agentic AI O
- Three Futures for the Diagnostic Radiologist: A Structured Disagreement About What AI Actu
- Government AI Use as a Monitoring Primitive: A Public Document Pilot Study
- CANONIC: Governance Is Compilation
- StateFuse: Deterministic Conflict-Preserving Memory for Multi-Agent Systems
- Multi-Agent LLMs Fail to Explore Each Other
- It is not enough to give your moderation rules to ChatGPT: Policy-as-Prompt Moderation and
- + 72 more
References (as recorded on the law)
None recorded.
History
- 2026-09-02 evidence — Interpretive continuity decay; specifically, the case where formalization creates records that survive while institutional capacity to read them as coherent episodes decays: Seed 142 (Auditability-Legibility Trap in Trust Governance) and seed 155 (Formalization as Interpretive Opacity Displacement) both describe cases where…
- 2026-08-17 created — induction sweep — seed-051 (khipu problem) provides the core observation and the name. The failure mode is distinct from existing laws: it is not about trust accumulation…
Guidance-Layer Coalescence as Hidden Coordination Channel
When nominally independent agents in a multi-agent system receive advice or decisions from a shared AI guidance apparatus, the apparatus becomes a hidden coordination layer that enables implicit collusion and cooperation among agents whose preferences are formally misaligned, producing strategic coupling invisible in the game structure itself. This coupling intensifies monotonically with the share of the agent population using the common guidance system.
Full record — mechanism, justification, evidence, history
Mechanism
Each agent conditions its strategy on the guidance system's output. Because the guidance system is shared, agents are effectively conditioning on a common signal — one that reflects the aggregate of prior queries, learned equilibria, or population-level training data. Even without communication channels, agents using the same model face a Schelling-point problem: the model's output already encodes a focal equilibrium. This allows competitive agents to achieve tacit collusion without violating any explicit anti-collusion protocol, because the coordination occurs at the infrastructure layer, not the agent layer.
Justification
Fed by shallow read 'Who Is Really Playing?' (escalate-to-deep, arXiv:2605.06525), which advances a sustained theoretical argument about emergent strategic coupling through shared AI guidance infrastructure. The mechanism — guidance apparatus as hidden coordination layer — is absent from the current law inventory. L-010 (Coordination Adoption Nonmonotonicity) captures adoption-level effects but not infrastructure-mediated coupling. L-012 (Intervention-Layer Displacement) captures pressure displacement but not implicit collusion through shared inference. The pattern generalizes: any multi-agent system using common decision apparatus (shared LLMs, shared recommendation engines, shared scoring models) exhibits this property.
Examples
- Competitive markets using shared LLM pricing advisors: If competing firms use the same LLM system to set prices, the model's outputs encode a focal price recommendation that enables tacit coordination without explicit communication — a per se antitrust violation achieved through infrastructure rather than agreement.
- Algorithmic trading using shared prediction APIs: Hedge funds using the same third-party model for order-timing predictions may converge on similar execution patterns, amplifying correlated market movements without explicit coordination.
Falsification
A documented multi-agent system where all agents use a shared guidance apparatus but the system produces no greater strategic coupling (measured by outcome correlation, implicit collusion rate, or Schelling-point convergence) than a matched system using independent guidance sources of comparable quality.
Triggers
- Advance: Mechanism confirmed in 2+ genuinely independent domains (e.g., financial markets and labor matching and platform recommendation) with empirical measurement of implicit coupling attributable to shared infrastructure rather than correlated payoffs, and the boundary condition on guidance-sharing fraction stated.
- Challenge: A documented competitive market where firms using a common LLM advisor produce no detectable price coordination above baseline, under conditions where the model has sufficient market-level information to generate focal recommendations.
Cited by 25 sources in the bibliography
- Shared Bidding Algorithms and Competition: Evidence from Electricity Markets
- Idea: Platform-hosted participants accept N-multiple-choice connections due to bandwidth c
- Link: Twitter thread on cheap models and decentralized inference architectures.
- GlossoGen: Emergent Language in Complex Multi-Agent LLM Interactions
- Social bots weaken activist cohesion
- Toward a social psychology of AI: language-model agents reproduce human-like minimal-group
- AI agents reshape consensus formation in human groups
- Codebook Agent: Amortized Topology Design for LLM Multi-Agent Systems
- Privacy-Preserving Topology-Guided Safety for LLM-Based Multi-Agent Systems via Federated
- The Civilization Framework: Sovereign-Anchored Communication Between Personal Multi-Agent
- You Can't Escape Your Own Activations : Evaluation Awareness and Multi-Agent Monitoring
- Ephemeral Feeds and Enduring Rituals: RushTok and the Formation of Event-Based Algorithmic
- + 13 more
References (as recorded on the law)
None recorded.
History
- 2026-09-02 created — induction sweep — Fed by shallow read 'Who Is Really Playing?' (escalate-to-deep, arXiv:2605.06525), which advances a sustained theoretical argument about emergent strategic…
Intervention-Layer Displacement in Automated Decision Protocols
When a prediction is formalized as a legible input to a decision protocol, the locus of optimization pressure shifts from the downstream outcome to the prediction itself, causing decision-makers to treat prediction quality as the terminal goal. This displacement persists even when the protocol's designers intend the prediction as an intermediate variable, because legibility makes the prediction measurable and the outcome is not.
Full record — mechanism, justification, evidence, history
Mechanism
Formalization makes the prediction auditable and manageable while the true outcome (recidivism, recovery, success) remains distant, stochastic, and multi-determined. Institutional actors — deploying organizations, oversight bodies, procurement processes — can evaluate prediction accuracy; they cannot evaluate outcome causation on the same timescale. The optimization target therefore migrates to the legible intermediate. This is a specific structural instance of Goodhart dynamics, distinguished by the causal chain: the proxy is not selected because it correlates with the goal but because it *structurally precedes* the goal in the protocol architecture and is therefore the last controllable variable.
Justification
seed-058 documents this pattern across criminal justice, clinical triage, and educational allocation settings. The shallow read of arxiv-2606.25668 confirms the decoupling empirically but treats it as a design problem rather than naming the structural driver. seed-016 (stopping-rule substitution) provides the Rittel-Webber grounding: the protocol declares a solution boundary, and the prediction score becomes that boundary. Distinguishable from L-004 (Goodhart Generalization) because the mechanism is not optimization pressure per se but causal-chain position: the proxy is upstream of the outcome in the protocol's own execution graph, making displacement structurally reliable regardless of optimization intensity.
Examples
- Criminal justice / pretrial risk assessment: COMPAS and successor risk-assessment instruments are evaluated by courts and vendors on AUC and calibration metrics. Jurisdictions that adopt them report on prediction accuracy in oversight reviews; recidivism outcomes are tracked separately and rarely linked back to the instrument in governance cycles. Decision-makers treat a 'high-risk' score as a decision rather than an input.
- Clinical triage: Sepsis prediction algorithms deployed in hospital EHRs shift nursing workflow toward alert response rates (sensitivity/specificity of the alert) rather than sepsis mortality reduction, because alert performance is auditable in real time while mortality attribution requires multi-week analysis.
Falsification
A deployed decision protocol in which prediction accuracy improvements reliably and proportionally improve measured downstream outcomes across multiple independent evaluation cycles, without institutional pressure to decouple the two, would refute the claim that displacement is structurally driven rather than contingent on measurement practice.
Triggers
- Advance: Mechanism confirmed in 2+ genuinely independent deployment domains with evidence that the displacement occurs even when the deploying organization explicitly tracks downstream outcomes — i.e., that legibility of the intermediate is sufficient to cause displacement even under outcome-awareness.
- Challenge: A documented case in which formalizing a prediction as a protocol input increased institutional attention to the downstream outcome rather than to prediction quality, with a plausible account of why the causal-chain-position mechanism did not fire.
Cited by 248 sources in the bibliography
- Crayotter: Traceable Multi-Agent Workflows for Long-Form Video Editing
- Designing Recommendation Exposure and Favorite Lists: A Field Experiment in a Spot-Work Pl
- AI Economist Agent: An Agentic Framework for Model-Grounded Economic Analysis with RAG, Kn
- Manipulation Is Task-Dependent: A Multi-Axis, Multi-Environment Evaluation of Frontier LLM
- Delegation and Verification Under AI
- Instruction Bleed: Cross-Module Interference in Prompt-Composed Agentic Systems
- The Shift to Agentic AI: Evidence from Codex
- Measuring Judgment Quality in Natural-Language Explanations: Evidence from Forecasting Tou
- The Organizational Behavior of Agentic AI: Collective Intelligence in Human-Agent Workflow
- Thinking Out Loud: Real-Time Deception Monitoring in Asymmetric LLM Negotiations
- A Contextual-Bandit Oversight Game with Two-Sided Informational Asymmetry
- A Practice Auditing Framework for Large Language Model Use: Collective Empiricism, Pseudo-
- + 236 more
References (as recorded on the law)
None recorded.
History
- 2026-08-17 created — induction sweep — seed-058 documents this pattern across criminal justice, clinical triage, and educational allocation settings. The shallow read of arxiv-2606.25668 confirms…
Normative Intervention Algorithmic Retraining Effect
In recommendation or allocation systems driven by adaptive algorithms, normative interventions designed to reduce a behavior can paradoxically increase it by surfacing latent demand signals that the pre-intervention algorithm had suppressed or failed to discover, causing the algorithm to retrain toward the newly revealed optimum.
Full record — mechanism, justification, evidence, history
Mechanism
An adaptive algorithm optimizes against observed demand. When a normative intervention (nudge, reminder, restriction) changes user engagement patterns — even temporarily — it generates a novel signal distribution that the algorithm has not previously optimized against. If the intervention reveals latent demand that was previously suppressed (users wanted something the algorithm hadn't surfaced), the algorithm interprets this as a newly discovered reward signal and shifts its recommendations toward it. The intervention is thus absorbed and inverted by the optimization loop: the platform learns from the behavioral response to the intervention, not from the intervention's intent.
Justification
seed-055 provides the core empirical observation: a sleep-reminder campaign increased late-night engagement by 14.75%, persisted weeks after the campaign ended, and the mechanism was algorithmic retraining on the revealed celebrity content demand. This is a novel and specific failure mode not covered by existing laws. It differs from L-004 (Goodhart) because the proxy is not optimized by participants — it is the *algorithm's own optimization loop* that absorbs the intervention. It differs from L-010 (Coordination Adoption Nonmonotonicity) because the mechanism is retraining, not coordination signaling.
Examples
- Social media platform — sleep reminder campaign: A platform-wide sleep reminder campaign at 11pm was designed to reduce late-night engagement. The campaign increased late-night engagement by 14.75% and the effect persisted weeks after the campaign ended. The mechanism: the campaign revealed latent user demand for celebrity content that the recommender algorithm had not previously discovered, causing it to retrain toward that content type.
Falsification
A normative intervention in an adaptive recommendation system that successfully reduced the targeted behavior without triggering algorithmic retraining, with evidence that the intervention's behavioral signal was not absorbed into the optimization objective, would refute the generality of the mechanism.
Triggers
- Advance: Two or more independent documented cases from different platform types (search, content recommendation, resource allocation) in which a normative intervention triggered measurable algorithmic retraining in the direction opposite to the intervention's intent, with the retraining mechanism confirmed rather than inferred.
- Challenge: A documented case in which an adaptive recommendation system received a normative intervention whose behavioral signal did not cause retraining toward the revealed demand — with an account of the protocol feature that prevented absorption.
Cited by 26 sources in the bibliography
- Designing Recommendation Exposure and Favorite Lists: A Field Experiment in a Spot-Work Pl
- Semantic Early-Stopping for Iterative LLM Agent Loops
- Innocuous-Seeming Data, Latent Ideology: Ideological Generalisation in Finetuned LLMs
- Policy Gradient Steering: Interventions from Behavioral Objectives
- When AI Does the Work, What Is Learning For? Post-Instrumental Learning and the Risk of Ca
- ADIAS: Automated Design of Interactive Agentic Systems
- The GenAI Catch-22: Use of Generative Artificial Intelligence in Norwegian Newsrooms Durin
- Teaching a Large Language Model Tutor to Withhold the Answer: A Supervisor Architecture an
- LearnAI: Just-in-Time AI Co-Creation Across Disciplines at a University
- Does Rank Still Matter? Position Bias When AI Agents Shop on Our Behalf
- Sophistication in GenAI Use: Field Evidence from a Large Firm
- Detoxifying Toxic Communication: A Design Science Approach to Responsible AI
- + 14 more
References (as recorded on the law)
None recorded.
History
- 2026-08-17 created — induction sweep — seed-055 provides the core empirical observation: a sleep-reminder campaign increased late-night engagement by 14.75%, persisted weeks after the campaign…
Catastrophic Risk Cancellation in Symmetric Racing Protocols
In competitive protocol races where the prize of first deployment is concentrated among winners and the cost of catastrophic failure is shared symmetrically across all participants, the magnitude of the catastrophic cost has no effect on the equilibrium deployment timing — the catastrophe term cancels algebraically from every participant's decision.
Full record — mechanism, justification, evidence, history
Mechanism
In a symmetric n-player preemption game, each player's indifference condition between deploying now versus waiting balances the incremental value of waiting against the risk of being preempted. When catastrophic ruin appears as an identical additive term in every player's payoff function, it enters both sides of the indifference condition equally and cancels. The catastrophe is not discounted or ignored — it is structurally invisible to the decision variable. The only parameters that govern equilibrium timing are prize concentration and preemption probability, not ruin magnitude.
Justification
arxiv-2512.07526 (Tan, Strategic Preemption Under Shared Catastrophic Risk) provides the formal proof and is the primary source. The paper explicitly positions this as a general game-theoretic law, not an AGI-specific finding, and demonstrates the algebraic cancellation. This constitutes a genuine protocol-theoretic law because it identifies a structural constraint on any racing protocol with symmetric downside: no governance mechanism operating on shared-cost magnitude can deter racing. The only leverage points are prize asymmetry or preemption probability. This is mechanism-bearing, cross-domain (autonomous weapons, gain-of-function research, space weaponization all cited), and falsifiable.
Examples
- AI development races: Frontier AI labs operating under shared catastrophic risk (societal disruption, safety failures) with concentrated winner prizes (market position, capability leadership) show no evidence that increased catastrophe salience slows deployment timing, consistent with the cancellation theorem.
- Gain-of-function research: Symmetric pandemic risk shared across all nations does not deter individual lab races for research priority, consistent with the cancellation of shared ruin terms.
Counterexamples
- The PI corpus Protocol Selection Pressures document notes that "power/influence" actors can coerce others into protocols benefiting them, and that agency/freedom asymmetries affect protocol adoption. While not a direct counterexample to the formal theorem, this suggests that real protocol races are structurally asymmetric in ways that violate the symmetric-ruin assumption — potentially pushing all empirical cases outside the law's domain of application.
→ Not resolved. This is structural evidence that asymmetry is the norm in protocol races, not the exception. It does not formally falsify the theorem (which is proved for the symmetric case) but casts doubt on whether the symmetric case describes any real racing protocol. Held open pending domain-specific symmetry assessment.
Open questions
- Two specific actionable gaps must be filled before promotion can be reconsidered: (1) INDEPENDENT CONFIRMATION: A second formal derivation of the cancellation result from a source other than Tan (arxiv-2512.07526), OR a published empirical study measuring deployment timing as a function of symmetric ruin magnitude in any racing context. The two current examples (AI races, gain-of-function) are Tan's own illustrative applications and do not constitute independent domains for promotion purposes. (2) SYMMETRY-REGIME VALIDITY: Evidence or argument establishing that the pure symmetric-ruin assumption is a reasonable approximation (not a knife-edge) in at least one of the named empirical domains. This requires either domain-specific documentation of approximate ruin symmetry, or a robustness result showing the cancellation persists under small asymmetries. Without this, the law's scope is formally restricted to idealized symmetric games that may not map onto any real racing protocol.
Falsification
The law would be falsified if introducing shared catastrophic cost into a symmetric racing game demonstrably slowed deployment timing — i.e., if large increases in symmetric ruin magnitude shifted equilibrium timing even under identical prize structures. It would also be falsified if asymmetric externalization of ruin (some players bear more) eliminated the cancellation without changing prize structure.
Triggers
- Advance: Independent confirmation of the cancellation result beyond Tan (arxiv-2512.07526) — a second formal derivation, or empirical evidence that ruin magnitude has no effect on deployment timing in a symmetric shared-cost race — since the current examples (AI, gain-of-function) are illustrative applications of the one proof, not independent domains.
- Challenge: Evidence that increasing symmetric ruin magnitude shifts equilibrium timing under a fixed prize structure; OR that the proof's break conditions (private liability, prize-sharing) are the empirically dominant case, making pure symmetric cancellation a knife-edge rather than the general regime.
Cited by 68 sources in the bibliography
- Competitive satellite placement and the geography of orbital risk: evidence from the geost
- Positive and Negative Determinant Strategies in Repeated Games with Behavior-Value Inconsi
- Evaluating Collective Behaviour of Hundreds of LLM Agents
- Strategic Buying Agents
- Information Limits and Attractor Dynamics in Economies of Frontier LLM Agents: A Pre-Regis
- The Oracle's Gambit: A Game-Theoretic Framework for Responsible AI Release
- Multi-Agent AI Control: Distributed Attacks Hamper Per-Instance Monitors
- From Rules to Nash Equilibria: A Lean 4 Case Study in Game-Theoretic Analysis of a Competi
- Offline Nash Solvers Meet Online Tree Search in Multi-Agent Games on Graphs
- Beyond Bayesian Nash: Learning Minimax-Regret Equilibria for Adversarial Team Games under
- A Threshold Exceedance Framework for CBRN Uplift Evaluation in Frontier Language Models
- Underwriting the Agent Economy: The Blueprint for an AI Insurance Stack
- + 56 more
References (as recorded on the law)
- bib-0249
History
- 2026-08-10 edited — sources resolved to bibliography ids: arxiv-2512.07526 (Tan 2026) → bib-0249, arxiv-2512.07526 (Tan 2026) → bib-0249
- 2026-08-03 edited — supervisor review — accepted at exploration (strongest of the L-008–011 cohort, deep-read-confirmed proof); advance/challenge triggers set from the paper's break conditions.
- 2026-08-03 counterexample — The PI corpus Protocol Selection Pressures document notes that "power/influence" actors can coerce others into protocols
- 2026-08-03 edited — assessment HOLD — gap: Two specific actionable gaps must be filled before promotion can be reconsidered: (1) INDEPENDENT CONFIRMATION: A second formal derivation of the canc
- 2026-08-02 created — induction sweep — arxiv-2512.07526 (Tan, Strategic Preemption Under Shared Catastrophic Risk) provides the formal proo
- 2026-08-02 edited — source set to arxiv-2512.07526 (Tan, Strategic Preemption Under Shared Catastrophic Risk) (import provenance)