L-009

Falling Behind Drives Unsafe Development in an Idealised AI Race Experiment

Source: econ.GN updates on arXiv.org — https://arxiv.org/abs/2607.26034 Date read: 2026-09-02 Connected to: L-009, seed-048 Kind: content Escalation: store-only Escalation rationale:

What this is

A behavioral economics framed experiment testing AI race dynamics through paired repeated-choice games between safe and unsafe development paths. Participants face uncertain termination horizons and competitive payoff structures designed to model competitive pressure and capability races. Primary finding: actors behind in the race adopt unsafe development strategies at elevated rates, even when collective rationality would favor coordination on safety.

What I took from it

This is empirical work on a game-theoretic phenomenon (the prisoner's dilemma structure of competitive safety), not a protocol systems phenomenon. The paper establishes that relative position (falling behind) is a behavioral trigger for unsafe action within a known, symmetric game. Crucially: the experimental frame is transparent to participants. There is no opacity, no legibility mismatch, no proxy capture—only rational defection under competitive pressure with perfect information about payoffs.

For the new nature inventory, this is confirmatory on L-009 (Catastrophic Risk Cancellation in Symmetric Racing Protocols) but does not extend it. The mechanism is elementary: winning matters more than safety when the prize is concentrated. The work does not examine what happens when: - The safety cost structure itself is opaque or subject to misestimation - Protocol enforcement makes certain unsafe moves legible in ways that create new optimization targets (seed-081) - Agents are not symmetric in their ability to detect or enforce safety norms - The race operates under formalized, computable compliance signals (L-014, L-008)

It is a clean, honest experimental confirmation of a baseline case. It does not generalize the mechanism into the regime where artificial protocols create novel incentive distortions.

Research connections

  • L-009: Confirms the baseline case—concentrated prize + competitive pressure → safety abandonment. Does not test whether asymmetric information, legibility capture, or proxy optimization amplifies the effect.
  • seed-048: Supports the existence of the phenomenon but under perfect information; does not interrogate how institutional or computational legibility changes the dynamics.

Seed

Seed title: none

Seed type: —

Seed text: —


RATIONALE FOR STORE-ONLY: This paper provides clean experimental evidence for a well-understood game-theoretic principle (competitive incentive misalignment). It does not introduce a novel mechanism, does not reveal a failure mode absent from the current inventory, and does not extend L-009 into the regime where protocol formalization or legibility asymmetry becomes operative. It is a good confirmation study, not an induction accelerant.