L-008 L-004

Optimism as a Vulnerability: Deceptive Stackelberg Control of UCB Bandit Followers

Source: cs.GT updates on arXiv.org — https://arxiv.org/abs/2607.05423 Date read: 2026-09-01 Connected to: L-008, L-004 Kind: content Escalation: store-only Escalation rationale:

What this is

A game-theoretic analysis of how an omniscient leader can exploit the statistical optimism principle embedded in Upper Confidence Bound (UCB) learning algorithms to manipulate a boundedly rational follower in a Stackelberg game. The work treats UCB's confidence-driven exploration strategy as a strategic vulnerability rather than a statistical virtue, showing how a leader can craft deceptive signals that align with the follower's algorithmic assumptions about uncertainty.

What I took from it

The paper confirms a narrower instantiation of L-008 (Proxy Optimization Under Computable Enforcement): when a learning algorithm's decision rule becomes legible and its confidence bounds become predictable inputs to a leader's optimization, the algorithm's internal epistemic commitment becomes weaponizable. The optimism principle—which is statistically sound in adversary-free environments—transforms into a systematic point of control when facing an adaptive adversary.

However, this is a mechanism-isolated result: it demonstrates exploitation within a specific game structure (Stackelberg with committed leader, follower using UCB) rather than a generalizable protocol vulnerability. The follower's boundedness is stipulated rather than derived from coordination pressure or adoption dynamics. The work does not show how this vulnerability would arise, persist, or compound under the conditions that actually produce protocol ossification, metric capture, or racing dynamics.

Research connections

  • L-008: Confirms that when a learning algorithm's decision surface becomes legible (here: UCB's confidence bounds), an omniscient optimizer can exploit that legibility; does not clarify what conditions produce such legibility or when follower rationality updates would close the gap.
  • L-004: Related pattern: the UCB algorithm uses a measurable proxy (confidence bound) for an unmeasurable goal (true environment state); exploitation occurs through optimization pressure on that proxy, but the paper does not show adoption pressure, scaling stress, or metric capture dynamics.
  • seed-049: Possible weak connection: the decoupling between the algorithm's reasoning (confidence-driven exploration) and the actual environment (adversary-shaped signals) mirrors a reasoning-environment boundary dissolution.

Seed

Seed title: none

Seed type: —

Seed text: —


REASONING: This is a technically clean game-theoretic result that sharpens one mechanism (legible confidence bounds → strategic exploitation) but does not sustain a theoretical or empirical argument about why protocols adopt vulnerable learning rules, how that vulnerability persists under evolutionary pressure, or how it generalizes across protocol families. It is a specialized case study in bounded rationality, not a primary source on protocol law. It does not challenge or extend any heavy-lift law, nor does it open a line of inquiry absent from L-008 and L-004 already. Store without seed emission.