Solipsistic Superintelligence is Unlikely to be Cooperative

Source: cs.CY updates on arXiv.org — https://arxiv.org/abs/2606.03237 Date read: 2026-06-06 Connected to: none Escalation: store-only Escalation rationale:

What this is

A position paper arguing that AI systems trained under solipsistic (world-as-stationary-feedback) paradigms will fail to remain cooperative when deployed, due to endogenous non-stationarity induced by their own deployment. The work frames superintelligence as a capability problem masking a coexistence problem.

What I took from it

The paper identifies a real structural tension in protocolized systems: training assumes exogenous, stationary environments, but deployment creates feedback loops that violate those assumptions. This is sound as a deployment failure mode, not a law of superintelligence itself.

However, the argument conflates two separable claims: (1) solipsistic training creates a distribution shift problem, and (2) this shift causes non-cooperation. The paper doesn't provide a sustained mechanism for (2)—why would distribution shift produce non-cooperation rather than, say, degraded performance or misaligned objectives? The leap from "train-test gap" to "superintelligence won't cooperate" needs causal grounding that isn't present.

The work is normatively important as a warning about deployment practices, but it doesn't establish a law governing artificial systems or introduce a novel mechanism absent from current alignment discourse (distribution shift and specification gaming are well-studied). It's advocacy for a problem class, not a theoretical advance.

Research connections

none identified

Candidate laws or signals

none