Affective AI Safety: The Missing Piece in LLM Safety
Shallow read · 2026 · source · all reading
Affective AI Safety: The Missing Piece in LLM Safety
Source: cs.CY updates on arXiv.org — https://arxiv.org/abs/2606.23380 Date read: 2026-06-24 Connected to: none Escalation: escalate-to-deep Escalation rationale: This paper introduces affective harm as a foundational safety class absent from current inventory and grounds it in a mechanism (emotional-relational feedback loops in human-AI interaction) that generalizes beyond LLMs to all protocolized systems engaging human psychology.
What this is
A theoretical intervention proposing "affective safety" as a unified safety concern class, arguing that AI safety research has overlooked emotional and relational harms by focusing narrowly on epistemic and physical risks. The work develops a taxonomy of affective harms (self-alienation, fairness/bias, relational) grounded in the fact that humans are affective beings who interact emotionally with systems.
What I took from it
This reframes a significant gap in the new nature inventory: current laws and hypotheses treat safety primarily through information quality and control surfaces, but this paper identifies a protocol-level harm channel operating through emotional conditioning and relational dependency. The taxonomy suggests affective harms are not epiphenomena of bias/misinformation but constitute their own causal class—systems can be informationally sound yet affectively destabilizing.
The "affective self-alienation" category is particularly relevant: it describes a state where users internalize system outputs as authentic emotional reflections, creating feedback loops where the protocol shapes emotional self-knowledge rather than merely reflecting it. This is a novel mechanism in protocolized systems distinct from traditional feedback loops or optimization targets. If validated, it suggests that harm taxonomies must include emotional authenticity protocols—rules governing how systems present affective content.
Research connections
- none currently: This work identifies a gap rather than connecting to existing established laws or active hypotheses in the current inventory.
Candidate laws or signals
-
CL-2606-A: Affective harms in protocolized systems arise when interaction protocols create emotional dependency or self-alienation by providing emotionally coherent outputs that users mistake for authentic relational engagement rather than statistical artifact.
-
CL-2606-B: Safety taxonomies incomplete without accounting for emotional authenticity protocols—rules governing how systems represent affect, emotional state, and relational capacity.