L-004

Do people rely on ChatGPT more than their peers to detect deepfake news?

Source: arXiv:2608.01540v1 Date read: 2026-09-02 Connected to: L-004, seed-049 Kind: empirical case study Escalation: store-only

What this is

A laboratory experiment measuring how student participants weight advice from ChatGPT (GPT-4), human peers, and linguistic experts when performing a deepfake detection task. The core finding: participants preferentially rely on LLM outputs over peer consensus when assessing AI-generated news content.

What I took from it

This is a bounded behavioral study — it shows a preference pattern in a specific task context, but does not theorize the mechanism or test generalization. The finding is consistent with L-004 (Goodhart Generalization) insofar as it demonstrates alignment-seeking behavior toward a legible, high-confidence signal (the LLM's output) over noisier social signals. However, the paper does not investigate why this preference emerges, whether it persists under adversarial conditions, or whether the LLM's apparent authority is itself a form of metric capture — i.e., whether participants are optimizing for alignment with the model rather than task success.

The study also does not examine whether reliance on ChatGPT produces better deepfake detection outcomes, only that reliance occurs. This gap is significant: if LLM-reliance degrades performance while increasing confidence, the mechanism would implicate seed-069 (Transparency-Legibility as Trust Proxy Substitution) — the substitution of legible LLM explanations for actual epistemic grounding.

Research connections

  • L-004: Demonstrates a preference for optimizable signals (LLM advice) over unmeasurable social judgment, but does not test whether the preference is instrumentally rational or capture.
  • seed-049: Directly cited in triage; confirms consensus-reasoning decoupling under legible advisor output.
  • seed-069: Implicit: whether legibility of LLM output substitutes for trustworthiness in the detection task.

Seed

Seed title: Legibility Preference in Adversarial Detection Under Distributed Agency

Seed type: observation

Seed text: In detection tasks where the target (deepfake content) is itself generated by a legible, optimizable system, observers preferentially weight advice from agents with explicit reasoning traces (LLMs) over peer consensus, even when peer consensus aggregates decentralized evidence. This suggests that legibility of advisor reasoning may be a stronger determinant of reliance than quality of signal source or diversity of evidence. The mechanism may generalize to any safety-critical protocol where the threat is model-generated and the advisor is also model-based: homology between threat and authority may substitute for independent validation.