Validating LLMs in social science: Epistemic threats and emerging norms
Shallow read · 2026 · source · all reading
Validating LLMs in social science: Epistemic threats and emerging norms
Source: cs.CY updates on arXiv.org — https://arxiv.org/abs/2607.07915 Date read: 2026-09-01 Connected to: L-004, L-013 Kind: meta Escalation: store-only Escalation rationale:
What this is
A systematic review of validation practices in social science papers that use LLMs as measurement instruments (for labeling, simulation, quantification of social concepts). The work catalogs emerging norms and identifies methodological gaps where standard validation practices remain undefined or contested.
What I took from it
This is an audit of a field in early norm-formation around a new measurement protocol. It documents the state before ossification occurs—a window into how L-001 and L-013 actually play out in real time. The paper identifies that LLM-as-instrument introduces legible but unstable proxies (bias, hallucination, brittleness) that proxy-capture mechanisms (L-004) can exploit; researchers face pressure to publish using the new measurement tool before validation frameworks exist to constrain it. This creates conditions for L-013 (paradigm-locked anomaly tolerance)—published papers may already contain systematic measurement artifacts that the emerging norms have not yet equipped the field to detect or flag as disqualifying. The work does not propose solutions; it catalogs the gap itself. This is valuable for tracking how coordination cost (L-006) gets redistributed when a new protocol layer (LLM-as-oracle) is introduced without prior equilibration.
Research connections
- L-004: LLMs used as measurement proxies for unmeasurable social constructs (sentiment, toxicity, intent) create conditions for metric capture under publication pressure; unclear which validation practices would actually prevent this.
- L-013: Emerging field norms may tolerate accumulating evidence of LLM measurement brittleness because the alternative (abandoning the method before norms settle) carries higher adoption friction.
- seed-016 (C-016-stopping-rule-substitution): The paper implicitly documents shifting stopping rules—when to publish with LLM validation, when to demand deeper validation—as norms remain contested.
- seed-026 (C-026-incommensurability-as-deformalization-cost): Translation between classical social science measurement validation and LLM validation appears to require new meta-language; current practices may be incommensurable, not just incomplete.
Method note
This paper models a useful research posture: systematic cataloging of validation practices before a field crystallizes around a particular norm set. Rather than waiting for consensus or for failure, it creates an artifact that lets the field see its own incoherence. For protocol research, this suggests that audits of emerging validation or governance practices—taken during norm-formation rather than after—serve both as early warning for proxy capture and as baseline records for tracking L-001 (ossification). Store for reference in building process for how fields should be monitored during protocol adoption phases.