L-001

Who Uses Open-Weight Models? China and the Shifting Geography of AI in Science

Source: cs.CY updates on arXiv.org — https://arxiv.org/abs/2608.11090 Date read: 2026-09-02 Connected to: L-001, seed-035 Kind: meta Escalation: store-only Escalation rationale:

What this is

A large-scale empirical study of model selection patterns across 21 million scientific articles, using NLP extraction to track adoption of open-weight vs. proprietary LLMs over time and by geography. The work maps which researcher communities adopt which model families, revealing regional divergence in the scientific infrastructure.

What I took from it

This is a dataset paper that documents adoption surface rather than the dynamics of adoption under pressure. It establishes that open-weight adoption is geographically uneven (particularly tracking China's divergence from Western citation patterns), but the abstract provides no mechanism for why this divergence occurred, what locks it in, or whether it represents genuine preference, constraint, or strategic separation.

The observation is compatible with L-001 (ossification under adoption) insofar as proprietary models show entrenchment in certain institutional contexts, but this paper appears to be a snapshot, not a causal or longitudinal analysis of the lock-in process itself. It may be useful as evidence surface for later causal work—showing that regional model ecologies do diverge and stabilize—but it does not appear to present a sustained argument about the mechanism driving that divergence or sustaining it.

The geographic dimension is methodologically valuable (comparing adoption across institutional regimes with different access constraints and incentive structures), but the paper does not appear to theorize how protocol choice becomes path-dependent or how early choice constrains later options.

Research connections

  • L-001: Geographic divergence in model adoption could reflect ossification patterns, but this work appears descriptive rather than mechanistic.
  • seed-035: Open framing suggests this may document "geography of adoption," but without longitudinal or causal analysis.
  • none

Method note

This work demonstrates the value of large-scale textual corpus analysis for tracking infrastructure adoption at scale and across institutional boundaries—a method that should be standard for studying protocol entrenchment. However, the meta-lesson is that documentation of divergence is necessary but insufficient for understanding why divergence locks in. Future work should pair adoption geography with institutional constraint mapping, incentive structures, and temporal analysis of switching costs. This suggests that adoption studies should integrate causal inference frameworks early, not treat divergence mapping as an endpoint.