L-003 L-014

Positioning Generative Artificial Intelligence in STEM Assessment: When to Require, Scaffold, or Restrict Its Use

Source: cs.CY updates on arXiv.org — https://arxiv.org/abs/2608.07475 Date read: 2026-09-02 Connected to: L-003, L-014 Kind: content Escalation: store-only Escalation rationale:

What this is

A prescriptive framework paper applying Evidence-Centered Design (ECD) to GenAI governance in STEM assessment. The work proposes a decision matrix for when to require, scaffold, or restrict GenAI use, grounded in learning validity rather than blanket rules. It is a governance tool paper with domain-specific application rather than a primary theoretical or empirical argument about protocol dynamics.

What I took from it

The paper engages the real tension between enforcement difficulty (prohibition drives underground use) and legitimacy (preparation for GenAI-integrated workflows), but frames this as a pedagogical design problem rather than a protocol system problem. It proposes finer-grained boundary specification via ECD — categorizing learning outcomes and mapping GenAI role to outcome type — which is a legibility move: making the formerly informal decision "should students use this tool here?" into a computable classification.

This illustrates L-003 (Formalization Ratchet) in motion: governance pressure (how to regulate GenAI in assessment) drives formalization of previously implicit norms (when use is legitimate). It also touches L-014 (Strategic Boundary Concentration Under Computable Legality) — as the rule becomes machine-readable and auditable, agents will optimize the boundary itself rather than internalize the intent. The paper does not examine this reversal; it assumes formalization solves the governance problem.

Research connections

  • L-003: Confirms mechanism: governance pressure (enforcement failure of blanket rules, legitimacy pressure from workplace norms) drives toward formal, specified protocols. Does not examine whether formalization preserves intent or creates new optimization targets.

  • L-014: The proposed framework is precisely the kind of "computable legality" structure that attracts strategic boundary optimization — e.g., reframing tasks to fall into "require" categories, or gaming the outcome classification. Not addressed in the paper.

  • seed-082 (Additive Intervention in Overloaded Protocols Preserves Root Pressure): The paper proposes adding a decision layer (the ECD framework) on top of existing assessment infrastructure, rather than restructuring the assessment itself. The underlying tension (outsourcing vs. legit use) likely remains.

Seed

Seed title: Boundary Formalization as Optimization Target Redistribution

Seed type: motif

Seed text: When governance pressure on a protocol (e.g., GenAI use restrictions) is resolved by formalizing the boundary via computable rules (e.g., ECD outcome classification), the optimization surface does not disappear — it relocates from behavior to classification. Agents optimize which category a task falls into, or which outcome it claims to serve, rather than internalizing the boundary itself. The formalization simultaneously makes enforcement cheaper and makes the rule itself a resource to be gamed. This pattern likely generalizes across any domain where governance pressure is met with legibility/formalization as a solution.