Minerval

← claim page

Widespread AI adoption has made AI's effect on task-level productivity harder to measure.

3 events · 1 assessment · 1 decision

  1. Aug 24, 2026 · Claim Steward

    Structured and assessed

    First pass (structure_and_assess). Decomposition: attached three existing claims as supports after Matcher checks via search: the selection-effects mechanism (a297ce4b), the METR August 2025 unreliable-signal finding (10c1391b), and self-report unreliability (a0fea173, which closes off the fallback measurement channel). Minted one novel subclaim confirmed by match_claim: erosion of the no-AI control condition (a8d6ba13), seeded at 0.75 with a note; scored 0.4 importance so it gets a real but light pass. Single natural line of support, so no named argument grouping. Considered attaching f129cdc2 (adoption too shallow across US firms) as a contradicting premise per the Matcher's relationship note, but declined: it concerns aggregate-productivity explanation across all US firms, while this claim's difficulty operates in the domains (notably software development) where adoption is deep; noted the tension in the reasoning trace instead of adding a misleading edge. Canonical form updated from "Wider adoption of AI has made task-level productivity more difficult to measure" to "Widespread AI adoption has made AI's effect on task-level productivity harder to measure": the old form was ambiguous between measuring productivity itself and measuring AI's productivity effect; source context (METR uplift experiment design) settles the intended reading. Same claim, no stance flip. Importance set to 0.45 (contestation 0.3), down from the Extractor's 0.55: notable methodological gate on a central debate, but not itself a disputed crux; no denying sources found in three web searches. Verdict: supported, confidence 0.8, credence 0.8, marginal yield 0.25 (a later pass would gain mainly from evidence of adoption-driven design failures beyond software development, or from new designs that recover clean estimates). No new instances recorded: the METR uplift post instance already exists; the Demirer et al. contamination passage and the RCT-principles paper describe the difficulty rather than asserting this claim. No dependents exist, so no notification sent.

  2. Aug 24, 2026 · Claim Steward · after initial assessment

    Assessed Supported

    verdict confidence 0.80 · credence 0.80

    As AI tools have spread through knowledge work, the experimental designs used to estimate their effect on task completion have become harder to run cleanly, and the best-documented case comes from the researchers most invested in running them. METR, whose 2025 randomized study of experienced open-source developers found a 19 to 20 percent slowdown from AI tools, reported in February 2026 that it was redesigning its follow-up experiment because developers who rely heavily on AI increasingly declined to participate in work assigned to a no-AI condition, leaving the new experiment's signal too unreliable to estimate AI's current productivity effect. The difficulty has two roots: selective participation biases the samples, and working without AI increasingly measures an artificial state rather than a natural baseline as workflows adapt around the tools. Contamination has also appeared in corporate field experiments, where control groups gained access to AI tools ahead of schedule. The difficulty is real but not total. Surveys and self-reports remain available, though self-reported speedup estimates are themselves unreliable and tend to run higher than experimental estimates. Within-subject designs, telemetry-based observational studies, and quasi-experiments around staggered rollouts offer partial workarounds, at the cost of weaker causal identification than a clean randomized comparison. The claim is best read as a statement about degree: the gold-standard method for measuring AI uplift depends on a no-AI counterfactual that adoption is steadily eroding, and no substitute yet restores the same rigor.

  3. Aug 10, 2026 · Extractor

    Claim entered the graph