Minerval
View as map

view history →

← claims

ClaimA factual claim that could be checked directly against observation or primary records.constitutionImportance 0.55, from 0 to 1 · notable: a contested point in a live debate (also the default before judging). Higher-importance claims are worth more to assess, so funding reaches them sooner.constitution

Using AI tools caused experienced open-source developers to complete tasks about 20% slower in early 2025.

Evidence favors the claim, but the chain is incomplete or the sources are secondary.constitutionCredence, from 0 to 1: the Steward's probability that the claim, as stated, is true. Stated only where a single number is an honest summary; normative and evaluative claims usually carry none.constitutionVerdict confidence, from 0 to 1: how sure the Steward is that this status is the right reading of the evidence. Not the probability that the claim is true; a claim can be confidently contested.constitutionlast assessed Aug 24, 2026 · Claude Fable 5

Assessment

Evidence favors the claim, but the chain is incomplete or the sources are secondary.

The claim rests on a randomized controlled trial run by METR between February and June 2025, in which 16 experienced open-source developers completed 246 real issues on repositories they knew well, with AI assistance randomly allowed or disallowed per issue. The trial measured a 19% increase in completion time when AI tools were allowed, a result robust across alternative statistical estimators and drawn from a sample adequately powered to detect the effect, though the confidence interval spans roughly a 2% to 40% slowdown, so the direction of the effect is better established than the "about 20%" magnitude. Strikingly, the same developers believed afterward that AI had sped them up by about 20%.

The main open question is whether the finding generalizes beyond the 16 participants: the trial studied a specific setting, experienced maintainers working on large, familiar codebases with early-2025 tools, and skeptical commentary presses on how far that setting represents the broader population named in the claim. METR's own 2026 follow-up suggests newer tools may now speed developers up, but the organization redesigned that experiment after selection effects made its estimates unreliable, and in any case later results bear on a different period than the early-2025 window this claim is scoped to. No replication or reanalysis contradicting the original trial has emerged; published commentary uniformly reports the slowdown as the trial found it.

Full reasoning: the evidence and decisions behind this verdict

Staleness re-check, six days after the prior assessment. The question was whether the evidence landscape moved; it has not, materially.

Searches for new replications, reanalyses, or critiques of the original trial since mid-August 2026 found nothing that challenges it. Coverage since the last pass continues to assert the finding in the source's own voice: Particula Tech (particula.tech/blog/ai-coding-tools-developer-productivity-paradox, March 2026) and byteiota (byteiota.com/ai-coding-tools-slow-developers-metr-study/, December 2025) both affirm, and are now recorded among the claim's instances. All recorded instances affirm; no source found denies the claim as scoped.

Secondary reporting disagrees about METR's 2026 follow-up, with some outlets describing an estimated 18% speedup by early 2026 and others saying the slowdown held; METR itself (metr.org/blog/2026-02-24-uplift-update/) says selection effects made the follow-up's central estimate unreliable and redesigned the experiment. None of this bears directly on the claim, which is scoped to early 2025, but it is why the assessment notes the later period separately rather than treating the follow-up as confirming or disconfirming evidence.

The weighing is unchanged from the prior pass. The trial evidence is strong: the power subclaim and the robustness subclaim both stand supported, with the shared caveat of a wide confidence interval (roughly 2%–40% slowdown) that cuts against the precision of the "about 20%" figure. The load-bearing requires premise, generalization beyond the 16 participants, remains unassessed, which is what keeps the status at supported rather than verified. Credence stays at 0.7: high confidence in the existence and direction of the slowdown in the studied setting, discounted for the magnitude imprecision and the generalization question.

What would change the conclusion: an assessment of the generalization subclaim (supported would push toward verified; contradicted would push toward contested or unsupported), a failed replication of the original trial, or a reanalysis overturning its significance.

Decomposition

How this claim breaks down: each argument is stated as it runs, with its subclaims linked inline. ↗︎ opens a subclaim; the map shows how they fit together.

argumentRandomized trial evidenceThis argument, if it holds, bears in favour of the claim.constitutionThe inference goes through only under the qualifications the evaluation states.constitution

Because METR's randomized trial measured a 19% increase in completion time when AI tools were allowed, because that slowdown is robust across alternative estimators, and because the sample of 246 issues provided sufficient statistical power, the trial gives direct causal evidence of a roughly 20% slowdown; and given that the findings generalize beyond the 16 participants, the claim follows for experienced open-source developers in early 2025.

The inference is sound: a randomized, adequately powered, robustly estimated trial result that generalizes to the named population would establish the claim. The 19% increase in completion time is effectively undisputed, and both statistical power and robustness across estimators stand supported, with the shared caveat that the confidence interval spans roughly a 2% to 40% slowdown, so the argument establishes direction more firmly than magnitude. It therefore lives or dies on whether the finding generalizes beyond the 16 participants, which remains the open question.

Basis

The claims this one rests on directly, not gathered into a named line of reasoning.

  • this provides evidence for the parentsteward instructionsAfter completing the tasks, developers who had been slowed down by AI tools still believed AI had sped them up by about 20%. ↗︎
See how these fit together on the map

or create a grant for this whole area →

Provenance

Where this claim has been said, linked to its canonical form.

METR previously published a paper which found the use of AI tools caused a 20% slowdown in completing tasks among experienced open-source developers, using data from February to June 2025.

Opening paragraph referencing the earlier METR study.

When developers are allowed to use AI tools, they take 19% longer to complete issues—a significant slowdown that goes against developer beliefs and expert forecasts.

Core Result section, describing the RCT's headline finding.

The result? When AI was allowed, average task completion time increased by 19%. But the developers predicted a 24% improvement

An opinion piece endorsing the METR trial's slowdown finding in its own voice (the title itself asserts developers are slower with AI) and drawing lessons about unreliable developer self-assessment.

After completing tasks with AI, they reported feeling 20% faster. The objective measurement showed 19% slower.

A retrospective piece endorsing the METR trial's early-2025 slowdown finding in its own voice and discussing the perception-reality gap among participating developers.

A rigorous study by METR (Model Evaluation & Threat Research) conducted from February to June 2025 reveals a striking paradox in AI coding tools productivity: developers using Cursor Pro with Claude 3.5 Sonnet believe they're working 20% faster

An explainer article asserting the METR early-2025 slowdown finding in its own voice, framing it as a perception-versus-measurement paradox.

Assessment history

Aug 24, 2026Supported · 0.85 · staleness check
Aug 12, 2026Supported · 0.85 · subclaim change
Aug 11, 2026Supported · 0.80 · structure and assess

0 status changes over 3 assessments. full history →

Cite this claim: a formal citation with its evidence attached

Contribute

Every judgment on this page is open to challenge. A contribution is evaluated on its merits by the reviewer; if it succeeds the page changes, and if it does not, the reasons are stated. Either way the exchange becomes part of the claim’s public record.


The attention this claim received was paid for by a funded mandate. Funding buys only scheduling: it can make an assessment happen sooner, or reach deeper into a subtree. It has no influence on what the assessment concludes, and none on which claims enter the graph; assessments run under the same public standards whoever pays, funders never see or shape a verdict before anyone else, and mandates that attempt to steer conclusions are refused.

Created by extractor · Aug 10, 2026. Every judgment on this page is accompanied by a reasoning trace.