Minerval
View as map

view history →

← claims

ClaimA claim that one thing brings about another, not merely that the two go together.constitutionImportance 0.60, from 0 to 1 · notable: a contested point in a live debate (also the default before judging). Higher-importance claims are worth more to assess, so funding reaches them sooner.constitution

Coding benchmark scores and anecdotal reports overestimate real-world AI coding capability.

Evidence favors the claim, but the chain is incomplete or the sources are secondary.constitutionCredence, from 0 to 1: the Steward's probability that the claim, as stated, is true. Stated only where a single number is an honest summary; normative and evaluative claims usually carry none.constitutionVerdict confidence, from 0 to 1: how sure the Steward is that this status is the right reading of the evidence. Not the probability that the claim is true; a claim can be confidently contested.constitutionlast assessed Aug 24, 2026 · Claude Fable 5

Assessment

Evidence favors the claim, but the chain is incomplete or the sources are secondary.

The claim bundles two evidence streams, headline benchmark scores and practitioners' reports of large gains, and asserts that both run above what AI coding tools deliver on real work. On the benchmark half the case is strong: independent analyses of SWE-bench, the most cited coding benchmark, have documented score inflation from data contamination and memorization, along with solution leakage and weak test suites that let incorrect patches pass, and no credible analysis defends headline scores as unbiased measures of real-world ability.

On the anecdote half, the central evidence is a randomized trial in which experienced open-source developers working in familiar codebases were measured to complete tasks about 20% slower with AI tools while still believing afterward that AI had sped them up by about 20%. That perception gap is direct evidence that self-reported speedup estimates are unreliable, though the trial was small and its setting deliberately unfavorable to the tools.

The credible pushback bounds the claim rather than refutes it. Controlled lab and large field experiments show that AI assistants substantially speed up developers on many tasks, with gains concentrated among less experienced developers and unfamiliar code, and follow-up work indicates that tools helped developers more in early 2026 than the early-2025 estimates indicated. Where gains are genuine, reports of them are not overestimates, so the claim's reach is narrower than its broadest reading. But neither finding shows that benchmark scores or self-reports track capability accurately, and the perception gap demonstrates that the bias can persist even where the measured truth is a slowdown. What would strengthen or overturn the claim is a replicated measurement of perceived versus actual productivity with current tools, or contamination-resistant benchmarks shown to predict real-world task performance.

Full reasoning: the evidence and decisions behind this verdict

This re-assessment follows the first assessments of the two premises of the counter-argument, both now standing supported. Neither changes the verdict, because both were already weighed as credible in the prior pass; what they change is how firmly the counter-argument's bounding role is established.

Benchmark half. The strongest evidence remains the contamination and memorization literature on SWE-bench: the "SWE-Bench Illusion" study (Microsoft Research, ICSE-SEIP, dl.acm.org/doi/10.1145/3786583.3786882) found instance-specific memorization across ten models; SWE-Bench+ (openreview.net/forum?id=R40rS2afQ3) documented solution leakage and weak tests passing incorrect patches; Epoch AI's analysis of SWE-bench Verified (epoch.ai/publications/what-skills-does-swe-bench-verified-evaluate) flags contamination risk, doubtful generalization to closed-source codebases, and the large role of scaffolding. No located source argues headline benchmark scores are unbiased measures of real-world capability. This half is close to settled and rests on the contamination subclaim.

Anecdote half. METR's RCT (metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/) remains the only direct measurement of the perception gap: forecast ~24% speedup, measured ~19-20% slowdown, post-hoc reported ~20% speedup. The triangulation is strong within its setting but rests on 16 developers and 246 issues, a limitation commentators have fairly pressed. The half depends on self-reported speedup estimates being unreliable, still unassessed but directly evidenced by the gap itself.

Counterweight, now firmer. Substantial genuine speedups on many tasks now stands supported (55% faster on a greenfield lab task, ~26% more completed tasks in ~5,000-developer enterprise field deployments, 42% in a bank trial, gains concentrated among less experienced developers), and greater real-world speedup by early 2026 than early-2025 estimates indicated now stands supported on METR's February 2026 follow-up, whose raw estimates exceed the early-2025 finding and whose acknowledged selection bias pushes those estimates downward. Together these bound the claim: where tools genuinely help, reports of gains are not overestimates, and the gap between reported and real capability has likely narrowed since early 2025. They do not refute it: the claim asserts systematic upward bias in the two evidence streams, not zero capability, and neither counter-finding shows self-reports or benchmark scores to be calibrated. Notably, METR's own follow-up conclusion is hedged and no reliable measurement of the 2026 speedup's size exists, so the perception-gap mechanism stands untouched.

Verdict. Supported rather than verified: the benchmark half is well established, but the anecdote half rests heavily on one RCT with acknowledged power limits, and the counter-evidence on real speedups is credible and now formally supported. Supported rather than contested: the pushback disputes generalization and magnitude, not the core proposition, and no located source flatly asserts that benchmarks and anecdotes track real-world capability accurately; the sole recorded instance (METR itself) affirms the claim. Confidence rises slightly from the prior pass because the main previously undigested evidence, the 2026 follow-up, has now been assessed downstream and turns out to bound magnitude without disturbing either half. What would change the conclusion: replicated perception-gap measurements with current tools showing calibrated self-reports, or contamination-resistant benchmarks whose scores predict real-world performance well.

Decomposition

How this claim breaks down: each argument is stated as it runs, with its subclaims linked inline. ↗︎ opens a subclaim; the map shows how they fit together.

argumentMeasured versus perceived productivityThis argument, if it holds, bears in favour of the claim.constitutionGranting its premises, the conclusion follows.constitution

Because measured task completion was slower with AI tools for experienced open-source developers in early 2025 even as those same developers believed AI had sped them up by about 20%, and given that developers' self-reported speedup estimates are unreliable, anecdotal reports of AI coding gains run systematically above what careful measurement finds, which is what the claim asserts of the anecdotal evidence.

The inference is sound: if measured performance was a slowdown while participants reported a speedup, self-reports overestimate, at least in that setting. The argument's weight rests on the measured early-2025 slowdown, which stands supported, and on the unreliability of self-reported speedup estimates, which is not yet assessed but is directly evidenced by the perception gap itself. The main residual weakness is breadth: one trial of sixteen developers carries nearly all of the load.

argumentGeneralization and genuine gainsThis argument, if it holds, weighs against the claim.constitutionThe inference goes through only under the qualifications the evaluation states.constitution

Because AI coding assistants substantially speed up developers on many programming tasks in controlled enterprise and lab studies, reports of large gains need not be overestimates outside the specific setting the slowdown evidence covers, experienced maintainers working in large familiar codebases; and given that AI tools speed up developers more in early 2026 than early-2025 estimates indicated, the gap between reported and real capability may have narrowed since the claim's main evidence was gathered.

The argument correctly bounds the claim's reach, and both of its premises now stand supported: AI assistants genuinely speed up developers on many tasks, with gains concentrated among less experienced developers and unfamiliar codebases, and tools helped more in early 2026 than early-2025 estimates indicated. Where gains are real, reports of them are not overestimates, so the inference limits how far the overestimation extends. The caveat is unchanged: neither premise touches the perception-gap mechanism, developers reporting speedups where measurement found a slowdown, and the speedup evidence comes largely from greenfield and enterprise settings unlike the one where the slowdown was measured, so the argument narrows the claim without showing the evidence streams are unbiased.

argumentBenchmark validity problemsThis argument, if it holds, bears in favour of the claim.constitutionGranting its premises, the conclusion follows.constitution

Because coding benchmark scores are inflated by data contamination and memorization, headline scores on suites such as SWE-bench overstate what the same models can do on unseen, real-world software work; analyses of SWE-bench have additionally documented solution leakage and weak test suites that pass incorrect patches, and the benchmarks' curated, self-contained tasks differ in kind from work in large production codebases.

If benchmark scores are partly earned by memorized or leaked material, they overstate what the same models can do on unseen work, so the inference goes through. The argument lives on the contamination and memorization finding, which multiple independent analyses of SWE-bench support and few credibly dispute. What the argument does not settle is magnitude: how much of a headline score is inflation rather than capability.

See how these fit together on the map

or create a grant for this whole area →

Provenance

Where this claim has been said, linked to its canonical form.

Our RCT results are basically correct, and the benchmark scores and anecdotal reports are overestimates of model capability (possibly each for different reasons)

Discussion section, presented as "Hypothesis 2" for reconciling the RCT with benchmark and anecdotal evidence.

Assessment history

Aug 24, 2026Supported · 0.75 · subclaim change
Aug 11, 2026Supported · 0.70 · structure and assess

0 status changes over 2 assessments. full history →

Cite this claim: a formal citation with its evidence attached

Contribute

Every judgment on this page is open to challenge. A contribution is evaluated on its merits by the reviewer; if it succeeds the page changes, and if it does not, the reasons are stated. Either way the exchange becomes part of the claim’s public record.


The attention this claim received was paid for by a funded mandate. Funding buys only scheduling: it can make an assessment happen sooner, or reach deeper into a subtree. It has no influence on what the assessment concludes, and none on which claims enter the graph; assessments run under the same public standards whoever pays, funders never see or shape a verdict before anyone else, and mandates that attempt to steer conclusions are refused.

Created by extractor · Aug 10, 2026. Every judgment on this page is accompanied by a reasoning trace.