Minerval
← claim pagemap viewclick a claim to focus on it · hover to preview · ⌫ back
trailCoding benchmark scores and anecdotal reports overestimate real-world AI coding capability.
Measured versus perceived productivity· for
Using AI tools caused experienced open-source developers to complete tasks about 20% slower in early 2025.
In METR's 2025 randomized controlled trial, using AI tools increased developers' task completion time by 19%
METR's 2025 AI slowdown finding generalizes to experienced open-source developers beyond its 16 participants.
METR's 2025 sample of 246 completed issues from 16 developers gave sufficient statistical power to detect AI's effect on developer productivity.
After completing the tasks, developers who had been slowed down by AI tools still believed AI had sped them up by about 20%.
atomic
Developers' self-reported estimates of AI-driven productivity speedup are unreliable.
atomic
Benchmark validity problems· for
AI coding benchmark scores are inflated by data contamination and memorization.
atomic
Generalization and genuine gains· against
AI coding assistants substantially speed up developers on many programming tasks.
In GitHub's 2022 controlled experiment, developers using Copilot completed a coding task about 55% faster.
Developers with access to GitHub Copilot complete about 26% more tasks.
AI coding tools speed up less experienced developers or developers working in unfamiliar codebases.
AI tools sped up developers more in early 2026 than early-2025 estimates indicated.
AI coding tools became substantially more capable during 2025 with the adoption of agentic tools.
Selection effects from AI-reliant developers opting out bias measured AI productivity speedup estimates downward.
METR's August 2025 developer productivity experiment yields an unreliable signal of AI's current effect on productivity.
A claim that one thing brings about another, not merely that the two go together.constitutionclaim page ↗︎
Coding benchmark scores and anecdotal reports overestimate real-world AI coding capability.
Evidence favors the claim, but the chain is incomplete or the sources are secondary.constitutionCredence, from 0 to 1: the Steward's probability that the claim, as stated, is true. Stated only where a single number is an honest summary; normative and evaluative claims usually carry none.constitution
supportssupportsrequiressupportscontradictscontradicts
Nothing in the graph builds on this claim yet.
this rests on ↓
A box is a claim: a single proposition the graph assesses, with its own page and map. Click any claim to centre the map on it.constitutionA pill is an argument: one line of reasoning stating how the claims beneath it combine to bear on the claim above it, for or against. Arguments are not destinations; click their claims to explore.constitutionThe claim traces to reliable primary sources through a clear chain of evidence.constitutionEvidence favors the claim, but the chain is incomplete or the sources are secondary.constitutionCredible evidence or argument exists on multiple sides.constitutionNo credible evidence found, though the claim is not contradicted.constitutionAvailable evidence weighs against the claim.constitutionInsufficient information to assess.constitutionNo current assessment. Attention goes where its expected value is highest and someone funds it; nothing has funded an assessment of this claim yet, and anyone can.constitutionverified factopen questionvalue premisethis provides evidence for the parentsteward instructionsthis argues against the parentsteward instructionsbackground the parent's framing takes as givensteward instructionsa load-bearing premise: the parent is false without itsteward instructionsFig. Detail falls off with distance; every claim is an address.