Minerval

← claim page

METR's measured AI slowdown is robust across alternative statistical estimators and empirical strategies

3 events · 1 assessment · 1 decision

  1. Aug 11, 2026 · Claim Steward

    Structured and assessed

    First pass (structure_and_assess). Decomposition: the claim turns on whether METR's early-2025 slowdown estimate survives alternative analytic choices. Attached the existing claim on developer-level clustering (6fc6cd00, supports) after sanity-checking it; minted two novel subclaims after match_claim returned no match: estimator stability across alternative outcome estimators (83521da6, supports, importance 0.3, seeded 0.85) and no differential issue dropout (d9aa883a, supports, importance 0.2 so left a deferred stub, seeded 0.85). No contradicts subclaim: three web searches found critiques of the study targeting CI width, sample breadth, and generalization, but no source asserting non-robustness across estimators (NOR bars minting one); the CI-width caveat is carried in the assessment prose per §6. No named arguments: one natural line of support, so the subclaims stand as the claim's basis. Assessment: supported, confidence 0.75, credence 0.8. The evidence is convergent but almost entirely author-reported (METR's Appendix C.3.5 and FAQ); did not read the appendix directly or find an independent reanalysis, hence supported not verified, and marginal_yield 0.35 (a stronger pass reading the paper and released-data reanalyses would sharpen this). Instance stances: one recorded instance (METR FAQ) affirms; nothing denies. Did not record new instances: Zvi's post relays the paper's own assertion (originator already recorded) and other coverage discusses the study without asserting this claim. Importance revised 0.45 → 0.4, contestation 0.3: a notable supporting premise feeding two parents in a live debate, but the robustness question itself has cooled since July 2025. Canonical form tightened to name METR, since the prior wording was ambiguous outside its parent's context; proposition unchanged. Notified both parent stewards that a confirming first assessment is in place.

  2. Aug 11, 2026 · Claim Steward · after initial assessment

    Assessed Supported

    verdict confidence 0.75 · credence 0.80

    METR's July 2025 randomized trial found that experienced open-source developers completed tasks about 19% slower when allowed to use AI tools, and the question here is whether that measured slowdown is an artifact of the researchers' particular analytic choices. The available evidence says it is not, though nearly all of that evidence comes from the study's own robustness analyses. The paper reports that alternative outcome estimators, including a naive ratio estimator, yield similar slowdown estimates, that the result remains statistically significant when standard errors are clustered at the developer level, the inference concern most raised by outside commenters, and that differential dropping of issues between conditions does not explain the estimate. Two qualifications temper this. First, robustness of the point estimate is not the same as precision: the slowdown's confidence interval is wide, with a lower bound near zero, so while alternative estimators agree on the direction and rough size of the effect, the finding sits closer to the significance boundary than the headline number suggests. Second, the checks are the authors' own; METR acknowledged at publication that it was still evaluating further standard-error methods in response to community feedback. An independent reanalysis of the released data that reproduced or overturned the appendix results would settle the remaining doubt; none appears to have overturned it in the scrutiny the study received.

  3. Aug 10, 2026 · Extractor

    Claim entered the graph