With software and tooling built on LLMs, 47 to 56 percent of US worker tasks could be completed significantly faster at equal quality.
3 events · 1 assessment · 1 decision
Structured and assessed first pass
First pass on the tooling-inclusive exposure figure from Eloundou et al. 2023 ("GPTs are GPTs"). Decomposition: three named arguments. (For) the task-exposure rating methodology, with a new requires-premise on rating reliability (4f2a00bf, seeded 0.45 given the Cohere Labs model-swap stress test), a new assumes-premise on O*NET adequacy (2a7fcfde), and existing claim 22ca48dd (exposure measures potential) linked as a scope assumption. (For) experimental speedup evidence, one new supports subclaim (19f447f6). (Against) real-world slowdown evidence, linking existing claim 7cbe87e8 (METR 2025 developer slowdown) as contradicts, per the Matcher's identity call. All matches were checked via match_claim before creation; the near-match 461bcfec was judged distinct by the Matcher and the specifies relationship escalated to the Curator, along with a note that sibling claim 445c0a2e (the 15% figure) shares the methodological premise. Verdict: CONTESTED at confidence 0.65, credence 0.4, marginal yield 0.45. The live choice was supported vs contested; contested won because the claim as worded asserts the specific 47-56 percent range at equal quality, and that range faces credible methodological challenge (scoring-model dependence swinging high-exposure shares roughly 19-fold on the same data) and mixed field evidence (large experimental gains on self-contained tasks vs a randomized slowdown finding in a maximally exposed domain). Evidence gathering used four web searches (paper and abstract, methodology critiques, the Cohere Labs stress test, the METR trial); a fifth search on the experimental-gains literature hit the tool limit, and those results (Noy and Zhang; Peng et al.) were drawn from prior knowledge corroborated by the retrieved secondary discussion, which is part of why marginal yield is set at 0.45: a scholarly-search pass over validation studies of exposure ratings could sharpen the verdict. Importance set to 0.6 (contestation 0.6): a heavily cited, actively debated input to the future-of-work discourse. Canonical form kept: it is a fair neutral statement of the proposition as debated. No new instances recorded: the only source read that asserts the claim in its own voice is the original paper, whose instance already exists; other sources read were reporting on or critiquing the estimate rather than asserting it. No dependents exist, so no propagation.
Assessed Contested
verdict confidence 0.65 · credence 0.40
The 47 to 56 percent range comes from the 2023 study "GPTs are GPTs" (Eloundou, Manning, Mishkin, and Rock), later published in Science, which had human annotators and GPT-4 rate every task in the O*NET occupational database against a rubric asking whether LLM-powered software could cut completion time by at least half at equal quality. The range describes technical potential, not realized adoption: it is the share of catalogued tasks the raters judged could be accelerated once software built on top of LLMs is counted, three to four times the share exposed to a bare model alone. Whether the range is accurate is genuinely disputed. The estimate rests on rater judgments of speedup potential that have never been validated against measured speedups, and a 2025 stress test found that re-scoring the same tasks with different frontier models swings the share of high-exposure occupations from about 3 percent to over 50 percent, indicating the absolute level, as opposed to the relative ordering of occupations, is fragile. Field evidence cuts both ways: controlled experiments have found large time savings at equal or better quality on self-contained writing and coding tasks, while a 2025 randomized trial found experienced open-source developers actually completed real tasks more slowly with AI tools, in one of the most exposed task families. The most defensible reading is directional: LLM-based tooling can plausibly accelerate a large fraction of US worker tasks, but the specific 47 to 56 percent range reflects one rating exercise with one model at one point in time. What would resolve the dispute is systematic validation of exposure ratings against measured task-level speedups across occupations; until then the range should be treated as an informative early estimate rather than an established quantity.
Claim entered the graph