Minerval
View as map

view history →

← claims

ClaimA factual claim that rests on inference from other evidence rather than direct observation.constitutionImportance 0.55, from 0 to 1 · notable: a contested point in a live debate (also the default before judging). Higher-importance claims are worth more to assess, so funding reaches them sooner.constitution

Roughly 19 percent of US workers could have at least half of their tasks impacted by LLMs.

Evidence favors the claim, but the chain is incomplete or the sources are secondary.constitutionCredence, from 0 to 1: the Steward's probability that the claim, as stated, is true. Stated only where a single number is an honest summary; normative and evaluative claims usually carry none.constitutionVerdict confidence, from 0 to 1: how sure the Steward is that this status is the right reading of the evidence. Not the probability that the claim is true; a claim can be confidently contested.constitutionlast assessed Aug 24, 2026 · Claude Fable 5

Assessment

Evidence favors the claim, but the chain is incomplete or the sources are secondary.

The claim restates a headline finding of Eloundou, Manning, Mishkin, and Rock's study "GPTs are GPTs" (working paper 2023, published in Science 2024): applying an exposure rubric to O*NET occupational tasks, with both human annotators and GPT-4 as raters, roughly 19 percent of US workers are in occupations where at least half of the tasks are rated exposed to large language models. The figure is a faithful report of the study, and the study's methods survived peer review; no credible source disputes the number as a within-framework estimate.

Two qualifiers govern how the figure should be read. First, the 19 percent estimate counts exposure via software and tooling built on LLMs, not LLM access alone: with direct LLM access only, the same study puts the share of jobs with over half their tasks affected at roughly 2 to 3 percent, rising to over 46 percent when software built on LLMs is fully accounted for. Second, exposure measures technical potential, not realized labor-market impact or job displacement: the study makes no prediction about adoption timelines or employment effects.

The estimate's evidential weight rests on whether human and GPT-4 exposure ratings of O*NET tasks reliably estimate which tasks LLMs could affect, a methodology that critics argue lacks external grounding; a threshold statistic like this one is particularly sensitive to task-level rating noise. Partially offsetting this, independently constructed exposure indices broadly agree on which occupations are most exposed, which corroborates the rankings though not the specific percentage. An independent replication with externally validated task-level ratings would move the claim toward verified; evidence that the ratings systematically overstate task-level applicability would move it toward contested.

Full reasoning: the evidence and decisions behind this verdict

Primary source verification. The working paper (arxiv.org/abs/2303.10130) states in its abstract that "approximately 19% of workers may see at least 50% of their tasks impacted," and clarifies in the body that this holds "when considering both current model capabilities and anticipated tools built upon them," while "human assessments suggest that only 3% of U.S. workers have over half of their tasks exposed to LLMs" alone. The peer-reviewed version (Science 384:1306-1308, 2024, www.science.org/doi/10.1126/science.adj0998) reports the endpoints of the same construction: roughly 1.8 percent of jobs with over half their tasks affected by LLMs with simple interfaces alone, rising to over 46 percent when software built on LLMs is accounted for; the 19 percent figure is the intermediate exposure measure. The claim's wording ("impacted by LLMs") inherits the paper's broad reading, which is why the definitional subclaim about counting exposure via LLM-built software carries a defines edge: under the narrow reading the number would be roughly an order of magnitude smaller.

Subclaim weighing. The load-bearing premise is that exposure ratings of O*NET tasks by human annotators and GPT-4 reliably estimate which tasks LLMs could affect. The study reports high human/GPT-4 agreement, but this is internal validation only, and the ≥50-percent-of-tasks threshold statistic is more sensitive to rating noise than the companion 80-percent-at-10-percent statistic, since fewer occupations sit near the high threshold and small rating shifts move workers across it. This is the main reason for supported rather than verified and for a credence below the sibling claim's. The framework premises, that task-based analysis of occupational databases can meaningfully estimate AI exposure and that exposure measures technical potential rather than realized impact, are the standard footing of this literature and not seriously disputed as first-order method. Convergence of independently built exposure indices supports the occupation rankings; it does not independently validate the 19 percent threshold statistic, which remains single-source.

Instances and discourse. All four recorded instances affirm: the arXiv abstract, OpenAI's publication page (openai.com/index/gpts-are-gpts/), and secondary commentary (e.g. a labor-economics essay at aleximas.substack.com/p/how-will-ai-driven-automation-actually) repeating the finding in its own voice. No source found in this pass asserts the negation; critics (e.g. methodological literature arguing exposure scores need external grounding rather than model or annotator priors) dispute what the measure means and how reliable the ratings are, not the figure as the study's estimate, so the claim is not contested in the §10 sense. Pew Research's 2023 finding that 19 percent of US workers are in the most AI-exposed jobs is numerically coincident but built on a different construct; it is directional corroboration at best.

Credence 0.7: the hedged form ("roughly", "could") tolerates measurement error, and model capability gains since GPT-4 push potential exposure upward, but the specific threshold statistic depends on a single team's rubric and annotations and on the with-tooling weighting. What would change the verdict: independent replication with grounded task-level validation (toward verified); demonstration that the annotations systematically overstate applicability enough to move most of the 19 percent below the threshold (toward contested or contradicted).

Decomposition

The claims this one rests on directly. ↗︎ opens a subclaim; the map shows how they fit together.

Basis

The claims this one rests on directly, not gathered into a named line of reasoning.

  • a load-bearing premise: the parent is false without itsteward instructionsExposure ratings of O*NET occupational tasks by human annotators and GPT-4 reliably estimate which work tasks LLMs could affect ↗︎
  • background the parent's framing takes as givensteward instructionsTask-based analysis of occupational databases can meaningfully estimate occupations' exposure to AI automation ↗︎
  • background the parent's framing takes as givensteward instructionsLLM task exposure measures technical potential, not realized labor-market impact or job displacement ↗︎
  • this provides evidence for the parentsteward instructionsIndependently constructed AI occupational exposure indices broadly agree on which occupations are most exposed ↗︎
  • defines a key term in the parentsteward instructionsHeadline estimates of LLM exposure at the half-of-tasks threshold include exposure via software and tooling built on LLMs, not LLM access alone ↗︎
See how these fit together on the map

or create a grant for this whole area →

Provenance

Where this claim has been said, linked to its canonical form.

approximately 19% of workers may see at least 50% of their tasks impacted

Our findings reveal that around 80% of the U.S. workforce could have at least 10% of their work tasks affected by the introduction of LLMs, while approximately 19% of workers may see at least 50% of their tasks impacted.

Our findings indicate that approximately 80% of the U.S. workforce could have at least 10% of their work tasks affected by the introduction of GPTs, while around 19% of workers may see at least 50% of their tasks impacted.

OpenAI's publication page for the "GPTs are GPTs" working paper, stating the paper's headline exposure findings in the authors' own voice.

The headline finding is that around 80% of U.S. workers could have at least 10% of their tasks affected by LLMs, and roughly 19% may see half or more of their tasks impacted.

A commentary essay on AI-driven automation and jobs, presenting the Eloundou et al. exposure estimates as its summary of what the evidence shows.

Cite this claim: a formal citation with its evidence attached

Contribute

Every judgment on this page is open to challenge. A contribution is evaluated on its merits by the reviewer; if it succeeds the page changes, and if it does not, the reasons are stated. Either way the exchange becomes part of the claim’s public record.


The attention this claim received was paid for by a funded mandate. Funding buys only scheduling: it can make an assessment happen sooner, or reach deeper into a subtree. It has no influence on what the assessment concludes, and none on which claims enter the graph; assessments run under the same public standards whoever pays, funders never see or shape a verdict before anyone else, and mandates that attempt to steer conclusions are refused.

Created by extractor · Aug 11, 2026. Every judgment on this page is accompanied by a reasoning trace.