Roughly 19 percent of US workers could have at least half of their tasks impacted by LLMs.
Assessment
Evidence favors the claim, but the chain is incomplete or the sources are secondary.
The claim restates a headline finding of Eloundou, Manning, Mishkin, and Rock's study "GPTs are GPTs" (working paper 2023, published in Science 2024): applying an exposure rubric to O*NET occupational tasks, with both human annotators and GPT-4 as raters, roughly 19 percent of US workers are in occupations where at least half of the tasks are rated exposed to large language models. The figure is a faithful report of the study, and the study's methods survived peer review; no credible source disputes the number as a within-framework estimate.
Two qualifiers govern how the figure should be read. First, the 19 percent estimate counts exposure via software and tooling built on LLMs, not LLM access alone: with direct LLM access only, the same study puts the share of jobs with over half their tasks affected at roughly 2 to 3 percent, rising to over 46 percent when software built on LLMs is fully accounted for. Second, exposure measures technical potential, not realized labor-market impact or job displacement: the study makes no prediction about adoption timelines or employment effects.
The estimate's evidential weight rests on whether human and GPT-4 exposure ratings of O*NET tasks reliably estimate which tasks LLMs could affect, a methodology that critics argue lacks external grounding; a threshold statistic like this one is particularly sensitive to task-level rating noise. Partially offsetting this, independently constructed exposure indices broadly agree on which occupations are most exposed, which corroborates the rankings though not the specific percentage. An independent replication with externally validated task-level ratings would move the claim toward verified; evidence that the ratings systematically overstate task-level applicability would move it toward contested.
Full reasoning: the evidence and decisions behind this verdict
Primary source verification. The working paper (arxiv.org/abs/2303.10130) states in its abstract that "approximately 19% of workers may see at least 50% of their tasks impacted," and clarifies in the body that this holds "when considering both current model capabilities and anticipated tools built upon them," while "human assessments suggest that only 3% of U.S. workers have over half of their tasks exposed to LLMs" alone. The peer-reviewed version (Science 384:1306-1308, 2024, www.science.org/doi/10.1126/science.adj0998) reports the endpoints of the same construction: roughly 1.8 percent of jobs with over half their tasks affected by LLMs with simple interfaces alone, rising to over 46 percent when software built on LLMs is accounted for; the 19 percent figure is the intermediate exposure measure. The claim's wording ("impacted by LLMs") inherits the paper's broad reading, which is why the definitional subclaim about counting exposure via LLM-built software carries a defines edge: under the narrow reading the number would be roughly an order of magnitude smaller.
Subclaim weighing. The load-bearing premise is that exposure ratings of O*NET tasks by human annotators and GPT-4 reliably estimate which tasks LLMs could affect. The study reports high human/GPT-4 agreement, but this is internal validation only, and the ≥50-percent-of-tasks threshold statistic is more sensitive to rating noise than the companion 80-percent-at-10-percent statistic, since fewer occupations sit near the high threshold and small rating shifts move workers across it. This is the main reason for supported rather than verified and for a credence below the sibling claim's. The framework premises, that task-based analysis of occupational databases can meaningfully estimate AI exposure and that exposure measures technical potential rather than realized impact, are the standard footing of this literature and not seriously disputed as first-order method. Convergence of independently built exposure indices supports the occupation rankings; it does not independently validate the 19 percent threshold statistic, which remains single-source.
Instances and discourse. All four recorded instances affirm: the arXiv abstract, OpenAI's publication page (openai.com/index/gpts-are-gpts/), and secondary commentary (e.g. a labor-economics essay at aleximas.substack.com/p/how-will-ai-driven-automation-actually) repeating the finding in its own voice. No source found in this pass asserts the negation; critics (e.g. methodological literature arguing exposure scores need external grounding rather than model or annotator priors) dispute what the measure means and how reliable the ratings are, not the figure as the study's estimate, so the claim is not contested in the §10 sense. Pew Research's 2023 finding that 19 percent of US workers are in the most AI-exposed jobs is numerically coincident but built on a different construct; it is directional corroboration at best.
Credence 0.7: the hedged form ("roughly", "could") tolerates measurement error, and model capability gains since GPT-4 push potential exposure upward, but the specific threshold statistic depends on a single team's rubric and annotations and on the with-tooling weighting. What would change the verdict: independent replication with grounded task-level validation (toward verified); demonstration that the annotations systematically overstate applicability enough to move most of the 19 percent below the threshold (toward contested or contradicted).
Decomposition
The claims this one rests on directly. ↗︎ opens a subclaim; the map shows how they fit together.
The claims this one rests on directly, not gathered into a named line of reasoning.
- requiresa load-bearing premise: the parent is false without itsteward instructions →Exposure ratings of O*NET occupational tasks by human annotators and GPT-4 reliably estimate which work tasks LLMs could affect ↗︎
- assumesbackground the parent's framing takes as givensteward instructions →Task-based analysis of occupational databases can meaningfully estimate occupations' exposure to AI automation ↗︎
- assumesbackground the parent's framing takes as givensteward instructions →LLM task exposure measures technical potential, not realized labor-market impact or job displacement ↗︎
- supportsthis provides evidence for the parentsteward instructions →Independently constructed AI occupational exposure indices broadly agree on which occupations are most exposed ↗︎
- definesdefines a key term in the parentsteward instructions →Headline estimates of LLM exposure at the half-of-tasks threshold include exposure via software and tooling built on LLMs, not LLM access alone ↗︎
Provenance
Where this claim has been said, linked to its canonical form.
approximately 19% of workers may see at least 50% of their tasks impacted
Our findings reveal that around 80% of the U.S. workforce could have at least 10% of their work tasks affected by the introduction of LLMs, while approximately 19% of workers may see at least 50% of their tasks impacted.
Our findings indicate that approximately 80% of the U.S. workforce could have at least 10% of their work tasks affected by the introduction of GPTs, while around 19% of workers may see at least 50% of their tasks impacted.
OpenAI's publication page for the "GPTs are GPTs" working paper, stating the paper's headline exposure findings in the authors' own voice.
The headline finding is that around 80% of U.S. workers could have at least 10% of their tasks affected by LLMs, and roughly 19% may see half or more of their tasks impacted.
A commentary essay on AI-driven automation and jobs, presenting the Eloundou et al. exposure estimates as its summary of what the evidence shows.
Cite this claim: a formal citation with its evidence attached
Contribute
Every judgment on this page is open to challenge. A contribution is evaluated on its merits by the reviewer; if it succeeds the page changes, and if it does not, the reasons are stated. Either way the exchange becomes part of the claim’s public record.
The attention this claim received was paid for by a funded mandate. Funding buys only scheduling: it can make an assessment happen sooner, or reach deeper into a subtree. It has no influence on what the assessment concludes, and none on which claims enter the graph; assessments run under the same public standards whoever pays, funders never see or shape a verdict before anyone else, and mandates that attempt to steer conclusions are refused.
Created by extractor · Aug 11, 2026. Every judgment on this page is accompanied by a reasoning trace.