Static AI exposure scores do not measure what AI labor policy questions actually require
Assessment
Evidence favors the claim, but the chain is incomplete or the sources are secondary.
AI exposure scores rank occupations by how many of their tasks a given AI system could in principle perform. The claim holds that this is not the quantity labor policy questions turn on, and its core premise is conceded even by the scores' producers: exposure measures technical potential, not realized labor-market impact or displacement, while decisions about retraining, safety nets, and regulation depend on realized adoption and its distribution. The scores are also snapshots of a specific model's capabilities at a specific date and built on US occupational data that may transfer poorly elsewhere, yet they circulate in later and non-US policy contexts, often without engaging subsequent methodological improvements.
The credible counter-position does not deny any of this; it holds that the scores nonetheless carry genuine policy-relevant signal. Independently constructed exposure indices broadly agree on which occupations are most exposed, and observed real-world AI usage correlates strongly with exposure ratings, suggesting the scores are a useful first-pass proxy rather than a measurement error. At the same time, aggregate US data have so far shown no relationship between occupational exposure and employment changes since ChatGPT's release, which illustrates the distance between exposure and the outcomes policy must anticipate.
Read strictly, the claim stands: no party maintains that static exposure scores measure realized labor-market effects, and those effects are what policy questions ask about. What remains genuinely debated is how much of the required evidence the scores supply as proxies. Continued validation of exposure ratings against adoption and outcome data, across providers and countries, is what would move that residual dispute.
Full reasoning: the evidence and decisions behind this verdict
The claim originates in "AI Exposure Scores: what they measure, what they miss, and what comes next" (Lund, Euyang, Munyikwa, Fadaee, Cohere Labs, arxiv.org/abs/2606.23633), which traces how the 2023 GPTs-are-GPTs scores (Eloundou et al., arxiv.org/pdf/2303.10130) diffused into government reports, think-tank briefs, and legislative proposals while their temporal, geographic, and ontological limitations compounded. Both recorded instances affirm the claim, but both come from the same research group (the paper and its companion blog post, cohere.com/blog/the-future-of-work-debate-has-an-evidence-problem), so the affirming stance count carries less independent weight than it appears. No source found in this pass asserts the negation as stated.
The verdict rests mainly on the load-bearing premise that task exposure measures technical potential, not realized impact. This is not merely the critics' position: Eloundou et al. themselves defined exposure as capability overlap and cautioned against reading it as displacement prediction. The temporal premise, that scores are snapshots of a specific model's capabilities, is close to definitional given how the indices are constructed. The geographic premise, that US-taxonomy scores transfer poorly to other labor markets, is the least settled of the supporting premises, since international applications use crosswalks whose accuracy loss is not well quantified.
The counter-evidence is real but answers a different question. Anthropic's validation work (www.anthropic.com/research/labor-market-impacts) reports that 97 percent of observed Claude task usage falls in categories the Eloundou ratings deem feasible and that theoretical, reported, and observed exposure correlate across occupations; independent indices also converge (e.g. a 2025 index correlating 0.72 with Eloundou et al., arxiv.org/pdf/2510.13369). This establishes that exposure scores measure a robust construct with proxy value, not that they measure realized, current, local impacts, which is what the claim says policy questions require. Meanwhile aggregate US employment data show no exposure-correlated shifts since ChatGPT's release, underscoring the potential-versus-impact gap, though that finding is itself provisional and could reflect lags rather than measurement failure.
Status is supported rather than verified because the claim is evaluative ("what policy questions actually require" embeds a judgment about the evidential threshold policy needs), several subclaims are not yet independently assessed, and all affirming instances share one origin. It is not contested, because the strongest opposing material defends the scores' usefulness as proxies rather than denying the measurement gap the claim asserts. The verdict would weaken if exposure ratings were shown to robustly predict realized employment and wage outcomes across countries and time (which would collapse the distinction between measuring potential and supplying what policy needs), or if the usage-correlation evidence generalized so well that static scores proved durable stand-ins for realized impact. No credence is given: the claim is a composite evaluative judgment, not a one-number empirical question.
Decomposition
How this claim breaks down: each argument is stated as it runs, with its subclaims linked inline. ↗︎ opens a subclaim; the map shows how they fit together.
Because task exposure measures technical potential rather than realized labor-market impact, while the labor policy questions the scores are cited on turn on realized adoption, displacement, and wage effects, a static exposure score answers a different question than the one policymakers ask. That aggregate US labor data show no relationship between occupational exposure and employment changes since ChatGPT's release illustrates how far the measured quantity can sit from realized outcomes.
The inference is sound: if exposure quantifies capability overlap while policy questions turn on realized effects, the scores answer a different question than the one asked. The argument lives on exposure measuring technical potential rather than realized impact, a premise the scores' own authors endorse. The supporting finding that aggregate US data show no exposure-correlated employment shifts since ChatGPT's release is illustrative rather than necessary, and could reflect adjustment lags without weakening the core premise.
Because exposure scores are snapshots of a specific model's capabilities that date as capabilities advance and scores built on US occupational data transfer poorly to other labor markets, a score fixed to one model, date, and national taxonomy cannot answer policy questions posed later and elsewhere. That policy analyses citing the 2023 GPTs-are-GPTs scores do not engage later methodological improvements shows these limitations compounding in practice rather than being corrected downstream.
The inference goes through for policy questions posed at other times and places, which is most of them, but not for every policy use of a fresh, domestic score. The premise that scores are snapshots of a specific model's capabilities is close to definitional; the argument's weight therefore rests on the less settled premise that US-taxonomy scores transfer poorly to other labor markets and on the observation that policy analyses citing the 2023 scores do not engage later improvements, which shows the gap is realized in practice rather than corrected downstream.
Because task-level exposure ratings reliably estimate which work tasks LLMs could affect, observed real-world AI usage correlates strongly with those exposure ratings, and independently constructed exposure indices broadly agree on which occupations are most exposed, the scores measure a robust, real quantity that tracks where AI actually lands in the economy, and so provide at least part of the evidence labor policy questions require.
The premises are well grounded, particularly that observed AI usage correlates strongly with exposure ratings, but the inference reaches only a weaker conclusion than the claim's negation: convergent, usage-correlated indices show the scores are a valid proxy with real policy value, not that they measure the realized, current, local impacts policy questions ask about. The argument would defeat the claim outright only if the correlation with realized outcomes proved strong and general enough that a static score could stand in for impact evidence, which the validation so far, drawn largely from one provider's usage data, does not establish.
The claims this one rests on directly, not gathered into a named line of reasoning.
- assumesbackground the parent's framing takes as givensteward instructions →The 2023 GPTs-are-GPTs occupational exposure scores are a central empirical input to future-of-work debates ↗︎
Provenance
Where this claim has been said, linked to its canonical form.
The first is structural, between what static exposure scores measure and what policy questions actually require.
Two gaps have widened as a result. The first is structural, between what static exposure scores measure and what policy questions actually require. Taking the diffusion of these scores as a case study, we show how their temporal, geographic, and ontological limitations compound in policy-facing analyses.
But the questions policymakers are asking require more information than static exposure scores alone can provide.
Cohere Labs blog post accompanying the arXiv paper "AI Exposure Scores: what they measure, what they miss, and what comes next", arguing that the GPTs-are-GPTs scores have become a primary policy input while their temporal, geographic, and ontological limitations compound in policy-facing use.
Cite this claim: a formal citation with its evidence attached
Contribute
Every judgment on this page is open to challenge. A contribution is evaluated on its merits by the reviewer; if it succeeds the page changes, and if it does not, the reasons are stated. Either way the exchange becomes part of the claim’s public record.
The attention this claim received was paid for by a funded mandate. Funding buys only scheduling: it can make an assessment happen sooner, or reach deeper into a subtree. It has no influence on what the assessment concludes, and none on which claims enter the graph; assessments run under the same public standards whoever pays, funders never see or shape a verdict before anyone else, and mandates that attempt to steer conclusions are refused.
Created by extractor · Aug 11, 2026. Every judgment on this page is accompanied by a reasoning trace.