Access to a generative AI conversational assistant increases customer support agents' productivity by roughly 14-15% on average.
Assessment
Evidence favors the claim, but the chain is incomplete or the sources are secondary.
The claim originates in the study "Generative AI at Work" by Brynjolfsson, Li, and Raymond, which followed the staggered rollout of a generative AI conversational assistant across more than 5,000 customer support agents at a Fortune 500 software firm. The working paper (2023) reported a 14% average increase in issues resolved per hour; the peer-reviewed version, published in The Quarterly Journal of Economics in 2025, revised the headline estimate to 15%. The figure of roughly 14-15% therefore rests on a single large, well-designed field study that has survived peer review at a top economics journal, with no published replication failures or credible methodological rebuttals to date.
The main caveat is generality: the evidence comes from one firm, one tool, and one modality (chat-based support), so the specific magnitude may not carry over to other support organizations or tools. The study also found that the average conceals substantial variation, with gains concentrated among novice and low-skilled agents, a pattern whose generality across task types remains contested. Independent deployments in other support settings finding similar magnitudes would strengthen the claim toward established fact; a failed replication or a credible critique of the study's identification strategy would unsettle it.
Full reasoning: the evidence and decisions behind this verdict
Evidence base. The claim originates in Brynjolfsson, Li and Raymond, "Generative AI at Work" (NBER working paper w31161, April 2023, www.nber.org/papers/w31161; also posted to SSRN, papers.ssrn.com/sol3/papers.cfm?abstract_id=4426942), which reported that access to the assistant increases productivity, measured as issues resolved per hour, by 14% on average among roughly 5,179 customer support agents. The peer-reviewed version (The Quarterly Journal of Economics 140(2): 889-942, May 2025, academic.oup.com/qje/article/140/2/889/7990658) revised the headline estimate to 15%. The canonical form's "roughly 14-15%" covers both published figures for the same finding.
Re-check (this pass). A fresh search for critiques, failed replications, or rebuttals of the estimate again found none; results were confined to the study's own venues (NBER, SSRN, QJE, institutional pages). All three recorded instances affirm the claim and none deny it. Nothing in the evidence landscape has moved since the prior assessment, so the verdict stands unchanged.
Weighing. The design (staggered difference-in-differences over a large agent panel, with the productivity gain decomposed into faster handling, more chats handled, and higher resolution rates) is strong for a field setting, and publication in a top economics journal after two years of review adds weight. The subclaim that gains fall far more to novice and low-skilled workers than to experienced ones stands contested, but only in its general form across task types; its contested status does not undercut the average effect here, which is compatible with any distribution of gains across workers.
Why supported rather than verified. The claim as stated is general ("customer support agents"), while the evidence is one firm, one tool, and one modality (chat support). Within the studied setting the estimate is close to verified; as a general proposition about customer support work, the evidence is a single study, however strong. Credence 0.75 reflects high confidence in the finding itself discounted by uncertainty about whether the specific magnitude generalizes.
What would change the conclusion. A failed replication or a credible methodological critique of the difference-in-differences identification (for example, evidence of differential pre-trends or selective rollout) would move the verdict toward contested. Two or three independent deployments in other support organizations finding similar magnitudes would move it toward verified.
Decomposition
The claims this one rests on directly. ↗︎ opens a subclaim; the map shows how they fit together.
The claims this one rests on directly, not gathered into a named line of reasoning.
- specifiesa more specific version of the parentsteward instructions →Generative AI assistance raises productivity far more for novice and low-skilled workers than for experienced, highly skilled workers. ↗︎
Provenance
Where this claim has been said, linked to its canonical form.
Access to the tool increases productivity, as measured by issues resolved per hour, by 14% on average
we study the staggered introduction of a generative AI-based conversational assistant using data from 5,179 customer support agents.
Access to AI assistance increases worker productivity, as measured by issues resolved per hour, by 15% on average, with substantial heterogeneity across workers.
Peer-reviewed published version of the study (QJE 140(2): 889-942), reporting the headline average productivity effect from the staggered introduction of a generative AI conversational assistant among 5,172 customer-support agents. The published estimate is 15%, revised from the working paper's 14%.
Access to the tool increases productivity, as measured by issues resolved per hour, by 14% on average, including a 34% improvement for novice and low-skilled workers but with minimal impact on experienced and highly skilled workers.
SSRN posting of the working paper studying the staggered introduction of a generative AI conversational assistant among customer support agents; the abstract asserts the 14% average productivity gain.
Assessment history
0 status changes over 2 assessments. full history →
Cite this claim: a formal citation with its evidence attached
Contribute
Every judgment on this page is open to challenge. A contribution is evaluated on its merits by the reviewer; if it succeeds the page changes, and if it does not, the reasons are stated. Either way the exchange becomes part of the claim’s public record.
The attention this claim received was paid for by a funded mandate. Funding buys only scheduling: it can make an assessment happen sooner, or reach deeper into a subtree. It has no influence on what the assessment concludes, and none on which claims enter the graph; assessments run under the same public standards whoever pays, funders never see or shape a verdict before anyone else, and mandates that attempt to steer conclusions are refused.
Created by extractor · Aug 10, 2026. Every judgment on this page is accompanied by a reasoning trace.