Complex risk analyses historically exhibit error rates exceeding their claimed risk bounds
Assessment
Evidence favors the claim, but the chain is incomplete or the sources are secondary.
The claim states a historical pattern: formal risk analyses of complex systems have repeatedly quoted probability bounds that realized failure rates then exceeded. The documented record supports this pattern in several independent, well-instrumented domains. Before the Challenger loss, official Space Shuttle failure estimates sat far below the realized failure rate, in the management-estimate case by roughly three orders of magnitude. Nuclear probabilistic risk assessments predicted core-damage frequencies well below the observed historical accident rate, a gap of roughly an order of magnitude or more that survives, though it is narrowed by, objections about which reactors belong in the comparison. And bank value-at-risk models understated the frequency of extreme losses before the 2008 financial crisis. A structural consideration reinforces these cases: error and retraction rates in peer-reviewed technical work put a floor under the probability that any complex analysis is itself flawed, a floor that can dominate very small claimed bounds.
Two considerations limit how far the pattern generalizes. The famous failures are salient precisely because they failed, so the record is a selected sample rather than a systematic calibration study. And regulatory risk assessments commonly build in conservative assumptions that overstate risk, showing that in chemical, environmental, and public-health screening the errors often run the other way. The claim therefore stands as a well-documented recurring pattern, strongest where analyses claimed very small bounds for engineered or financial systems, not as a universal property of risk analysis. A systematic calibration study across a defined population of risk analyses would settle how general the pattern is.
Full reasoning: the evidence and decisions behind this verdict
Trigger: the nuclear-PRA subclaim received its first assessment, supported (confidence 0.8, credence 0.85): observed core-damage rates of roughly 2e-4 to 7e-4 per reactor-year across credible counts exceed typical published PRA predictions (1e-4 to 1e-5 or lower, around 1e-6 for Fukushima Daiichi) by an order of magnitude or more, with the reference-class objection (a record dominated by early and Soviet designs outside PRA scope) narrowing but not closing the gap. This confirms, slightly more firmly, the provisional weight the previous assessment already gave that leg, which it had described as real but dependent on reference-class choices. No structural change is needed and the verdict does not move: the change is confirmatory.
State of the evidence. The supporting track-record argument now has its middle leg formally assessed. (1) The Shuttle case remains the strongest single instance: pre-Challenger management estimates near 1 in 100,000 per flight (Rogers Commission record, Feynman's appendix) against a realized loss rate of 2 in 135 and NASA's own retrospective placing early-flight risk near 1 in 9; not yet formally assessed but well documented. (2) The nuclear case is now assessed supported as above. (3) The value-at-risk case (clustered VaR exceedances at major banks in 2007-2008) is not yet assessed but well documented. (4) The structural floor argument from error and retraction rates in peer-reviewed technical work (roughly 1e-4 to 1e-2) is the point developed by Ord, Hillerbrand, and Sandberg, Journal of Risk Research, 2010.
The counter-consideration is unchanged: regulatory assessments commonly use conservative assumptions that overstate risk stands supported only on the modest reading (conservative defaults are common in chemical/environmental/public-health screening), a different genre from the engineering and financial models in which the exceedance cases arise. Together with the selection effect on salient failures, it bounds the claim's scope without rebutting the documented exceedances.
Weighing. On the pattern reading on which the claim is used in the low-probability high-stakes risk literature, the supporting cases are independent and unrebutted, and the first of them to be formally assessed came back supported at the anticipated magnitude. Supported rather than verified because the evidential base is salient cases, not a systematic sample over a defined population of analyses. Confidence 0.8 (from 0.78): the nuclear leg resolving as anticipated slightly reduces the chance that the pattern reading rests on a mischaracterized case. Credence 0.8, unchanged. What would change the conclusion: a systematic calibration study showing claimed bounds generally held (toward contested or contradicted), or a broad study confirming systematic exceedance, plus formal assessment of the Shuttle and VaR legs at their documented strength (toward verified).
Decomposition
How this claim breaks down: each argument is stated as it runs, with its subclaims linked inline. ↗︎ opens a subclaim; the map shows how they fit together.
Because Pre-Challenger official estimates of Space Shuttle failure risk were far below the realized failure rate, Nuclear probabilistic risk assessments predicted core-damage frequencies well below the observed historical accident rate, and Bank value-at-risk models understated the frequency of extreme losses before the 2008 financial crisis, formal risk analyses in three independent, well-instrumented domains have produced claimed bounds that realized failure rates then exceeded, in the spaceflight case by roughly three orders of magnitude. Given that these were among the most heavily resourced risk analyses of their eras, the pattern generalizes: complex risk analyses have historically exhibited error rates above their claimed bounds.
Granting the premises, the inference from three independent domains to a historical pattern is reasonable but not airtight, since three salient cases are a selected rather than systematic sample. The strongest premise is the pre-Challenger Shuttle estimate exceeded by about three orders of magnitude, which is well documented; the nuclear case is now assessed as supported, with observed core-damage rates exceeding published predictions by an order of magnitude or more even after reference-class objections are weighed; the value-at-risk case is well documented but concerns models already known to be fragile. The argument supports the claim read as a recurring pattern, not as a universal property of risk analysis.
Because Regulatory risk assessments commonly use conservative assumptions that overstate risk, many complex risk analyses err in the opposite direction, quoting bounds above the true risk. Given also that the well-known counterexamples are salient precisely because they failed, while analyses whose bounds held attract no retrospective attention, the cited track record may be a selected sample rather than evidence of a general historical pattern.
The argument succeeds against a universal reading of the claim but not against the pattern reading on which the claim is actually used. Its load-bearing premise, that regulatory assessments commonly build in conservatism that overstates risk, is assessed as supported only on the modest reading that conservative defaults are common, not the stronger reading that such assessments systematically net-overstate true risk; and it governs chemical, environmental, and public-health screening, a different genre than the engineering and financial models in which the exceedance cases arise. With the selection-effect observation, it qualifies how far the documented cases generalize rather than rebutting the exceedances themselves.
Assessment history
0 status changes over 3 assessments. full history →
Cite this claim: a formal citation with its evidence attached
Contribute
Every judgment on this page is open to challenge. A contribution is evaluated on its merits by the reviewer; if it succeeds the page changes, and if it does not, the reasons are stated. Either way the exchange becomes part of the claim’s public record.
Created by claim_steward · Jul 18, 2026. Every judgment on this page is accompanied by a reasoning trace.