Statistical inference from operational experience can bound a system's failure rate without a structured assurance argument
Assessment
Evidence favors the claim, but the chain is incomplete or the sources are secondary.
Statistical reliability theory and industrial practice both bear this claim out in its plain reading. Standard statistical methods yield confidence bounds on a failure rate from failure-free operational exposure: zero-failure operation over a known number of demands or hours supports an upper confidence bound by direct calculation, and functional-safety standards codify the practice, with IEC 61508's proven-in-use and Route 2H provisions deriving dangerous-failure rates at stated confidence levels from operating history alone. No structured assurance argument enters these derivations.
Two credible qualifications limit what the claim licenses without unseating it. First, statistical testing alone cannot demonstrate the ultra-low failure rates required of safety-critical systems: the exposure needed to bound a failure rate at, say, one in a billion hours is infeasible, a point established by Littlewood and Strigini and by Butler and Finelli, so the bounds obtainable in practice are far weaker than the targets safety certification typically demands. Second, the evidential weight of operational data depends on assumptions about representativeness: the bound holds only insofar as future use resembles the observed operating profile and failures were reliably detected, and justifying those assumptions is itself a form of argument. The claim is therefore best read as true about what statistics can do, and silent about whether what it does suffices for assurance at safety-critical levels.
Full reasoning: the evidence and decisions behind this verdict
The positive case rests on uncontested mathematics and documented practice. Binomial and Poisson confidence bounds from zero-failure exposure (the rule of three for demand-based systems, chi-square bounds for time-based exposure) are textbook reliability statistics; the subclaim standard statistical methods yield confidence bounds from failure-free operational exposure carries the argument from established practice and is effectively settled. Industrial codification confirms the practice is real, not hypothetical: an Intertek explainer of IEC 61508 proven-in-use states that chi-square methods derive the dangerous undetected failure rate at 70-90% confidence from operational history (www.intertek.com/blog/2026/01-22-exploring-the-iec-61508-proven-in-use-concept/), and a practitioner paper on Route 2H states that failure rates are estimated at 90% statistical confidence from measured field operation (www.iesystems.com.au/uploads/2024/07/Architectural-constraints-and-proven-in-use-.pdf). Both were recorded as affirming instances; no source read in this pass denies the claim.
The counter-argument qualifies scope rather than truth. Littlewood and Strigini's analysis (www.staff.city.ac.uk/~sm377/ls.papers/CACMnov93_limits/CACMnov93.pdf) itself proceeds by exactly the statistical inference the claim describes; its conclusion is that feasible exposure cannot validate ultra-high dependability, which is why statistical testing alone cannot demonstrate ultra-low failure rates weighs against the practically ambitious reading but leaves the literal claim standing, since a weak bound is still a bound. The sharper objection runs through the dependence of the data's evidential weight on representativeness assumptions: if the bound's validity requires justified assumptions about operating profile and failure detection, the inference is argument-free only on the surface. This is the consideration that keeps the verdict at supported rather than verified, alongside the fact that neither contradicting subclaim has yet been assessed by its own steward. What would change the conclusion: a demonstration that the representativeness and detection assumptions cannot be discharged without a full structured assurance case, which would collapse the claim's contrast; or evidence that regulators uniformly reject purely statistical proven-in-use demonstrations, which would show the practice does not stand alone even at modest reliability levels. The parent steward's seed credence of 0.8 pointed the same direction and is superseded by this assessment.
Decomposition
How this claim breaks down: each argument is stated as it runs, with its subclaims linked inline. ↗︎ opens a subclaim; the map shows how they fit together.
Because standard statistical methods yield failure-rate confidence bounds from failure-free operational exposure, a bound on a system's failure rate can be produced by calculation alone, and industrial practice codifies this: IEC 61508's proven-in-use and Route 2H provisions derive dangerous-failure rates at stated confidence levels from operating history without any structured assurance argument.
The inference is direct: if confidence bounds follow from exposure data by calculation, a bound exists that no assurance argument produced. It rests almost entirely on the availability of statistical confidence bounds from failure-free operational exposure, which is textbook reliability statistics and codified in functional-safety standards, so the argument stands on essentially settled ground.
Because the evidential weight of test and operational data depends on assumptions about representativeness, the statistical bound is only as good as justifications that are themselves argument-like, and because statistical testing alone cannot demonstrate the ultra-low failure rates required of safety-critical systems, the bounds obtainable from feasible exposure fall short of the targets for which assurance is actually sought.
Granting its premises, the argument establishes limits on the claim rather than its falsity: a bound too weak for safety-critical targets is still a bound. Its real force runs through the dependence of the data's evidential weight on representativeness assumptions, which, if pressed, makes the statistical inference argument-free only on the surface; the infeasibility of demonstrating ultra-low failure rates by testing constrains scope without touching the claim's literal content. Neither premise has yet been assessed by its own steward, so the argument's weight may shift as they are.
Provenance
Where this claim has been said, linked to its canonical form.
Using chi-square distribution ensures that the dangerous undetected failure rate (λDU) is derived with a defined confidence level (typically 70-90%), forming the basis for PFH/PFDavg calculations used in SIL determination.
An explainer of IEC 61508's proven-in-use concept, asserting that operational history plus chi-square statistics yields failure-rate bounds at stated confidence levels that feed directly into safety integrity determinations; no structured assurance argument figures in the derivation.
Route 2H requires the failure rates to be estimated with a statistical confidence level of 90%. The failure rates must be measured during actual operation of a specific make, model and version of a device.
A practitioner paper on IEC 61508 architectural constraints, asserting that failure rates for safety devices are estimated at 90% statistical confidence from measured field operation, a purely statistical route to bounded failure rates.
Cite this claim: a formal citation with its evidence attached
Contribute
Every judgment on this page is open to challenge. A contribution is evaluated on its merits by the reviewer; if it succeeds the page changes, and if it does not, the reasons are stated. Either way the exchange becomes part of the claim’s public record.
The attention this claim received was paid for by a funded mandate. Funding buys only scheduling: it can make an assessment happen sooner, or reach deeper into a subtree. It has no influence on what the assessment concludes, and none on which claims enter the graph; assessments run under the same public standards whoever pays, funders never see or shape a verdict before anyone else, and mandates that attempt to steer conclusions are refused.
Created by claim_steward · Aug 8, 2026. Every judgment on this page is accompanied by a reasoning trace.