Verification · Validation · Falsification · Uncertainty

Kill or validate: every claim needs a way to fail.

BenchEWS separates publication, mathematical specification, correct software computation, and empirical validation. None of these levels substitutes for another.

PUBLISHED THEORY ≠ EMPIRICAL VALIDATION ≠ IMPLEMENTED SOFTWARE

01 / Evidence Path

No signal without a named measurement chain.

  1. 01System and research question
  2. 02Boundary and observables
  3. 03Data and provenance
  4. 04Model and estimator
  5. 05Null and competing models
  6. 06Preregistered criteria
  7. 07Reproducible run
  8. 08Evidence, uncertainty, and abstention

02 / Published Framework

BenchEWS Individual v0.2: ten hard falsification criteria

The published theoretical framework specifies ten explicit conditions under which its temporal-lead hypothesis should be considered unsupported or falsified. It reports no empirical data and is not a validated diagnostic instrument.

02A / Programme-level falsifiability

The positioning paper states how the architecture must fail.

The published v1.0 study narrows the BenchEWS claim to a non-scalar O/F/D/R/V profile and specifies conditions that would weaken, replace, or reject it. It is a positioning, formalization, and falsifiability study—not an empirical validation of a working diagnostic instrument.

Read and cite the positioning study

03 / Temporal Validity

Estimator latency is part of the claim

Window length, smoothing, sampling, and detection delay can create an apparent temporal lead. Observed lead and estimator-latency-corrected lead must therefore be tested separately.

01

Parameter recovery fails

02

Null or competing models explain the data equally well or better

03

Temporal lead disappears after latency correction

04

The result is not robust to plausible windows and preprocessing

05

The persistence measure changes direction or meaning across regimes

06

Signal–response coupling is not identifiable

07

Perturbation recovery is not reproducible

08

Out-of-sample performance does not exceed the preregistered null

09

The result depends on undisclosed settings

10

Independent replication contradicts the central relationship

04 / Decision States

Bounded outcomes instead of artificial certainty

  • VERIFIEDComputation agrees with its specification and reference cases.
  • SUPPORTED IN CASEA hypothesis passes named preregistered tests in a bounded case.
  • FALSIFIED / UNSUPPORTEDA hard criterion fails or a competing model performs better.
  • INDETERMINATEData, measurement quality, or identifiability are insufficient.
  • NOT TESTEDMissing evaluation must not be interpreted as indirect success.

05 / Self-Correction Validation

Earlier only counts when it is also reliable.

Evaluation must cover discrimination, calibration, lead time, incremental information, robustness, cross-domain transfer, and reproducibility.

Δtlead = tbaseline − tBenchEWSΔtlead > 0

A positive lead is not an improvement if false positives, instability, or estimator bias increase.

Inspect falsification and recovery criteria

06 / Software Boundary

Theory, implementation, and evidence remain separate.

Studio 1.4.0 is released. Studio 2.0 remains functionally separate and in release preparation. BenchEWS Individual, Quality-Driven Propagation, and Adaptive Reopening belong to the planned Studio 3.0 research horizon as published theory—not retroactively to Studio 2.0.