Research · Falsification as method
Every result is assumed wrong until it survives.
Research maturity is mostly the willingness to kill your own findings. This program generated candidate books above 1.6 Sharpe, then subjected them to leakage, selection, cost, and forward tests — and suspended the contaminated ones rather than cosmetically adjusting them. What survives is a conservative, pre-registered floor and a stronger candidate under honest forward validation.
The gates a candidate must clear
Deflated Sharpe Ratio
Every candidate is deflated against the honest, inherited trial count (~11,300 program-wide looks), not a narrowed after-the-fact denominator. The bar is DSR ≥ 0.90.
Purged CV + embargo
Residual-FF4 labels, a 35-day embargo, and purged cross-validation so overlapping-horizon leakage can't inflate the fit.
Sealed holdout before selection
An era is sealed and blind before any variant is screened on it — the cleanest out-of-sample point the program has.
Cross-regime IC
Information coefficient must stay positive across 2015–21, 2022–23, and 2024–26 — no single-regime survivors.
Pessimistic-cost primary
The pessimistic cost arm is scored as the primary result; a lever must beat a matched-mean-gross constant-exposure control, not a straw baseline.
Pre-registration
Weights and success criteria are fixed and committed before the next anchor's evidence exists, behind a hashed signal-spec fence.
Self-invalidation trail — headlines suspended, not adjusted
| Claimed headline | Why it failed audit | Honest resolution |
|---|---|---|
| Strategy C — 1.20 / 1.25 Sharpe | 3,472-cell grid on one sample; label contamination | corrected to ~0.63 gross / ~0.71 net |
| Grand backtest — 1.434 / 27.95% | future-split market-cap leak; a third of picks existed only through it | suspended — do not quote |
| Clean Arm A NAV — 1.2667 | NAV series did not track its own positions (β 0.48 vs 1.03) | retracted |
| Trial-035 ML — 1.224 | optimistic cost ~100–300 bp/yr | no validated edge vs raw-quality control (t=0.28) |
| Two-sleeve target — 1.30 / 25% | Sharpe leg fails by 0.42 (SE 0.35) | gate FAIL: 0.882 / 25.1% |
The registered fence pins spec metadata (names, polarity, horizons, input families) under a
canonical SHA-256; scratch_* research lives outside the methodology fence — no
registry mutation, no trial-count accounting, no promotion claim — and nothing crosses
without a committed one-page amendment memo. Fence .
Live data freshness — the pipeline behind the record
| Dataset | Latest data | Severity | Rows |
|---|