Research · Falsification as method

Every result is assumed wrong until it survives.

Research maturity is mostly the willingness to kill your own findings. This program generated candidate books above 1.6 Sharpe, then subjected them to leakage, selection, cost, and forward tests — and suspended the contaminated ones rather than cosmetically adjusting them. What survives is a conservative, pre-registered floor and a stronger candidate under honest forward validation.

~11,300 recorded looks DSR-gated pre-registered

The gates a candidate must clear

Deflated Sharpe Ratio

Every candidate is deflated against the honest, inherited trial count (~11,300 program-wide looks), not a narrowed after-the-fact denominator. The bar is DSR ≥ 0.90.

Purged CV + embargo

Residual-FF4 labels, a 35-day embargo, and purged cross-validation so overlapping-horizon leakage can't inflate the fit.

Sealed holdout before selection

An era is sealed and blind before any variant is screened on it — the cleanest out-of-sample point the program has.

Cross-regime IC

Information coefficient must stay positive across 2015–21, 2022–23, and 2024–26 — no single-regime survivors.

Pessimistic-cost primary

The pessimistic cost arm is scored as the primary result; a lever must beat a matched-mean-gross constant-exposure control, not a straw baseline.

Pre-registration

Weights and success criteria are fixed and committed before the next anchor's evidence exists, behind a hashed signal-spec fence.

Self-invalidation trail — headlines suspended, not adjusted

Claimed headlineWhy it failed auditHonest resolution
Strategy C — 1.20 / 1.25 Sharpe3,472-cell grid on one sample; label contaminationcorrected to ~0.63 gross / ~0.71 net
Grand backtest — 1.434 / 27.95%future-split market-cap leak; a third of picks existed only through itsuspended — do not quote
Clean Arm A NAV — 1.2667NAV series did not track its own positions (β 0.48 vs 1.03)retracted
Trial-035 ML — 1.224optimistic cost ~100–300 bp/yrno validated edge vs raw-quality control (t=0.28)
Two-sleeve target — 1.30 / 25%Sharpe leg fails by 0.42 (SE 0.35)gate FAIL: 0.882 / 25.1%

The registered fence pins spec metadata (names, polarity, horizons, input families) under a canonical SHA-256; scratch_* research lives outside the methodology fence — no registry mutation, no trial-count accounting, no promotion claim — and nothing crosses without a committed one-page amendment memo. Fence .

Live data freshness — the pipeline behind the record

DatasetLatest dataSeverityRows
The honest ceiling is operational, not magical alpha. The 1.30 / 25% target is judged unreachable on the current stack; the credible unlevered envelope is ~1.0–1.15 Sharpe / 17–20% CAGR. Real capital stays no-go until the forward paper evidence clears its own pre-registered protocol.