# NN meta-label on the vol-gated dip-z: cross-index / market-state features (pre-registration) Written 2026-09-23, **before any model was fitted or any result looked at**. The only thing inspected before this was the trade count per year of the gated strategy (2021 91, 2022 238, 2023 43, 2024 60, 2025 119, 2026 102; win rate 62.6 %), which was needed to size the walk-forward blocks. Code: `research/nn_cross_index.py`. ## Universe and label - Trades: the production rule exactly as `vol_filter_test.collect_filtered(..., "vol pct >= 0.50")` produces them (SP500, NAS100, US30, DAX40; H4; z20 <= -1.5; GK sigma(30) expanding percentile >= 0.50; exit close >= SMA20 or 10 bars; stop 3 x ATR14; next-open fill; spread from the bar file). Generator `backtest.simulate` reused unchanged. - Label (classifier): net R-multiple > 0. Secondary label (regressor, sizing only): R. - Primary training universe = the gated trades. Secondary (reported, NOT part of the pass bar): train on the UNGATED dip-z trades (1,243, twice the data), score the gated ones. ## Timing / leakage control All four H4 files sit on the same 00/04/08/12/16/20 server-time grid. For a signal bar of symbol s opening at t (closing t+4h), another index's features are taken from its last bar with open time <= t, which closes at or before t+4h — i.e. already closed when the signal bar closes. When DAX is shut (US evening) this is its last closed bar (stale, causal). Correlation/dispersion use closes forward-filled onto the union grid with the same rule. ## Feature sets (all at the signal bar i = entry_i - 1) Per index k (fixed order SP500, NAS100, US30, DAX40): `z20_k`, `vp_k` (GK sigma30 expanding pctile), `r5_k` = (c-c[-5])/ATR14, `r30_k` = (c-c[-30])/ATR14. **CROSS (primary)** = the 16 per-index columns above + `breadth15` (# indices with z20<=-1.5) + `breadth10` (# with z20<=-1.0) + own `dd120` ((120-bar high - close)/ATR14) + own `slope200` ((SMA200 - SMA200[-20])/ATR14) + `corr60` (mean pairwise correlation of the 4 indices' H4 log returns over the last 60 union-grid bars) + `disp20` (cross-sectional std of the 4 indices' 20-bar log returns) + index one-hot (4). = 26 columns. **BAR** (the old in-terminal set, baseline): z20, (c-SMA200)/ATR, ret1/ATR, ret5/ATR, down streak, position in 20-bar range, range/ATR, close-in-bar, gap/ATR, ATR14/ATR100, efficiency ratio(10), variance ratio (4-bar vs 1-bar over 60), regime (c>SMA200), dow, hour + one-hot. **BOTH** = CROSS ∪ BAR (measures the increment). ## Models (hyper-parameters fixed now, not tuned) - **MLP (primary)**: StandardScaler -> MLPClassifier(hidden=(16,), alpha=1.0, solver=lbfgs, max_iter=2000), mean of 5 seeds. Median imputation from the training block. - RF: 500 trees, min_samples_leaf=15, max_features=sqrt. - GB: GradientBoosting(150 trees, depth 2, lr 0.05, subsample 0.7). - LR: L2 logistic, C=0.1 (linear reference). - Regressor for the R-sizing secondary: MLPRegressor, same architecture. ## Walk-forward Expanding window, **yearly refit** (6-monthly would give blocks of 20-60 trades). Validation blocks: 2023, 2024, 2025, 2026-01..08. Training = every trade whose signal bar is before the block start AND whose exit bar is >= 12 bars (of its own symbol) before the block start (purge + embargo 12 bars). Pooled model with index one-hot. ## Statistics - Per block: AUC and validation n. Mean AUC over blocks with n >= 100 (expected: only 2025 and 2026 qualify — this is stated in advance). - **Pooled OOS AUC** over all 2023-26 validation rows, 95 % bootstrap CI (2,000 resamples) and a label-permutation p-value (1,000). - Sanity: (1) labels shuffled inside every training block -> AUC must be ~0.5; (2) canary: CROSS + `canary` = first-bar return of the trade (c - o of the fill bar)/ATR, i.e. future data -> AUC must be clearly high, proving the harness sees signal. - Permutation importance of the primary model: drop in pooled OOS AUC when a column is shuffled within its block (20 repeats). ## Portfolio (OOS 2023-01-01 .. 2026-08-31, halves split at 2024-11-01) Rebuilt from the per-bar P (every gated signal bar gets a P from its block's model), so a skipped signal lets the symbol take a later signal exactly as the EA would: entries = gated & (P >= cut) -> `backtest.simulate`. Risk 0.25 %/trade, open-risk cap 0.75 % (sum of open stop-risk). **Friday flat and swap are NOT modelled** (the Python tooling does not; it affects all variants alike). - (a) baseline gated. - (b) filter P >= cut, cuts 0.45 / **0.50 (primary)** / 0.55. - (c) sizer: multiplier 0.5 / 1.0 / 1.5 by tercile of P, tercile edges from time-ordered 3-fold out-of-fold predictions on the training block (causal). ret/DD is ~scale-free, so average-risk drift does not bias the comparison; realized mean multiplier is reported. - Secondary: sizer from the MLP regressor's predicted R terciles. Per half: trades, trades/month, return (compounded at exit), maxDD exit-based AND mark-to-market (open positions marked at every H4 close, additive), ret/DD = return / MTM maxDD. ## PASS BAR (primary = MLP on CROSS, gated universe) AUC criterion: pooled OOS AUC >= 0.55 with bootstrap 95 % CI lower bound > 0.50, AND mean per-block AUC over blocks with n >= 100 >= 0.55. - **Filter PASS** = AUC criterion AND cut-0.50 filter ret/DD (MTM) > baseline in BOTH halves AND >= 2 trades/month in both halves. - **Sizer PASS** = AUC criterion AND tercile sizer ret/DD (MTM) > baseline in BOTH halves. Anything else is FAIL. Other models / cuts / the ungated-trained variant are reported for context and cannot turn a FAIL into a PASS.