169 行
12 KiB
Markdown
169 行
12 KiB
Markdown
# NN meta-label, cross-index / market-state features: RESULTS
| |||
| |||
Pre-registration: `NN_PLAN.md` (written before any fit). Code: `python research/nn_cross_index.py`
| |||
(~12 min). Everything below section "Verdict" is the script's verbatim output.
| |||
| |||
## Verdict: FAIL
| |||
| |||
The primary model (MLP on the CROSS set, gated universe, yearly walk-forward) has pooled OOS
| |||
AUC 0.564 with a 95% bootstrap CI of [0.498, 0.630] (20,000 resamples). The lower bound
| |||
does not exceed 0.50, so the AUC criterion fails and so do both applications. The per-block
| |||
AUCs are 0.433 / 0.613 / 0.558 / 0.598. None of the other 11 model x feature-set cells has
| |||
a CI that excludes 0.5, and on the 2025 block (n = 118) every cross/bar model except the MLP
| |||
is below 0.5. On its face the cut-0.50 filter improves ret/DD (MTM) in both halves:
| |||
H1 1.68 -> 2.15, H2 3.24 -> 4.15, at 2.8 and 7.7 trades per month. The robustness checks
| |||
(section 7) say this is luck. With a different set of 5 MLP seeds, the same spec gives AUC
| |||
0.529 [0.464, 0.594] and the filter falls to 1.82 / 3.03. Skipping the same 35% of signal
| |||
bars at random (200 draws) beats the baseline in both halves 12% of the time. The primary
| |||
filter's result is 7.5% (H1) / 14% (H2) into the tail of that random distribution, which is
| |||
not distinguishable from chance. The tercile sizer is worse than the baseline in both halves
| |||
(1.66 / 2.92). Training on the ungated trades (twice the data) gives AUC 0.526. The MLP
| |||
regressor's corr(predicted R, realised R) is +0.067. The harness checks behave as expected:
| |||
shuffled labels give 0.504, and the future-return canary gives 0.609 from the model and 0.680
| |||
on its own. So the pipeline can see a real signal, and it does not see one in these features.
| |||
Cross-index / market-state inputs join signal-bar state as "no separable information" at the
| |||
sample size this data allows (324 OOS trades, 2023-26).
| |||
| |||
Caveats: Friday flat and swap are not modelled, as in portfolio.py. MTM marks positions at
| |||
H4 closes. Section 6 of the first run printed a PASS because it drew a second 2,000-resample
| |||
bootstrap that gave CI [0.501, 0.635], while the table from the same run showed [0.498,
| |||
0.630]. The fix (one 20,000-resample CI used by both the table and the verdict) is a
| |||
correction of Monte-Carlo noise, not a change of the bar. Either way, a criterion that flips
| |||
on bootstrap noise is not a pass.
| |||
| |||
---
| |||
| |||
loaded SP500: 8749 bars 2021-01-04..2026-08-31, gated signal bars 574
| |||
loaded NAS100: 8778 bars 2021-01-04..2026-08-31, gated signal bars 618
| |||
loaded US30: 8755 bars 2021-01-04..2026-08-31, gated signal bars 588
| |||
loaded DAX40: 8657 bars 2021-01-04..2026-08-31, gated signal bars 562
| |||
gated trades 653, ungated 1243
| |||
| |||
## 1. Walk-forward AUC (gated universe, label R>0)
| |||
| |||
| model | features | per-block AUC (train n / val n) | mean AUC n>=100 | pooled OOS AUC [95% CI] | perm p |
| |||
|---|---|---|---|---|---|
| |||
| MLP (PRIMARY) | CROSS | 2023 0.433 (326/43); 2024 0.613 (372/61); 2025 0.558 (429/118); 2026 0.598 (551/102) | 0.578 | 0.564 [0.498, 0.630] | 0.033 |
| |||
| RF | CROSS | 2023 0.690 (326/43); 2024 0.672 (372/61); 2025 0.398 (429/118); 2026 0.507 (551/102) | 0.452 | 0.518 [0.453, 0.587] | 0.286 |
| |||
| GB | CROSS | 2023 0.519 (326/43); 2024 0.570 (372/61); 2025 0.425 (429/118); 2026 0.599 (551/102) | 0.512 | 0.523 [0.460, 0.587] | 0.242 |
| |||
| LR | CROSS | 2023 0.692 (326/43); 2024 0.668 (372/61); 2025 0.437 (429/118); 2026 0.514 (551/102) | 0.476 | 0.540 [0.475, 0.607] | 0.114 |
| |||
| MLP | BAR | 2023 0.762 (326/43); 2024 0.559 (372/61); 2025 0.439 (429/118); 2026 0.555 (551/102) | 0.497 | 0.538 [0.474, 0.602] | 0.126 |
| |||
| RF | BAR | 2023 0.729 (326/43); 2024 0.610 (372/61); 2025 0.311 (429/118); 2026 0.499 (551/102) | 0.405 | 0.488 [0.419, 0.552] | 0.611 |
| |||
| GB | BAR | 2023 0.787 (326/43); 2024 0.581 (372/61); 2025 0.350 (429/118); 2026 0.599 (551/102) | 0.474 | 0.532 [0.467, 0.599] | 0.162 |
| |||
| LR | BAR | 2023 0.681 (326/43); 2024 0.714 (372/61); 2025 0.278 (429/118); 2026 0.556 (551/102) | 0.417 | 0.491 [0.422, 0.562] | 0.605 |
| |||
| MLP | BOTH | 2023 0.438 (326/43); 2024 0.550 (372/61); 2025 0.415 (429/118); 2026 0.587 (551/102) | 0.501 | 0.493 [0.426, 0.558] | 0.578 |
| |||
| RF | BOTH | 2023 0.720 (326/43); 2024 0.606 (372/61); 2025 0.331 (429/118); 2026 0.508 (551/102) | 0.420 | 0.494 [0.423, 0.560] | 0.546 |
| |||
| GB | BOTH | 2023 0.704 (326/43); 2024 0.652 (372/61); 2025 0.410 (429/118); 2026 0.608 (551/102) | 0.509 | 0.549 [0.481, 0.617] | 0.079 |
| |||
| LR | BOTH | 2023 0.604 (326/43); 2024 0.688 (372/61); 2025 0.362 (429/118); 2026 0.572 (551/102) | 0.467 | 0.516 [0.452, 0.582] | 0.294 |
| |||
| |||
## 2. Sanity checks (MLP, CROSS)
| |||
| |||
| check | per-block AUC | pooled OOS AUC [95% CI] |
| |||
|---|---|---|
| |||
| labels shuffled in training | 2023 0.583; 2024 0.520; 2025 0.444; 2026 0.518 | 0.504 [0.436, 0.565] |
| |||
| canary: + first-bar return (future) | 2023 0.519; 2024 0.632; 2025 0.579; 2026 0.694 | 0.609 [0.546, 0.672] |
| |||
| |||
## 3. Secondary: MLP trained on the UNGATED dip-z trades, scored on gated (not in pass bar)
| |||
| |||
per-block: 2023 0.664 (472/43); 2024 0.561 (691/61); 2025 0.409 (884/118); 2026 0.582 (1093/102)
| |||
| |||
pooled OOS AUC 0.526 [0.460, 0.591], perm p 0.210
| |||
| |||
## 4. Permutation importance, primary model (drop in pooled OOS AUC, 20 repeats)
| |||
| |||
| feature | AUC drop | sd |
| |||
|---|---|---|
| |||
| disp20 | +0.0319 | 0.0107 |
| |||
| dd120 | +0.0280 | 0.0073 |
| |||
| corr60 | +0.0264 | 0.0147 |
| |||
| r30_NAS100 | +0.0244 | 0.0138 |
| |||
| vp_NAS100 | +0.0214 | 0.0142 |
| |||
| slope200 | +0.0189 | 0.0074 |
| |||
| vp_SP500 | +0.0174 | 0.0092 |
| |||
| breadth10 | +0.0160 | 0.0102 |
| |||
| z20_NAS100 | +0.0144 | 0.0093 |
| |||
| r5_NAS100 | +0.0128 | 0.0098 |
| |||
| |||
## 5. Portfolio, OOS 2023-01 .. 2026-08 (risk 0.25 %, open-risk cap 0.75 %, no Friday flat / swap)
| |||
| |||
| variant | window | n | /mo | mean risk x | return | maxDD exit | maxDD MTM | ret/DD (MTM) |
| |||
|---|---|---|---|---|---|---|---|---|
| |||
| BASELINE gated | H1 2023-01..2024-10 | 83 | 3.8 | 1.00 | +2.64% | 1.50% | 1.57% | 1.68 |
| |||
| BASELINE gated | H2 2024-11..2026-08 | 209 | 9.5 | 1.00 | +7.90% | 2.35% | 2.44% | 3.24 |
| |||
| BASELINE gated | OOS all | 292 | 6.6 | 1.00 | +10.74% | 2.35% | 2.38% | 4.52 |
| |||
| MLP-CROSS (PRIMARY) filter P>=0.45 | H1 2023-01..2024-10 | 68 | 3.1 | 1.00 | +3.24% | 1.41% | 1.43% | 2.27 |
| |||
| MLP-CROSS (PRIMARY) filter P>=0.45 | H2 2024-11..2026-08 | 178 | 8.1 | 1.00 | +7.74% | 2.13% | 2.21% | 3.51 |
| |||
| MLP-CROSS (PRIMARY) filter P>=0.45 | OOS all | 246 | 5.6 | 1.00 | +11.23% | 2.13% | 2.14% | 5.25 |
| |||
| MLP-CROSS (PRIMARY) filter P>=0.50 | H1 2023-01..2024-10 | 62 | 2.8 | 1.00 | +3.07% | 1.41% | 1.43% | 2.15 |
| |||
| MLP-CROSS (PRIMARY) filter P>=0.50 | H2 2024-11..2026-08 | 169 | 7.7 | 1.00 | +8.45% | 1.96% | 2.04% | 4.15 |
| |||
| MLP-CROSS (PRIMARY) filter P>=0.50 | OOS all | 231 | 5.3 | 1.00 | +11.77% | 1.96% | 1.98% | 5.95 |
| |||
| MLP-CROSS (PRIMARY) filter P>=0.55 | H1 2023-01..2024-10 | 59 | 2.7 | 1.00 | +2.70% | 1.41% | 1.43% | 1.89 |
| |||
| MLP-CROSS (PRIMARY) filter P>=0.55 | H2 2024-11..2026-08 | 164 | 7.5 | 1.00 | +8.09% | 2.01% | 2.07% | 3.92 |
| |||
| MLP-CROSS (PRIMARY) filter P>=0.55 | OOS all | 223 | 5.1 | 1.00 | +11.02% | 2.01% | 2.01% | 5.47 |
| |||
| MLP-CROSS (PRIMARY) sizer 0.5/1/1.5 | H1 2023-01..2024-10 | 74 | 3.4 | 1.07 | +2.75% | 1.49% | 1.65% | 1.66 |
| |||
| MLP-CROSS (PRIMARY) sizer 0.5/1/1.5 | H2 2024-11..2026-08 | 194 | 8.8 | 0.98 | +6.56% | 2.15% | 2.24% | 2.92 |
| |||
| MLP-CROSS (PRIMARY) sizer 0.5/1/1.5 | OOS all | 268 | 6.1 | 1.00 | +9.48% | 2.15% | 2.18% | 4.34 |
| |||
| |||
### Context (not in pass bar)
| |||
| |||
| variant | window | n | /mo | mean risk x | return | maxDD exit | maxDD MTM | ret/DD (MTM) |
| |||
|---|---|---|---|---|---|---|---|---|
| |||
| MLP-BAR filter P>=0.50 | H1 2023-01..2024-10 | 73 | 3.3 | 1.00 | +3.84% | 1.49% | 1.55% | 2.47 |
| |||
| MLP-BAR filter P>=0.50 | H2 2024-11..2026-08 | 180 | 8.2 | 1.00 | +7.39% | 2.15% | 2.25% | 3.29 |
| |||
| MLP-BAR filter P>=0.50 | OOS all | 253 | 5.8 | 1.00 | +11.52% | 2.15% | 2.17% | 5.31 |
| |||
| MLP-BOTH filter P>=0.50 | H1 2023-01..2024-10 | 72 | 3.3 | 1.00 | +2.72% | 1.41% | 1.43% | 1.90 |
| |||
| MLP-BOTH filter P>=0.50 | H2 2024-11..2026-08 | 174 | 7.9 | 1.00 | +7.32% | 2.23% | 2.33% | 3.14 |
| |||
| MLP-BOTH filter P>=0.50 | OOS all | 246 | 5.6 | 1.00 | +10.23% | 2.23% | 2.27% | 4.50 |
| |||
| RF-CROSS filter P>=0.50 | H1 2023-01..2024-10 | 74 | 3.4 | 1.00 | +3.00% | 1.36% | 1.38% | 2.17 |
| |||
| RF-CROSS filter P>=0.50 | H2 2024-11..2026-08 | 195 | 8.9 | 1.00 | +7.61% | 2.21% | 2.30% | 3.32 |
| |||
| RF-CROSS filter P>=0.50 | OOS all | 269 | 6.1 | 1.00 | +10.84% | 2.21% | 2.23% | 4.86 |
| |||
| GB-CROSS filter P>=0.50 | H1 2023-01..2024-10 | 71 | 3.2 | 1.00 | +2.25% | 1.36% | 1.38% | 1.63 |
| |||
| GB-CROSS filter P>=0.50 | H2 2024-11..2026-08 | 192 | 8.7 | 1.00 | +6.35% | 2.33% | 2.43% | 2.61 |
| |||
| GB-CROSS filter P>=0.50 | OOS all | 263 | 6.0 | 1.00 | +8.75% | 2.33% | 2.38% | 3.67 |
| |||
| LR-CROSS filter P>=0.50 | H1 2023-01..2024-10 | 71 | 3.2 | 1.00 | +2.79% | 1.41% | 1.43% | 1.95 |
| |||
| LR-CROSS filter P>=0.50 | H2 2024-11..2026-08 | 195 | 8.9 | 1.00 | +8.30% | 2.40% | 2.58% | 3.21 |
| |||
| LR-CROSS filter P>=0.50 | OOS all | 266 | 6.0 | 1.00 | +11.32% | 2.40% | 2.51% | 4.50 |
| |||
| MLP-CROSS ungated-trained filter P>=0.45 | H1 2023-01..2024-10 | 72 | 3.3 | 1.00 | +2.98% | 1.14% | 1.30% | 2.28 |
| |||
| MLP-CROSS ungated-trained filter P>=0.45 | H2 2024-11..2026-08 | 170 | 7.7 | 1.00 | +7.59% | 2.30% | 2.46% | 3.08 |
| |||
| MLP-CROSS ungated-trained filter P>=0.45 | OOS all | 242 | 5.5 | 1.00 | +10.80% | 2.30% | 2.39% | 4.51 |
| |||
| MLP-CROSS ungated-trained filter P>=0.50 | H1 2023-01..2024-10 | 69 | 3.1 | 1.00 | +3.11% | 1.00% | 1.05% | 2.97 |
| |||
| MLP-CROSS ungated-trained filter P>=0.50 | H2 2024-11..2026-08 | 161 | 7.3 | 1.00 | +6.84% | 2.14% | 2.32% | 2.95 |
| |||
| MLP-CROSS ungated-trained filter P>=0.50 | OOS all | 230 | 5.2 | 1.00 | +10.17% | 2.14% | 2.25% | 4.52 |
| |||
| MLP-CROSS ungated-trained filter P>=0.55 | H1 2023-01..2024-10 | 65 | 3.0 | 1.00 | +2.33% | 1.03% | 1.20% | 1.94 |
| |||
| MLP-CROSS ungated-trained filter P>=0.55 | H2 2024-11..2026-08 | 156 | 7.1 | 1.00 | +7.00% | 2.16% | 2.28% | 3.07 |
| |||
| MLP-CROSS ungated-trained filter P>=0.55 | OOS all | 221 | 5.0 | 1.00 | +9.50% | 2.16% | 2.23% | 4.26 |
| |||
| MLP-CROSS ungated-trained sizer 0.5/1/1.5 | H1 2023-01..2024-10 | 84 | 3.8 | 0.87 | +2.44% | 1.16% | 1.27% | 1.93 |
| |||
| MLP-CROSS ungated-trained sizer 0.5/1/1.5 | H2 2024-11..2026-08 | 199 | 9.1 | 0.93 | +7.03% | 2.02% | 2.06% | 3.41 |
| |||
| MLP-CROSS ungated-trained sizer 0.5/1/1.5 | OOS all | 283 | 6.4 | 0.91 | +9.64% | 2.02% | 2.01% | 4.79 |
| |||
| MLPRegressor-R sizer | H1 2023-01..2024-10 | 76 | 3.5 | 1.08 | +2.86% | 1.49% | 1.65% | 1.73 |
| |||
| MLPRegressor-R sizer | H2 2024-11..2026-08 | 190 | 8.6 | 1.05 | +6.92% | 2.47% | 2.53% | 2.73 |
| |||
| MLPRegressor-R sizer | OOS all | 266 | 6.0 | 1.06 | +9.97% | 2.47% | 2.47% | 4.05 |
| |||
| |||
MLPRegressor: OOS corr(predicted R, realised R) = +0.067 (n 324)
| |||
| |||
## 6. Pre-registered verdict
| |||
| |||
- AUC criterion (pooled >= 0.55, CI lo > 0.50, mean n>=100 blocks >= 0.55): pooled 0.564 [0.498, 0.630], mean big-block 0.578 -> FAIL
| |||
- Filter (cut 0.5) ret/DD > baseline in both halves & cadence >= 2/mo: yes -> FILTER FAIL
| |||
- Sizer ret/DD > baseline in both halves: no -> SIZER FAIL
| |||
| |||
## 7. Robustness checks added AFTER seeing section 6 (they cannot create a PASS; they test whether one is real)
| |||
| |||
- canary column alone, OOS rows: univariate AUC 0.680
| |||
- MLP-CROSS, seeds 100-104 instead of 0-4: pooled OOS AUC 0.529 [0.464, 0.594]; blocks 2023 0.421; 2024 0.537; 2025 0.495; 2026 0.607
| |||
| |||
| variant | window | n | /mo | mean risk x | return | maxDD exit | maxDD MTM | ret/DD (MTM) |
| |||
|---|---|---|---|---|---|---|---|---|
| |||
| seeds 100-104 filter P>=0.50 | H1 2023-01..2024-10 | 66 | 3.0 | 1.00 | +2.60% | 1.41% | 1.43% | 1.82 |
| |||
| seeds 100-104 filter P>=0.50 | H2 2024-11..2026-08 | 180 | 8.2 | 1.00 | +7.20% | 2.30% | 2.38% | 3.03 |
| |||
| seeds 100-104 filter P>=0.50 | OOS all | 246 | 5.6 | 1.00 | +9.99% | 2.30% | 2.32% | 4.31 |
| |||
| |||
- random skip of 35.0% of gated signal bars (the primary filter's skip rate), 200 seeds:
| |||
| |||
| window | baseline ret/DD | primary filter ret/DD | random-skip median [5%, 95%] | share of random >= primary | share of random > baseline |
| |||
|---|---|---|---|---|---|
| |||
| H1 2023-01..2024-10 | 1.68 | 2.15 | 1.51 [0.90, 2.24] | 7.5% | 31.0% |
| |||
| H2 2024-11..2026-08 | 3.24 | 4.15 | 3.18 [2.21, 4.79] | 14.0% | 45.0% |
| |||
| OOS all | 4.52 | 5.95 | 4.36 [3.07, 6.26] | 7.5% | 44.5% |
| |||
| |||
- random skips that beat baseline in BOTH halves: 12.0%
|