Warrior_EA/research/AFML_RESULTS.md
AnimateDread 5a4b36bb61 Add AFML parts A, B, C and HCC history decoder
- Implemented AFML part A for testing the dip-z book against search artifacts, including PBO, DSR, and CPCV metrics.
- Developed AFML part B to generate time and tick bars from M1 broker data, including return distribution statistics.
- Created AFML part C to build a pipeline for dip-z primary analysis, incorporating features and a random forest model for classification.
- Added HCC history decoder to read and process broker M1 `.hcc` files, ensuring proper handling of data structure and integrity.
2026-09-29 20:33:42 -04:00

5 KiB

AFML campaign - results (2026-09-27)

Plan and pass criteria: AFML_PLAN.md (written before any result). Scripts: afml_overfit.py (A), hcc.py + afml_bars.py (B), afml_meta.py (C).

Verdict: 0 of 3 pre-registered tests passed. The book stays as it is. It remains a real edge with a modest Sharpe, and 5.7 years of H4 is too short to prove it significant after the search. No AFML component improves it.

A. Is the book a product of the search? - FAIL on the letter, informative in substance

1,200-config neighbourhood (z threshold x window x max bars x vol gate x stop), 4 indices, H4 2021-01..2026-08, 0.25% risk.

test result criterion
configs with positive total 97.5% -
configs passing the prop screen 341 / 1,200; book ranks 137th on ret/DD -
PBO (CSCV, S=16, 12,870 splits) 0.255; the IS winner's OOS Sharpe is still > 0 in 96.6% of splits FAIL (<= 0.20)
PSR vs 0 (no deflation) 0.976 -
DSR, N = 1,200 / 5,000 0.38 / 0.28 FAIL (>= 0.95)
DSR at effective N (mean pair corr 0.57; eigen-entropy N_eff = 9) 0.75-0.81 at N = 10-20 FAIL
CPCV (10 groups, k=2, purge 5d, embargo 1%): the SELECTION PROCEDURE 9 paths, maxDD 2.2%..11.4%; 3/9 breach 5%; median ret/DD 4.81 FAIL

Reading: the family is profitable almost everywhere. The edge is not a search artifact, but the book's specific parameters are not better than the best of ~10-20 effectively independent tries would be by luck. Weekly Sharpe 0.13 (annualised 0.94) with skew -1.7 and kurtosis 12 cannot clear a deflated bar in 296 weeks. The deep D1 2008-26 test (ret/DD 7.74, different timeframe) is the independent evidence that the family is real. CPCV shows that re-picking the config on ret/DD is dangerous (paths up to 11% DD). Keep the plateau config and do not re-optimise it.

Prop-relevant, fixed book (stationary block bootstrap of weekly returns, 5,000 draws; weekly aggregation understates intra-week DD, so these are floors):

block 4 wk block 12 wk
maxDD median / p90 / p99 over 5.7 yr @0.25% 3.2% / 4.9% / 7.4% 2.8% / 4.3% / 6.0%
P(maxDD > 5%) 9.2% 3.6%
1-yr window, P(+8% before -5%) @0.25 / 0.50 / 1.00% risk 1% / 34% / 68% 0% / 32% / 69%
1-yr window, P(-5% before +8%) @0.25 / 0.50 / 1.00% risk 0% / 7% / 23% 0% / 4% / 20%

At 0.25% the book almost never fails and almost never reaches an 8% target within a year.

B. Bars sampled by activity (tick bars) - FAIL

Broker M1 .hcc 2022-01..2026-09, rebuilt decoder (hcc.py). The feed's tick volume is not stable: SP500 34M ticks in 2022, 10.5M in 2024, 40M in 2026 (8 months); DAX40 258M in 2022 -> 13M in 2023. A fixed tick count would size bars by the broker's plumbing, so the threshold adapts: EWMA(20) of prior days' tick volume divided by H4 bars per day (a logged deviation, and AFML's own remedy). Volume and dollar bars are impossible (CFD real volume = 0). The M1-built time H4 reproduces the exported H4 exactly (ungated 237/237, 242/242, 254/254, 242/241 trades).

Statistics - AFML's claim holds on our data: excess kurtosis 25->8 (SP500), 17->8, 22->12, 23->17; JB 2-10x lower; |z|>4 tails 0.7% -> 0.5%.

Trading - it does not help this edge:

gated book @0.25% n /mo bp t ret/DD
time FULL 366 6.4 42.0 4.78 6.19
tick FULL 332 5.8 13.6 1.40 1.38
time 2022-23 / 2024-26 61.7 / 32.9 7.30 / 3.81
tick 2022-23 / 2024-26 4.4 / 17.9 0.05 / 1.36

Across a 54-config grid tick bars win on ret/DD in only 15 (median 1.80 vs 2.97). Likely reason (not tested separately): the edge is overnight drift plus daily reversal, and both run on the clock. Tick bars fold the night into a few long bars and cut the median hold from 35 h to 23 h.

C. The whole pipeline (FFD + triple barrier + meta-label + uniqueness weights + purged CV) - FAIL

1,183 ungated dip-z events, 4 indices, H4 2021-26, base rate 0.60. FFD d = 0.10 on all four (chosen on pre-2023 data, 503 taps).

  • Purged 5-fold AUC 0.528, 95% CI 0.494..0.564.
  • Walk-forward 6-month refits: 0.525, 0.556, 0.530, 0.571, 0.390, 0.614, 0.477; pooled 0.521. The vol percentile alone scores 0.542.
  • Walk-forward book 2023-01..2026-08: ungated ret/DD 2.60; vol gate 2.89; meta-label filter 0.85; meta-label AND gate 0.46. The model destroys value.

This is the fourth meta-label failure on this primary (after ALGLIB forest/MLP, the cross-index NN, and the volatility forecast). The strongest single input was the 5-bar return / ATR (AUC 0.576 unfitted). It is a lead, not a result, and it is the same "path into the dip" idea that faded before.

What this changes

Nothing in the EA. The deploy config stays at the centre of a profitable plateau. The honest description is now "real family edge, Sharpe ~0.9, not provably better than its neighbours". The 0.25% risk sizing gives low breach risk and slow target progress. That trade-off is a sizing decision for the operator, not a research question.