Warrior_EA/research/AFML_RESULTS.md
AnimateDread 5a4b36bb61 Add AFML parts A, B, C and HCC history decoder
- Implemented AFML part A for testing the dip-z book against search artifacts, including PBO, DSR, and CPCV metrics.
- Developed AFML part B to generate time and tick bars from M1 broker data, including return distribution statistics.
- Created AFML part C to build a pipeline for dip-z primary analysis, incorporating features and a random forest model for classification.
- Added HCC history decoder to read and process broker M1 `.hcc` files, ensuring proper handling of data structure and integrity.
2026-09-29 20:33:42 -04:00

95 lines
5 KiB
Markdown

# AFML campaign - results (2026-09-27)
Plan and pass criteria: `AFML_PLAN.md` (written before any result). Scripts:
`afml_overfit.py` (A), `hcc.py` + `afml_bars.py` (B), `afml_meta.py` (C).
**Verdict: 0 of 3 pre-registered tests passed. The book stays as it is. It remains a
real edge with a modest Sharpe, and 5.7 years of H4 is too short to prove it
significant after the search. No AFML component improves it.**
## A. Is the book a product of the search? - FAIL on the letter, informative in substance
1,200-config neighbourhood (z threshold x window x max bars x vol gate x stop),
4 indices, H4 2021-01..2026-08, 0.25% risk.
| test | result | criterion |
|---|---|---|
| configs with positive total | **97.5%** | - |
| configs passing the prop screen | 341 / 1,200; book ranks 137th on ret/DD | - |
| PBO (CSCV, S=16, 12,870 splits) | **0.255**; the IS winner's OOS Sharpe is still > 0 in 96.6% of splits | FAIL (<= 0.20) |
| PSR vs 0 (no deflation) | 0.976 | - |
| DSR, N = 1,200 / 5,000 | 0.38 / 0.28 | FAIL (>= 0.95) |
| DSR at effective N (mean pair corr 0.57; eigen-entropy N_eff = 9) | 0.75-0.81 at N = 10-20 | FAIL |
| CPCV (10 groups, k=2, purge 5d, embargo 1%): the SELECTION PROCEDURE | 9 paths, maxDD 2.2%..11.4%; 3/9 breach 5%; median ret/DD 4.81 | FAIL |
Reading: the family is profitable almost everywhere. The edge is not a search
artifact, but the book's specific parameters are not better than the best of
~10-20 effectively independent tries would be by luck. Weekly Sharpe 0.13
(annualised 0.94) with skew -1.7 and kurtosis 12 cannot clear a deflated bar in 296
weeks. The deep D1 2008-26 test (ret/DD 7.74, different timeframe) is the
independent evidence that the family is real. CPCV shows that *re-picking* the
config on ret/DD is dangerous (paths up to 11% DD). **Keep the plateau config and
do not re-optimise it.**
Prop-relevant, fixed book (stationary block bootstrap of weekly returns, 5,000
draws; weekly aggregation understates intra-week DD, so these are floors):
| | block 4 wk | block 12 wk |
|---|---|---|
| maxDD median / p90 / p99 over 5.7 yr @0.25% | 3.2% / 4.9% / 7.4% | 2.8% / 4.3% / 6.0% |
| P(maxDD > 5%) | 9.2% | 3.6% |
| 1-yr window, P(+8% before -5%) @0.25 / 0.50 / 1.00% risk | 1% / 34% / 68% | 0% / 32% / 69% |
| 1-yr window, P(-5% before +8%) @0.25 / 0.50 / 1.00% risk | 0% / 7% / 23% | 0% / 4% / 20% |
At 0.25% the book almost never fails and almost never reaches an 8% target within a year.
## B. Bars sampled by activity (tick bars) - FAIL
Broker M1 `.hcc` 2022-01..2026-09, rebuilt decoder (`hcc.py`). **The feed's tick
volume is not stable: SP500 34M ticks in 2022, 10.5M in 2024, 40M in 2026 (8 months);
DAX40 258M in 2022 -> 13M in 2023.** A fixed tick count would size bars by the
broker's plumbing, so the threshold adapts: EWMA(20) of prior days' tick volume
divided by H4 bars per day (a logged deviation, and AFML's own remedy). Volume and
dollar bars are impossible (CFD real volume = 0). The M1-built time H4 reproduces
the exported H4 exactly (ungated 237/237, 242/242, 254/254, 242/241 trades).
Statistics - **AFML's claim holds on our data**: excess kurtosis 25->8 (SP500),
17->8, 22->12, 23->17; JB 2-10x lower; |z|>4 tails 0.7% -> 0.5%.
Trading - **it does not help this edge**:
| gated book @0.25% | n | /mo | bp | t | ret/DD |
|---|---|---|---|---|---|
| time FULL | 366 | 6.4 | 42.0 | 4.78 | 6.19 |
| tick FULL | 332 | 5.8 | 13.6 | 1.40 | 1.38 |
| time 2022-23 / 2024-26 | | | 61.7 / 32.9 | | 7.30 / 3.81 |
| tick 2022-23 / 2024-26 | | | 4.4 / 17.9 | | 0.05 / 1.36 |
Across a 54-config grid tick bars win on ret/DD in only 15 (median 1.80 vs 2.97).
Likely reason (not tested separately): the edge is overnight drift plus daily reversal,
and both run on the clock. Tick bars fold the night into a few long bars and cut the
median hold from 35 h to 23 h.
## C. The whole pipeline (FFD + triple barrier + meta-label + uniqueness weights + purged CV) - FAIL
1,183 ungated dip-z events, 4 indices, H4 2021-26, base rate 0.60. FFD d = 0.10 on
all four (chosen on pre-2023 data, 503 taps).
- Purged 5-fold AUC **0.528, 95% CI 0.494..0.564**.
- Walk-forward 6-month refits: 0.525, 0.556, 0.530, 0.571, **0.390**, 0.614, 0.477;
pooled 0.521. The vol percentile alone scores 0.542.
- Walk-forward book 2023-01..2026-08: ungated ret/DD 2.60; **vol gate 2.89**;
meta-label filter 0.85; meta-label AND gate 0.46. The model destroys value.
This is the fourth meta-label failure on this primary (after ALGLIB forest/MLP,
the cross-index NN, and the volatility forecast). The strongest single input was
the 5-bar return / ATR (AUC 0.576 unfitted). It is a lead, not a result, and it is
the same "path into the dip" idea that faded before.
## What this changes
Nothing in the EA. The deploy config stays at the centre of a profitable plateau.
The honest description is now "real family edge, Sharpe ~0.9, not provably
better than its neighbours". The 0.25% risk sizing gives low breach risk and slow
target progress. That trade-off is a sizing decision for the operator, not a
research question.