Warrior_EA/research/VOLNORM_PLAN.md

34 行
2.2 KiB
Markdown

# Tick-volume normalization (Option A) - pre-registration, 2026-09-30
Written before any result. Script: `volnorm.py`. Feed: broker M1 `.hcc`, SP500 / NAS100 / US30 /
DAX40, 2022-01..2026-08 (real volume is 0 on CFDs, so tick volume is all there is).
## The defect being tested
Three places compute relative volume as `bar / mean(last 20 bars)`: `SignalNeural::BuildFeatures`
(3 inputs), `CWarriorVote::MgmtFeatures` (1 input) and `FracDiff::VolumeAt`. On H1/H4 the last 20
bars span several sessions, so the ratio mostly answers "which hour is this?" - the cash open reads
as high volume every day, the overnight bars as low. Plus the feed's tick counts drift up to 20x
between years (AFML_RESULTS section B).
Candidate: `rvol_tod = tickvol / median(same hour-of-day slot, previous 60 sessions)`, causal
(prior slots only), log-transformed.
## Tests and pass criteria
| # | question | metric | pass |
|---|---|---|---|
| 1 | does the old feature encode the clock? | R^2 of log(old rvol) on hour-of-day dummies | reported; expect large |
| 2 | does the new one stop encoding it, and stop drifting? | same R^2; per-year sd of the yearly means of the log feature | new R^2 < 0.05 and year-mean spread smaller than old, all 4 indices |
| 3 | is it still measuring activity? | Spearman(log rvol, next-bar abs return / ATR) | new >= old - 0.02 on every index (must not destroy information) |
| 4 | does it separate paying dips from failing ones? | ungated dip-z events (z20 <= -1.5, exit on close >= SMA20 or 10 bars): mean bp top vs bottom tercile of the feature, stationary bootstrap CI of the difference, pooled over indices | pooled CI excludes 0, same sign in >= 3/4 indices. Compared against old-feature terciles. |
Test 4 is the only edge claim. Tests 1-3 are mechanical and say whether the input is repaired, which
is worth having even if 4 fails (the Wyckoff/NN inputs then stop being confounded by the clock).
## Rules
- No parameter is tuned: window 60 sessions, median, log. One run per bar size (H4 is the book's
timeframe; H1 reported as a robustness read, not a second chance).
- Result is reported whichever way it falls. A failed test 4 does not remove the repair from 1-3.
- Events overlap in time; the bootstrap is block-based (block = 10 events).