Warrior_EA/docs/Wyckoff/WYCKOFF_V21_EVALUATION.md

149 lines
12 KiB
Markdown

# Wyckoff indicators v2.1 / v5.7 — four proposed refinements, measured
Date 2026-09-30. Covers `ADWyckoffEventStream` (5.6 → 5.7), `ADWyckoffSignificantBarInversion`,
`ADWyckoffFailedStructure`, `ADShorteningOfThrust` (2.0 → 2.1), MQL5 and SQX Java.
**Method in one paragraph.** Every proposal was built as an input whose default reproduces the previous build
bit-for-bit, in MQL5 and in Java. The compiled MQL5 was then run in the MT5 Strategy Tester on four regimes and every
buffer dumped next to its bars; the real SQX Java blocks were replayed over the identical bars and diffed; and each
change was judged on whether the signals it removes are worse than the ones it keeps. Nothing below was inferred
from a Python re-implementation: both sides are the shipped code. Raw output is in
[v21_evaluation_data/](v21_evaluation_data/); the tooling is [tests/mt5/README.md](../../tests/mt5/README.md).
## Verdicts
| # | Proposal | What was built | Measured | Verdict |
| --- | --- | --- | --- | --- |
| 4 | **SOT: closed-bar confirmation** | `ADShorteningOfThrust.mq5` evaluates closed bars only, forming bar = 0, replay once per new bar. Optional `ConfirmMode=1`: a swing reverses on a CLOSE, not a wick | The v2.0 forming-bar reading disagreed with the finished bar on **7.1 % / 5.1 % / 3.4 %** of bars with the bar 1/4, 2/4, 3/4 formed — nearly as many wrong readings as there are signals (the finished bar carries one on 8.8 % of bars). `ConfirmMode=1` only lowers that to 4.9 / 4.6 / 2.7 % | **Closed-bar: adopted, always on.** `ConfirmMode=1`: opt-in, off |
| 1 | **EventStream volatility filter** | `SpikeATR` (0 = off), `SpikeScope` | 2–4 % of bars are >3 ATR spikes; at 3 ATR **0.2–0.7 %** of event codes change; the spring/LPS/SOS entries removed/added (42/36 pooled) show no edge either way | Opt-in, off |
| 2 | **SBI time-of-day relative volume** | `TodMode` 0/1/2, `TodDays` | Mode 1 flattens the hour-of-day concentration (top-hour share 0.17→0.10, 0.20→0.09, 0.11→0.09, 0.13→0.07); forward edge unchanged (ON−OFF z ≤ 0.6). Mode 2 discards ~45 % of signals for no gain | Opt-in, off. If used, use mode 1 |
| 3 | **FailedStructure location filter** | `LocationATR` (0 = off) | Removes 21 % (1 ATR) – 48 % (0.5 ATR) of failures; the removed ones were **not worse** than the kept ones (kept−removed z between −1.1 and +0.4) | Opt-in, off |
**Why "off" for three of four.** On these data the unfiltered signals themselves carry almost no stand-alone
forward edge (−0.05…+0.09 ATR, standard error 0.02–0.15), so a filter has nothing to sharpen, and no filter moved
the edge by more than the noise. That is *absence of demonstrated benefit*, not proof of harm: these outputs are
conditions meant to be combined inside SQX strategies, where a stand-alone edge is a weak test. The new inputs are
`@Parameter`s with Builder ranges, so the Builder's own robustness pipeline can judge them in context; the defaults
just keep every existing strategy and every golden file unchanged.
## What the audit found before changing anything
* **Only SOT ever evaluated the forming bar.** EventStream, SBI and FailedStructure already stopped at
`rates_total-2` and skip recomputing when no bar has closed. SOT replayed the *whole history on every tick* and
published bar 0, whose high/low keep moving — and the swing threshold (`max(0.5·ATR, 0.3·range)`) moves with them.
* **No lookahead anywhere.** Every array the four `OnCalculate` bodies read (`time, open, high, low, close,
tick_volume`) is `ArraySetAsSeries(true)`; every buffer is series; `volume`/`spread` are never read.
The SOT ZigZag is a causal state machine — a pivot is stamped on the bar that *detects* it, never back-dated.
* **"Ignore re-anchoring on >3 ATR bars" cannot be applied literally.** The book's selling/buying climax is *by
definition* the widest, heaviest bar of the move, so filtering every >3 ATR bar out of anchoring deletes the
events the indicator exists to find (scope 1 measures this: 0.4–1.6 % of event codes change, including climaxes
and their ARs). The default scope therefore only stops a spike from *stretching* structure — the forming range's
provisional extreme (which becomes the Creek/Ice at the AR) and the Phase B boundary.
* **SBI's volume test is deliberately ordinal** (v2.0 removed the 1.2×/1.5× multipliers). The RVOL change keeps it
ordinal — relative volume must beat the two preceding bars' relative volumes — so no threshold is reintroduced.
* **FailedStructure's structure edges already are balance edges**; the location filter asks a stricter thing
(proximity to the window's volume-profile POC / value-area high / low, or VWAP). The profile is 32 bins across
the window, each bar's tick volume spread over the bins its range touches, value area 70 %, grown from the POC
toward the heavier neighbour (upper on ties).
## Data and metric
| Set | Instrument | Bars | Evaluated from |
| --- | --- | --- | --- |
| H1a | SP500 H1 | 7,206 | 2018.01.01 (5,734 bars: the Feb-2018 volatility spike and the Q4 sell-off) |
| H1b | SP500 H1 | 15,807 | 2025.09.05 (5,894 bars) |
| M15 | SP500 M15 | 63,178 | 2025.09.05 (23,543 bars) |
| EUR | EURUSD H1 | 23,104 | 2008.01.08 (16,877 bars: the 2008–10 crisis) |
The earlier bars of each set warm the indicators up. **Edge** = mean forward return in the signal's direction minus
the unconditional forward return, in units of the signal bar's ATR(14), at 4/12/24 bars; the standard error ignores
overlap between neighbouring signals, so it is optimistic — read |z| < 2 as "no evidence". A change is credited only
if the signals it *removes* carry less edge than the ones it *keeps*. Multiple horizons × changes were examined; the
single |z| > 2 (SOT `ConfirmMode=1` kept-vs-removed, h4: +2.16) does not replicate at h12/h24 (+0.63/−0.88) and is
treated as noise.
## Java ↔ MQL5 parity
The MQL5 is the compiled indicator run by the Strategy Tester (`tests/mt5/ADWykProbe.mq5`); the Java is the real
block class driven bar by bar (`tests/java/wyckoff`, test doubles for `IndicatorBlock`/`ChartData`/`DataSeries`/
`SQTime`, everything else from SQX's own jars). Tolerance 1e-6, forming bar excluded, SOT's 30-bar
`PLOT_DRAW_BEGIN` blanking excluded.
| Indicator | Configurations compared | SP500 H1a / H1b / M15 | EURUSD H1 |
| --- | --- | --- | --- |
| SBI | TodMode 0, 1, 2 | **0 / 0 / 0 differing bars** | 7 bars (0.03 %) |
| FailedStructure | LocationATR 0, 0.5, 1, 2 | **0 / 0 / 0** | 0 |
| SOT | ConfirmMode 0, 1 | **0 / 0 / 0** | 5 bars (0.02 %) |
| EventStream | Spike (0), (3, scope 0), (4, scope 0), (3, scope 1) | **0 / 0 / 0** | 111 bars (0.48 %), 9 of them event codes |
The v2.1 build with every new input at its default is **identical to the v2.0 build on every closed bar** of 63,178
M15 bars, all four indicators, in one tester pass (`regression_and_truncation.txt`).
### Finding outside the brief: FX ties are decided by floating-point noise
The EURUSD residue is **pre-existing** (the logic is untouched) and is not a Java bug. The book's rules are strict
ordinal comparisons, and at five digits ties are common: a bar range of exactly 12 pips against another of exactly
12 pips, or a close at exactly ⅔ of the range. MT5 history stores some quotes one ulp off the nearest double
(`1.4652000000000003`), and the two platforms then resolve the same mathematical tie in opposite directions:
* 2008-07-07 06:00, range `0.0012999999999998568` vs prev `0.0012999999999998568` — a true tie; MT5 calls the bar
wider, Java (strict IEEE) does not.
* 2008-01-15 00:00, close at exactly ⅔ of the range, `cp = 0.6666666666665022 < 0.6666666666666666` in IEEE — Java
says "not in the upper third", MT5 says it is.
Neither side is consistently right; the tie is simply noise. **Snapping prices with `NormalizeDouble` was tried and
made it worse** (51 vs 5 differing SBI quality readings on EURUSD), because `NormalizeDouble` does not return the same double as parsing
the decimal. The clean fix is tolerance-based comparison helpers (`GT/GE/LT/LE` with a relative epsilon of about
1e-9, which is far below any real tick) applied to every ordinal test — a semantic change across several thousand lines that
was **not** made here. SP500, which the SQX golden files use, is exact. Worth doing before trusting bar-exact
MT5↔SQX agreement on forex.
## No lookahead, no repaint
* **Truncation invariance (MQL5).** Each indicator run to 2026.01.31 and to 2026.09.04: **12,285 shared closed bars
× 4 indicators × 2 settings (defaults and options on) — 0 differences.** A future-data leak or a value that
changes when later bars arrive would show here. (`regression_and_truncation.txt`, section B.)
* **Forming bar.** All four MQL5 indicators publish 0 on bar 0.
* **Old SOT behaviour, quantified** (`sot_forming_bar_repaint.txt`): 411 sampled H1 bars, block replayed with the
last bar 1/4, 2/4, 3/4 formed. 1/4 formed: 29 differences, of which 12 were signals shown and gone by the close
and 8 absent until the close (the rest changed magnitude).
## Benchmarks
Java, 63,178 M15 bars, median of 5 JVM runs, defaults, HEAD vs new: SBI 148 → 146 ms, FailedStructure 195 → 195 ms,
SOT 154 → 157 ms, EventStream 351 → 373 ms (2.3–5.9 µs per bar; the difference is noise). Allocation identical
(1.0 MB, 1.0 MB, 7.6 MB, 21.8 MB). Options on cost more only when they fire: FailedStructure `LocationATR` builds a
profile per reported failure (+30–60 % run time at 0.5 ATR, still 0.27 s for 63k bars).
MT5, same bars, first full pass: SBI 31–78 ms, FailedStructure 31–109 ms, SOT 32–78 ms, EventStream 310–390 ms
(timer granularity 15.6 ms); v2.0 and v2.1 are within noise of each other. Terminal memory grew 13 MB (SBI, FS),
8–10 MB (SOT), 40 MB (EventStream) per 63k bars, an upper bound that includes the probe's own copies.
**The real MT5 saving is per tick, not per pass.** v2.0 SOT re-ran the whole-history pass on *every tick*; v2.1
returns immediately until a bar closes. At ~70 ms per pass on 63k bars (5–10 ms on a normal chart) that is the
difference between a busy terminal and an idle one during fast markets.
## What changed
| Area | Files |
| --- | --- |
| MQL5 | `ADWyckoffEventStream.mq5` (5.7), `ADWyckoffSignificantBarInversion.mq5`, `ADWyckoffFailedStructure.mq5`, `ADShorteningOfThrust.mq5` (2.1) — new inputs appended **last** so positional `iCustom` calls keep working |
| Java | the four blocks + `ADWyckoffCommon.PivotParams.closeConfirm` (default false) |
| SQX → MT5 code export | `user/extend/Code/MetaTrader5/blocks/{ADWyckoffEventStream,ADWyckoffSignificantBarInversion,ADWyckoffFailedStructure,ADShorteningOfThrust}.tpl` now forward the new inputs (otherwise a strategy tuned with them would silently run with defaults in MT5) |
| Indicator Tester | `tests/WyckoffTesterConfig.xml` rows and `scripts/update_wyckoff_tester_config.py` carry the new trailing parameters at their defaults |
| Tooling | `tests/mt5/*`, `tests/java/wyckoff/*`, `scripts/wyckoff_probe_lib.py`, `eval_wyckoff_changes.py`, `wyckoff_regression_report.py`, `wyckoff_repaint_test.py` |
## Not done — needs the operator
* **Install sync.** SQX compiles the *installation's* snippet tree, which is a hand copy: the four blocks and
`ADWyckoffCommon.java` must be copied into `SQX_144_2953_win_20260601\user\extend\Snippets`, and the tree SQX will really
compile checked with `bash scripts/check_wyckoff_compile.sh --install` (the repo-tree check can pass while the
install fails).
The install's `IndicatorTester.xml` is a separate file; `scripts/sync_wyckoff_tester_rows.py` splices the repo rows
in, with SQX closed. The MT5 `.tpl` changes are under `user/extend/Code`, also a copy.
* **SQX Indicator Tester goldens** were not regenerated: with defaults nothing they contain changes, except that
SOT's forming bar is now 0 in MT5.
* `scripts/generate_ad_mt5_tpls.py` still describes the pre-v5 input lists (it was already stale); the `.tpl` files
were edited directly and the generator was left alone. The MetaTrader4 / JForex / PseudoCode / EasyLanguage
`.tpl` for these blocks were already out of step with the Java and were not touched.
* The tolerance-comparison hardening described above.