# Volatility Meta-Label Pipeline — Architectural Review & Execution Blueprint Reviewed tree: `MQL5/Shared Projects/Warrior_EA` (fleet terminal `10CE948A…`), 51,445 lines. Date: 2026-09-19. --- ## 0. BLOCKER — READ FIRST: the source tree is 17 days stale and the newer work has no source The working copy at `Documents/Workspaces/Warrior_EA` is **empty**. The only Warrior source on this machine is `MQL5/Shared Projects/Warrior_EA`, and **every file in it is frozen at 2026-09-02 15:33**. The deployed binary the tester actually runs — `MQL5/Experts/Warrior/Warrior_EA.ex5` — is dated **2026-09-13 09:13**. The source that produced it is gone. Confirmed absent from disk, machine-wide: | Module | Campaign it belongs to | |---|---| | `Signals/SignalDipBuy.mqh` | the dip-buy edge (the only surviving edge) | | `System/DipMeta.mqh` | the ALGLIB meta-label | | `Expert/WarriorExpert.mqh` | the session-aware `Refresh()` fix | | `WarriorJournal.mqh` | the journal + its lookback-leak fix | Corroborating evidence that this tree predates the Sep-11 work: `System/TradeChecks.mqh` contains the `TC*` helpers but **no `TCCanOpen()`** — the single entry gate added on Sep 11. A machine-wide search found no archive, zip, or backup. The `.ex5` is compressed, so no symbol names can be recovered from it (a control search for `Warrior` in the binary also returned nothing, so that is inconclusive rather than proof of absence). **Consequence for this plan.** Sections 1–3 below are correct against the Sep-2 architecture and every line reference is real. But the dip-buy signal, the meta-label scaffolding, and the session-aware refresh — the three things this pipeline would attach to — are not in the tree I can read. **Recover or reconstruct the Sep-13 source before executing Phase 3.** Phases 1 and 2 are safe to start now; they touch files that do exist. --- ## 1. Architectural review — where the premise and the code differ Three corrections, because they change what the work actually is. ### 1a. There is no min-max layer, and no sigmoid time encoding `Expert/Features/FeatureBuilder.mqh` (1,357 lines) is the only feature path. What it actually emits: | Feature group | Transform | Line | |---|---|---| | bar geometry | `(close-open)/atr`, `(high-open)/atr`, `(low-open)/atr` | 997–999 | | trend position | `donchPos20/50` — rank in range, bounded `[-1,1]` | 1053–1054 | | displacement | `(close - close[20])/atr`, clamped ±10 | 1059 | | mean extension | `(close - SMA20)/atr`, clamped ±10 | 1063 | | leg state | `dir*(close - legStart)/atr` | 1086 | | volume | `vol/volBase`, absorption, `vol×range` — all ratios | 1119–1129 | | time | **cyclical `sin`/`cos`** on hour, day-of-week, month | 1134–1147 | | ATR | `atr/close`, not raw ATR | 1148+ | There is no `MinMax`, no `Normalize`, and no sigmoid anywhere in the tree (grepped). Time is already cyclically encoded, which is the correct choice — a sigmoid on hour would break the 23:00→00:00 wrap. ### 1b. The real defect is the opposite of the stated one: these inputs are OVER-differenced Every column above is an **integer-order difference, d = 1**, ATR-scaled. That is stationary — and it is memory-less. Measured on a synthetic log-price random walk (`research/fracdiff.py`, validated below): ``` d taps obs adf_t corr-to-level 0.0 1 6000 -4.362 1.0000 <- raw level: all memory, fails ADF in general 0.3 2275 3726 -6.947 -0.0106 0.5 927 5074 -17.799 0.0020 1.0 2 5999 -53.934 -0.0004 <- WHAT THE EA FEEDS TODAY: memory ~= 0 ``` So the network is handed a stationary series with the price level scrubbed out of it. The one partial exception is `smaExtension` (`close − SMA20`), which is a crude low-order memory term — and notably it is one of the four columns the fleet keep-screen voted 24/24 to retain (comment at FeatureBuilder.mqh:1018). That is weak corroboration that retained level information is what the model was missing. **FDF is therefore the right patch — but framed as recovering memory, not as fixing non-stationarity.** ### 1c. The majority-class trap is structural, not a tuning artifact `Expert/AIBase/Inference.mqh:147-153`: ```cpp ENUM_SIGNAL CExpertSignalAIBase::Argmax3(double pBuy, double pSell, double pNeutral) { if(pBuy > pSell && pBuy > pNeutral) return Buy; if(pSell > pBuy && pSell > pNeutral) return Sell; return Neutral; // also the fallback on any tie } ``` Neutral is both the modal class **and** the tiebreak. No loss weighting fixes a tiebreak. A 2-class expansion/chop head removes the bucket entirely. Note the label is **already** a magnitude question, not a direction one — `LegRideLabel()` (Labels.mqh:253) asks "does riding this leg pay ≥ `LEG_LABEL_MIN_RIDE_ATR`", then signs the answer with the leg's own direction. The proposed target keeps the magnitude question, puts it on a fixed clock, and drops the sign. **This is a smaller change than it appears.** --- ## 2. Violations of the stationarity / time-horizon rules found in the tree | # | Where | Finding | Severity | |---|---|---|---| | V1 | FeatureBuilder.mqh 997–1186 | Every price column is `d=1` ATR-normalised → correlation to level ≈ 0. Memory destroyed. | High | | V2 | Inference.mqh 147 | `Argmax3` ties resolve to Neutral — majority class is also the tiebreak. | High | | V3 | Labels.mqh 253 + LegState.mqh 60 | Label horizon is **event-driven** (leg flip), capped at `LEG_LABEL_MAX_RIDE_BARS = 200`. On H4 that is ~33 days. Nothing constrains it to 48–72h, and nothing constrains it to inside the week. | High | | V4 | ExpertCustom.mqh ~878 | Friday liquidation fires only if a tick lands inside a **±1 minute** window (`MathAbs(nowMinOfDay - targetMinOfDay) <= 1`). A thin Friday close or a shut CFD session means no tick, no liquidation, position carried over the weekend. | **Critical** | | V5 | ExpertCustom.mqh ~888 | Schedule is evaluated in `TimeCurrent()` = broker server time. US and EU DST switch on different dates, so any NY-referenced instant drifts by an hour twice a year. | Medium | | V6 | tree-wide | **No swap logic exists at all.** `DEAL_SWAP` is read in the journal (TradeJournalManager.mqh:214) for P&L only. Nothing anywhere reads `SYMBOL_SWAP_*`. | Medium | | V7 | Variables/ConfidenceBridge.mqh 15–29 | Confidence is computed, published (`PublishAIVote`) and journalled, but **gates nothing** — the confidence-scaled management was removed 2026-08-25. There is a bus with no consumer. | Info — this is the hook | V4 is the one to fix regardless of whether the rest of this plan proceeds. --- ## 3. Execution blueprint ### Phase 0 — Recover the source (blocking for Phase 3) 1. Check the FLEET terminal's MetaEditor recent-files and any VCS/Dropbox history for the Sep-13 tree. 2. If unrecoverable, reconstruct `SignalDipBuy.mqh` from the recorded spec: D1, long-only, `DIP_ZSCORE` entry, inputs `DipEntry/DipZ/DipExitMA/DipTrendMA/DipMaxBars`, run alone. 3. Re-establish a source-of-truth location that is **not** only inside the terminal directory. ### Phase 1 — Fractional differentiation (`research/fracdiff.py`, `mql5_patches/FracDiff.mqh`) 1. **Pick `d` per symbol on real data**, not on one series. Run `min_ffd()` over each fleet instrument and over each era separately. Take the **smallest `d` that passes ADF in every era**, not the pooled minimum — this codebase has already been burned by pattern quality inverting across eras. 2. **Respect the tap budget.** Measured widths: | d | tau=1e-5 | tau=1e-4 | tau=1e-3 | |---|---|---|---| | 0.1 | 4076 | 503 | 62 | | 0.3 | 2275 | 388 | 66 | | 0.5 | 927 | 200 | 44 | **The stdlib series wrappers return 0.0 in silence past shift 1023.** At `tau=1e-5`, every `d ≤ 0.45` exceeds that ceiling and would multiply real weights by silent zeros. `FracDiff.mqh` therefore calls `CopyClose`/`CopyTickVolume` **directly** and never touches `m_Close.GetData()`. Default `tau = 1e-4` keeps all `d ≥ 0.1` under ~500 taps. 3. **Add FDF columns alongside the existing ones first, do not replace.** Extend the `names[]`/`widths[]` table at FeatureBuilder.mqh:441 with `ffd_close`, `ffd_volume`. Let the existing fleet keep-screen vote on them, exactly as it voted the ZigZag geometry columns out. Replace `d=1` columns only for those the screen actually drops. 4. Raise `HistoryBars` warm-up by `CFracDiff::WarmupBars()` — the first `taps` bars have no valid output and must not be emitted as zeros. ### Phase 2 — The volatility target (`research/vol_target.py`, `mql5_patches/VolEstimators.mqh`, `VolMetaLabel.mqh`) 1. **Run both targets in parallel and compare**, per the brief: - (a) Garman-Klass as a continuous regression target, - (b) the binary 48–72h expansion label. Score both with walk-forward refits. **Require the refit series, never a first fit** — a first-fit AUC of 0.68–0.77 on ~4 months has already been proven meaningless here. 2. **Use Yang-Zhang, not GK, for anything that gates a multi-day hold.** GK assumes the bar opened where the last one closed; index CFDs gap across the daily maintenance break and the weekend, so GK understates the variance the position is actually exposed to. GK stays as the efficient intra-bar estimator and as a feature. Both are implemented. 3. **Wire `PathBars()` into the existing overlap machinery** (`CLabelOverlap`, `EffectiveSampleSize()`, `PurgeBars()` in Labels.mqh). A 72h label on H4 spans 18 bars; a raw N overstates significance by ≈ √18 ≈ 4.2×. 4. **Budget for the clip.** Measured on synthetic H4 with a Friday 16:00 NY flat: **43% of triggers are dropped** for not fitting a 48h horizon inside the week (78 of 85 drops). These are dropped as `VOL_UNRESOLVED`, never labelled chop — labelling them 0 would manufacture a class that correlates with day-of-week, which the network already reads through its sin/cos features. It would learn the calendar. 5. Change the head from 3-class softmax to 2-class; `Argmax3` and its Neutral tiebreak go away. ### Phase 3 — The meta-label gate (`mql5_patches/VolMetaLabel.mqh`, `SwapWindow.mqh`) 1. **Do not add the gate as a voting signal module.** `CExpertSignal` has no per-side veto, and returning `EMPTY_VALUE` silences the entire ensemble rather than one side. A meta-label is a veto, so it belongs on the entry path as one gate — the same shape as `TCCanOpen()`. Hook: `CExpertCustom::Processing()` immediately before the `CheckOpenLong()/CheckOpenShort()` call (ExpertCustom.mqh:528). 2. `CVolMetaGate::Allowed()` blocks when `P(expansion) < 0.70`. **An unavailable reading also blocks** — a filter that degrades to "allow" stops filtering while the log still says it is on. 3. **Fix V4 with `CWeeklyFlatLatch`**: fire on the first tick at-or-after the deadline and stay armed until actually flat, instead of a ±1-minute tick lottery. Call it from `OnTick` **and** `OnTimer` so a dead tick stream delays the close rather than cancelling it. 4. **Fix V5 with `SWNewYorkTime()`**: derives NY wall-clock from `TimeGMT()` + US DST rules. Verified against an independent reference implementation across **78,888 hourly instants, 2018–2026: 0 mismatches**. 5. **Do not hardcode Wednesday for the triple swap.** It is per-symbol and the broker publishes it as `SYMBOL_SWAP_ROLLOVER3DAYS`; several index and metal CFDs bill on Friday. `SWTripleSwapDay()` reads it. 6. `SWTripleSwapCost()` prices the carry per `SYMBOL_SWAP_MODE` and **returns false on any mode it cannot price exactly** — which blocks the override rather than treating drag as zero. --- ## 4. What was validated, and how | Claim | Method | Result | |---|---|---| | FFD weights / recurrence | `d=1` must reduce to `[1,-1]` | ✓ exactly | | ADF estimator | 600-rep Monte Carlo of the null | 5th pct **−2.848** vs textbook −2.86; rejection 4.67% ≈ 5% | | Garman-Klass accuracy | 4,000 bars × 500 intraday steps, known σ | recovered **−4.2%** (discrete-sampling bias, caveat 3) | | GK efficiency | variance of estimator vs close-to-close | **9.97×** (theory ~7.4×) | | GK non-negativity | algebra + 200k adversarial degenerate bars | non-negative; floor matches `0.1137·ln(C/O)²` to 1e-15 | | Expansion label horizon | synthetic H4, 3 weeks | all spans within **[48,72]h**; 43% clipped | | NY DST offset | independent brute-force reference, 2018–2026 | **0 / 78,888 mismatches** | Two of my own drafting errors were caught by these checks and corrected in place: an initial test series went negative before a `log()` (producing NaN-spliced false stationarity), and my first draft claimed single-bar GK could be negative — it cannot, and the real hazard is that a corrupt bar yields a *plausible positive* variance instead. That corrected caveat is now in both the Python and MQL5 headers. --- ## 5. Files delivered ``` research/fracdiff.py FFD weights, fixed-width transform, ADF scan, d-selection research/vol_target.py GK / Parkinson / Rogers-Satchell / Yang-Zhang + expansion label mql5_patches/FracDiff.mqh CFracDiff — 1024-ceiling-safe, log-price, series-order mql5_patches/VolEstimators.mqh GK/Parkinson/YZ with explicit EMPTY_VALUE and H