forked from animatedread/Warrior_EA
Root cause of the B=32 regression, and it predates F4 entirely. Every Adam
kernel stored v already square-rooted and then fed that stored value back in
as if it were the variance:
v_new = sqrt(b2 * v_old + (1 - b2) * g^2)
That recursion has a fixed point at v ~= b2 = 0.999 for ANY gradient below
unit scale, so the denominator stops tracking the gradient and Adam degrades
into plain SGD with lr = lt. Measured against the shipped WarriorCPU.dll
(batch_accum_check.cpp, TestOptimizerScaleInvariance), 4000 steps of a
constant gradient: 3285x less displacement at |g|=1e-5 than at |g|=1, where
a scale-invariant optimizer gives the same distance for both. After the fix
all six magnitudes read 1.199 and v tracks |g| exactly.
It hit conv/LSTM specifically because they sit behind a batch-norm with
running variance ~2.6e+05, so their gradients arrive divided by ~500 - deep
in the degraded regime - while the dense stack near the loss stayed in the
working one. In situ on SP500 H1: lstm1 dW/W 2.62/10.0/7.14% -> 0.024/0.022/
0.003%, conv1 decaying to 0.000% by era 30. NeuronBatchNorm.mqh already
squared v back for gamma/beta and its comment named the kernels as wrong,
which is exactly why gamma/beta kept training while the stages behind froze.
Persisted .nnw needs no migration - v keeps its std-dev meaning.
Also, the two ways F4 exposed it, both mine:
- No LR compensation for B fewer steps per era. sqrt(B) for adaptive methods
(Krizhevsky 2014; Granziol et al. 2022), applied once in
InitialEtaForOptimizer(). Linear scaling (Goyal et al. 2017) is for SGD.
- Plateau patience denominated in eras, so raising B made the ladder 32x more
impatient in its only unit. PAI converged at era 41 on ~49k updates where
the same config had been finding new bests at era 1028.
TrainPlateauPatienceEras() stretches it by the same sqrt(B).
TRAIN_BATCH_SIZE 32 -> 8 so the patience stretch stays affordable (8 -> 23
eras per stage, not 8 -> 45). Both helpers are identities at B=1.
Deploy gate: DEPLOY_MIN_SIDE_RECALL_PCT (10%) folded into tradeableOK. The
perceptron reported Sell:0% recall in all 41 eras, cleared the floor on Buy
alone at 36.6% vs 34% chance, deployed, and sprayed buy arrows. Folded into
the ranking key rather than checked at deploy time so a one-sided era cannot
become best-so-far in the first place.
Deinit: the arrow purge now runs BEFORE ExtPanel.Destroy(), an unbounded
CAppDialog teardown that sat ahead of it - the same ordering inversion the
rule there exists to prevent. CONV was force-terminated 4.8 s into OnDeinit
(vs ~1.1 s for the three that finished) having reached none of its cleanup,
so its arrows stayed on the chart. Steps are now timed in the log.
PurgeChart's verification rescan filtered on OBJ_ARROW, the same blind spot
as the bulk delete, so "persisted 10 ... cleared 0" passed silently. It now
walks every object type and reports the object counts when both are zero.
Both build variants compile 0 errors / 0 warnings; both DLLs rebuilt.
FORCES A RETRAIN (already forced by N1) and both DLLs must ship with the .ex5.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
441 lines
31 KiB
Markdown
441 lines
31 KiB
Markdown
# Training pipeline audit — plateau escape, stability, normalization, capacity
|
|
|
|
Date: 2026-08-09. Scope: everything between the feature vector and the deployed checkpoint —
|
|
input normalization, derived topology, optimizer/learning-rate mechanics, the era loop's
|
|
correction mechanisms (checkpoint restore, eta decay, plateau ladder, shadow EMA), and the
|
|
selection/deploy gates. Every claim below carries a file:line reference; nothing is quoted
|
|
from memory.
|
|
|
|
---
|
|
|
|
## 0. Verdict in one paragraph
|
|
|
|
The pipeline is structurally sound and unusually well-defended (purged OOS split, logit-adjusted
|
|
loss, family-wise deploy gate, checkpoint restore on regression). The instability you see between
|
|
eras has three real, fixable contributors: **(1)** training is pure batch-size-1 online Adam, so
|
|
per-era weight trajectories are intrinsically noisy; **(2)** the learning-rate machinery operates
|
|
inside a 3x total dynamic range (1e-4 … 3e-4), so the plateau ladder's "warm restart" is a no-op
|
|
whenever eta is already at its ceiling — which is exactly the state a non-regressing plateau is in;
|
|
**(3)** the pass-2 shuffle uses MQL5's 15-bit `MathRand()`, which makes the Fisher-Yates provably
|
|
non-uniform on any era with more than 32,768 queued samples (every full-history H1 run). Fixing
|
|
these will make training *smoother and more monotone*. What they will not do is make OOS accuracy
|
|
climb "better and better": the repo's own measurements (noise-floor MI verdict, lag-profile verdict,
|
|
26 classic-pattern models at chance on 178k bars) say the directional information in these features
|
|
is at or below the noise floor. A neural net does not find patterns where there is no mutual
|
|
information; it finds patterns where nobody *looked*, which is a different thing. The honest
|
|
formulation of your goal is: IS error should fall monotonically (it can), OOS should rise to the
|
|
data's information ceiling and hold there stably (it can), and the family-wise gate decides whether
|
|
that ceiling is tradeable (so far it has said no).
|
|
|
|
---
|
|
|
|
## 1. Direct answers to your questions
|
|
|
|
### "We use Adam as the main optimizer and switch to SGD for fine-tuning — right?"
|
|
|
|
**No.** There is exactly one optimizer per model, chosen by the `TrainingOptimizer` input
|
|
(default `ADAM`, [Variables/Inputs.mqh:210](../Variables/Inputs.mqh#L210)), and it runs for the
|
|
whole life of the model. An Adam→SGD switch *used to exist* ("replay-only optimizer override")
|
|
and was removed on 2026-07-28 for good reason: it forced the entire net to SGD during pass 2,
|
|
and pass 2 is the **only** place `backProp()` runs — so "Adam with SGD fine-tuning" was in fact
|
|
"SGD at Adam's learning rate, always", with Adam's moments never updated. The removal note is at
|
|
[Training.mqh:680-689](../Expert/AIBase/Training.mqh#L680-L689). Do not reintroduce a switch;
|
|
the literature on optimizer switching (SWATS) shows mixed results, and this codebase has already
|
|
been burned once by a half-switched optimizer.
|
|
|
|
### "Should we play with learning rate and momentum when we keep regressing?"
|
|
|
|
The pipeline already adjusts the learning rate on regression — but inside a range so narrow the
|
|
adjustments barely do anything. Current mechanics:
|
|
|
|
| Mechanism | Trigger | Effect | Where |
|
|
|---|---|---|---|
|
|
| eta decay | selection score drops >5pp below best (`ETA_DECAY_REGRESSION_PCT`) | best checkpoint restored, `eta *= 0.7`, floor 1e-4 | [Training.mqh:1545-1554](../Expert/AIBase/Training.mqh#L1545-L1554) |
|
|
| recovery bump | new best era | `eta /= 0.7`, cap at ceiling | [Training.mqh:1482](../Expert/AIBase/Training.mqh#L1482) |
|
|
| warm restart | 8 / 16 eras with no new best (plateau stages 1, 2) | `eta = m_etaCeiling` | [Training.mqh:1588-1602](../Expert/AIBase/Training.mqh#L1588-L1602) |
|
|
| terminal | 24 eras with no new best (stage 3) | deploy best checkpoint, stop | [Training.mqh:1604-1662](../Expert/AIBase/Training.mqh#L1604-L1662) |
|
|
|
|
Constants: ceiling = `AdamLearningRate` = 3e-4 ([AI/Network.mqh:68](../AI/Network.mqh#L68)),
|
|
`ETA_MIN` = 1e-4, `ETA_DECAY_FACTOR` = 0.7
|
|
([ExpertSignalAIBase.mqh:295-297](../Expert/ExpertSignalAIBase.mqh#L295-L297)).
|
|
|
|
**Finding F2 (high impact): the plateau ladder's only escape is frequently a no-op.**
|
|
eta starts at the ceiling and only ever *leaves* the ceiling via the regression-decay branch.
|
|
A run that plateaus *without* a >5pp regression therefore sits at eta == ceiling, and both
|
|
"warm restarts" assign the value eta already has. In that state the ladder is not an escape
|
|
mechanism at all — it is a 24-era countdown to deploy. Even in the best case (eta floored at
|
|
1e-4), the restart is a 3x kick; SGDR-style restarts that actually relocate a model typically
|
|
span 10-100x. Additionally, escaping a basin generally requires a rate *larger* than the one
|
|
that settled into it — restoring the original rate revisits the same basin.
|
|
|
|
**Momentum:** `AdamBeta1` = 0.9 / `AdamBeta2` = 0.999, book defaults
|
|
([AI/Network.mqh:68-77](../AI/Network.mqh#L68-L77)); bias correction is implemented correctly
|
|
per-neuron (`lt = eta*sqrt(1-b2^t)/(1-b1^t)`,
|
|
[NeuronOCLBase.mqh:564,613](../AI/Impl/NeuronOCLBase.mqh#L613)). Hand-tuning beta1 was already
|
|
tried (0.8 experiment) and correctly reverted. The momentum problem is not the constants — see F3.
|
|
|
|
**Finding F3 (medium): checkpoint restore keeps stale Adam moments.**
|
|
`CaptureWeights()/RestoreWeights()` snapshot weights only — documented deliberately at
|
|
[AI/Network.mqh:686-689](../AI/Network.mqh#L686-L689). After a restore, the very next updates are
|
|
driven by first/second moments accumulated *along the rejected trajectory*, i.e. the optimizer
|
|
immediately pushes the restored weights back toward the state that just got rolled back. The delta
|
|
clip and decayed eta bound the damage but don't remove the direction. This is a plausible direct
|
|
cause of the "restore → regress again → restore again" oscillation you're watching. Zeroing the
|
|
moment buffers (and the per-neuron `t` counters) on restore — and on plateau warm restarts — is
|
|
cheap and is standard practice in Adam-with-restarts setups.
|
|
|
|
### "Do we have enough neurons and layers?"
|
|
|
|
The shape is *derived from the data*, not configured, and the derivation is deliberate about being
|
|
small: the first dense layer gets at most **one weight per in-sample bar**
|
|
([Topology.mqh:809-841](../Expert/AIBase/Topology.mqh#L809-L841)), snapped down a
|
|
{16,32,64,…,1024} ladder; depth comes from a geometric taper to the output
|
|
([Topology.mqh:669-689](../Expert/AIBase/Topology.mqh#L669-L689)); conv filters = half the
|
|
receptive-field window snapped to {4..32}; LSTM hidden solves `4H(H+in+1) ≤ isBars`
|
|
([Topology.mqh:771-805](../Expert/AIBase/Topology.mqh#L771-L805)).
|
|
|
|
Two honest observations:
|
|
|
|
- **Capacity is not why training plateaus.** More neurons buy IS fit. The OOS ceiling is set by
|
|
how much label-relevant information the features carry, and the project's own measurements put
|
|
that at/near the noise floor for direction ([[project_lag_profile_verdict]],
|
|
[[project_mi_noise_floor_verdict]], [[project_classic_patterns_no_edge]]). A bigger net finds
|
|
the noise faster.
|
|
- **Finding F6 (low-medium): with the full default feature set the first layer is pinned at its
|
|
16-unit floor.** ~60+ features/bar × 20 bars ≈ 1,200+ inputs; at ~25-30k in-sample bars the
|
|
budget `isBars/(inputWidth+1)` lands near or below 16, so the derived width bottoms out and the
|
|
warning at [Topology.mqh:833-839](../Expert/AIBase/Topology.mqh#L833-L839) may already be firing.
|
|
Check the startup `config -` line. If it reads "N dense from 16 units", the honest options are
|
|
(a) prune feature groups the research track has already measured as edge-free, which widens the
|
|
budget per input, or (b) accept the pinch — *not* raise the budget: 1 param/sample is already
|
|
generous for this signal-to-noise regime.
|
|
|
|
---
|
|
|
|
## 2. Bug found
|
|
|
|
**Finding F1: biased Fisher-Yates — `MathRand()` is 15-bit.**
|
|
[Training.mqh:699](../Expert/AIBase/Training.mqh#L699): `int sJ = MathRand() % (sIdx + 1);`
|
|
MQL5's `MathRand()` returns 0..32767. For every `sIdx ≥ 32768` the modulo is the identity, so
|
|
`sJ` is uniform on [0, 32767] only: a slot in the upper region can never swap with another upper
|
|
slot, and each upper element gets exactly one forced exchange with the lower region. For
|
|
`sIdx` just under 32768 there is ordinary modulo bias. The queue holds one entry per IS bar
|
|
(every bar queued exactly once, [Training.mqh:616-625](../Expert/AIBase/Training.mqh#L616-L625)),
|
|
and the training window is *all available history* — an H1 symbol easily exceeds 32,768 IS bars,
|
|
so the shipped configuration hits this every era. The shuffle exists precisely to break the
|
|
correlated same-class gradient runs documented at the `MAX_WEIGHT_DELTA` comment
|
|
([AI/Network.mqh:350-356](../AI/Network.mqh#L350-L356)); a non-uniform shuffle partially
|
|
re-admits them, era after era, in the same pattern (the bias is deterministic in structure).
|
|
|
|
**Fix (small):** compose a 30-bit value — `int r30 = (MathRand() << 15) | MathRand();` then
|
|
`sJ = r30 % (sIdx + 1);` (or rejection-sample to kill the residual modulo bias — at 30 bits
|
|
against ~100k the residual is negligible). Same fix applies to
|
|
[AutoTune.mqh:455](../Expert/AIBase/AutoTune.mqh#L455) if its array can exceed 32k (check).
|
|
|
|
---
|
|
|
|
## 3. Why eras "visually regress" — the full causal list
|
|
|
|
1. **Batch-size-1 online Adam (F4, the big one).** One weight update per sample
|
|
([NeuronBatchNorm.mqh:60-63](../AI/NeuronBatchNorm.mqh#L60-L63) states it plainly: "pure online
|
|
SGD - one weight update per sample, never a batched pass"). Gradient noise at batch 1 is maximal;
|
|
the end-of-era weight state is a random variable with substantial variance, so consecutive eras
|
|
genuinely differ even with identical data. Every defense in the file (delta clip, weight decay,
|
|
shadow EMA, checkpoint restore) is compensating for this at the symptom level.
|
|
**Literature-backed fix: gradient accumulation into mini-batches of 16-64.** Noise scales
|
|
~1/sqrt(B); B=32 cuts update variance ~5-6x, smooths the loss trajectory, makes Adam's second
|
|
moment estimate meaningful, and typically *speeds up* wall-clock convergence (fewer, better
|
|
steps). This is the single highest-impact stability change available. Implementation: accumulate
|
|
output-layer gradients across B samples before calling the update kernels — the shuffle,
|
|
logit-adjust and sample-weight machinery are all unaffected.
|
|
2. **Stale Adam moments after checkpoint restore (F3, above).**
|
|
3. **No-op warm restarts (F2, above)** — the ladder never actually perturbs a stuck run, so a run
|
|
oscillates in the 5pp dead zone (no restore, no decay: [Training.mqh:1452](../Expert/AIBase/Training.mqh#L1452))
|
|
until the timer expires.
|
|
4. **Batch-norm statistics move on every forward pass, including OOS scoring (F5).** Documented as
|
|
deliberate ([NeuronBatchNorm.mqh:70-79](../AI/NeuronBatchNorm.mqh#L70-L79)): pass 3 scoring
|
|
advances the EMA mean/variance, so (a) the OOS number partly measures BN drift rather than the
|
|
trained function, and (b) the same weights score differently depending on what was scored
|
|
before them. `SetBatchNormFrozen` already exists and is already used by
|
|
`ValidateCpuInference` ([Persistence.mqh:243-246](../Expert/AIBase/Persistence.mqh#L243-L246))
|
|
for exactly this reason. **Fix: freeze BN stats for the duration of pass 3** (and the OOS
|
|
simulation walk), unfreeze after. This changes measurement only, not live behavior, and makes
|
|
the selection metric a pure function of the checkpoint — which is what a metric that decides
|
|
deployment should be.
|
|
5. **Part of the "regression" is measurement noise, by design.** The per-era selection score at
|
|
~11k directional calls has ~0.4pp standard error, more at low-coverage eras; the 5pp
|
|
dead zone deliberately ignores oscillation inside it; and `dOosForecast` is a 10,000-sample EMA
|
|
([NetBuild.mqh:16](../AI/Impl/NetBuild.mqh#L16)) that carries momentum across eras. Two eras
|
|
that "look" different on the chart (NMS arrows redraw per era) can be statistically identical.
|
|
|
|
---
|
|
|
|
## 4. Normalization audit — result: sound, two representation nits
|
|
|
|
- **Price-unit features:** ATR-normalized at source, with the rationale spelled out at
|
|
[Features.mqh:477-488](../Expert/AIBase/Features.mqh#L477-L488). Correct and cross-symbol stable.
|
|
Bars with no ATR or EMPTY_VALUE are rejected, not zero-filled.
|
|
- **Clamps:** every block clamps to single digits (±10 swing, ±5 volume, ±2 delta); one
|
|
whole-bar sanitize gate zero-fills anything non-finite or >1e4
|
|
([Features.mqh:921-937](../Expert/AIBase/Features.mqh#L921-L937)) — this is the fix for the
|
|
2026-08-02 BN NaN-latch and it is correctly placed. BN has its own 1e6 input bound and a
|
|
1e-4 std floor ([NeuronBatchNorm.mqh:103-109](../AI/NeuronBatchNorm.mqh#L103-L109)).
|
|
- **Layer-level:** batch norm ON by default (window 1000), placed between dense pairs and before
|
|
the head ([Topology.mqh:1072-1134](../Expert/AIBase/Topology.mqh#L1072-L1134)), which absorbs
|
|
the remaining cross-family scale differences. EMA-form stats with cold-start bias ramp — correct
|
|
for an online regime.
|
|
- **Nit N1: `EventCode` (and `StructuralPhase`) are categorical codes fed as ordinals**
|
|
([Features.mqh:886-890](../Expert/AIBase/Features.mqh#L886-L890)). The network is told event 9
|
|
is "three times" event 3. A first dense layer can partially untangle this, but it wastes
|
|
capacity the 16-wide entrance doesn't have. Options: one-hot the code (costs ~10-20 inputs —
|
|
conflicts with F6), or collapse to the 2-3 axes that matter (event fired y/n, direction ±1,
|
|
phase 1-5 as its own scaled scalar). Low urgency; do it with the next forced retrain, since it
|
|
changes the input contract (fingerprint `WIN`-style version bump).
|
|
- **Nit N2:** heterogeneous natural ranges remain (RSI/100 in [0,1] vs ±10 clamps vs raw Wyckoff
|
|
readings "with no natural range", [Features.mqh:922-923](../Expert/AIBase/Features.mqh#L922-L923)).
|
|
With BN on this is mostly harmless; without BN it would matter. Fine as is *because* BN defaults on —
|
|
note the coupling if BN is ever disabled for an experiment.
|
|
|
|
---
|
|
|
|
## 5. Loss / head / labels — audited, no defects found
|
|
|
|
Joint softmax+CCE gradient over sigmoid-bounded outputs with `CLASS_LOGIT_SCALE` 6.0 temperature
|
|
([NetForward.mqh:387-401](../AI/Impl/NetForward.mqh#L387-L401),
|
|
[AI/Network.mqh:366-377](../AI/Network.mqh#L366-L377)); logit adjustment applied backward-only with
|
|
per-era re-measured priors ([Training.mqh:340-345](../Expert/AIBase/Training.mqh#L340-L345)); label
|
|
smoothing 0.9/0.05/0.05 sums to 1.0; AdamW-style decoupled decay at 1e-3 with the 0.01-failure
|
|
history documented; per-step delta clip 0.1; saturation-derivative floor 1e-3. The bounded-head +
|
|
temperature design is unconventional but internally consistent, and the comments correctly forbid
|
|
unbinding the head without retuning the two constants calibrated against it
|
|
([Topology.mqh:1144-1162](../Expert/AIBase/Topology.mqh#L1144-L1162)). Leave this stack alone.
|
|
|
|
---
|
|
|
|
## 6. Prioritized action plan
|
|
|
|
| # | Change | Effort | Expected effect | Risk |
|
|
|---|---|---|---|---|
|
|
| 1 | **F1** 30-bit shuffle RNG ([Training.mqh:699](../Expert/AIBase/Training.mqh#L699)) | trivial | removes a deterministic bias present in every full-history era | none |
|
|
| 2 | **F4** mini-batch gradient accumulation (B≈32) in `backProp` | medium | largest stability win; smoother, likely faster convergence | moderate (touches update path; verify dW/W in situ per [[feedback_verify_in_situ_not_offline]]) |
|
|
| 3 | **F3** zero Adam moments + `t` on `RestoreWeights()` and on plateau warm restarts | small | stops restore→re-regress oscillation | low |
|
|
| 4 | **F2** give restarts amplitude: restart eta to 3-10x ceiling for a bounded number of eras (delta clip already bounds per-step damage), widen `ETA_MIN` to 1e-5; optionally cosine-decay within the cycle (SGDR) | small | turns the ladder into a real escape instead of a timer | low-moderate |
|
|
| 5 | **F5** freeze BN stats during pass-3 scoring + OOS sim | small | selection metric becomes a pure function of the checkpoint | none (measurement only) |
|
|
| 6 | **F6/N1** prune measured-dead feature groups; re-encode EventCode | medium | frees entrance capacity; cleaner representation | forces retrain (fingerprint bump) |
|
|
|
|
Order matters: do 1, 3, 5 together (cheap, independently safe), then 2, then 4 — because 4's effect
|
|
can't be evaluated while per-sample noise (2) dominates the era-to-era variance.
|
|
|
|
### Implementation status (2026-08-09, same day — compiled clean, 0 errors 0 warnings)
|
|
|
|
- **F1 DONE** — `ShuffleRandomIndex()` (30-bit) added in ExpertSignalAIBase.mqh, used by the pass-2
|
|
queue shuffle (Training.mqh) and the MI-null block shuffle (AutoTune.mqh).
|
|
- **F3 DONE** — `CNet::ResetOptimizerState()` (NetWeights.mqh) + per-neuron-type overrides
|
|
(dense/conv/LSTM/batch-norm, both OCL and legacy scalar hierarchies; `ZeroOptimizerBuffer()` in
|
|
BufferDouble.mqh). Called on the mid-run regression restore, on every plateau warm restart, and on
|
|
the deploy-time restore in FinalizeTrainRun (online learning continues from that net). Weights,
|
|
BN running statistics, and gamma/beta are untouched — optimizer state only.
|
|
- **F2 DONE** — `PLATEAU_RESTART_BOOST 5.0`: restarts now jump eta to 5x the ceiling (1.5e-3) with a
|
|
geometric anneal back to the ceiling over `PLATEAU_PATIENCE_ERAS` eras (SGDR-style bounded cycle),
|
|
cleared early on any new best; `ETA_MIN` widened 1e-4 → 1e-5. Each restart also resets optimizer
|
|
moments (F3), so the kick explores instead of replaying the plateau's momentum.
|
|
- **F5 DONE** — BN running statistics frozen for the whole pass-3 OOS scoring walk
|
|
(`SetBatchNormFrozen(true/false)`), with a defensive unfreeze in FinalizeTrainRun for the
|
|
stop-mid-pass path. The selection metric is now a pure function of the checkpoint. The OOS
|
|
continual-learning *simulation* is deliberately NOT frozen — it evaluates live adaptive behaviour,
|
|
where the moving statistics are part of what is being simulated.
|
|
- **F4 DONE** (second pass, same day). True mini-batch accumulation, `TRAIN_BATCH_SIZE = 32`.
|
|
Design: the O(n²) outer product is native (two new kernels/exports per backend —
|
|
`AccumulateWeightGrad`, `AccumulateWeightGradConv`, plus a generic `AccumulateBufferInto`), while
|
|
the optimizer step is **host-side MQL5, shared by every tier** (`ApplyAccumToBlock`), so there is
|
|
one Adam/SGD implementation instead of four that can drift. Both DLLs rebuilt. Notes:
|
|
- The LSTM needs no outer-product kernel — by the update pass its `WeightsGradient` already holds
|
|
the sample's full dW. It could *not* simply be left un-zeroed between samples, because
|
|
`CPU_LSTMSeqBackward`/`DML_LSTMSeqBackward` **memset that buffer on entry**; hence the separate
|
|
accumulator plus elementwise add.
|
|
- Batch-norm's gamma/beta accumulate in host arrays, not new `BatchOptions` slots — `BN_OPT_STRIDE`
|
|
is baked into every persisted `.nnw` via `getWeightsBN`.
|
|
- Scoped to pass 2 only; `OnlineLearnStep` keeps immediate per-sample updates. Every save,
|
|
checkpoint and scoring boundary calls `FlushBatch()`, which scales by the *real* sample count so a
|
|
short trailing batch still takes a correctly-sized step.
|
|
- Degrades, never fails: `CNet::BatchSize()` returns 1 (with a one-time log line) on a tier that
|
|
cannot accumulate — an OpenCL device where the kernels didn't build, or the legacy scalar tier a
|
|
DLL-free Market build can land on. Those train exactly as before.
|
|
- **Verified offline**: `DirectML/batch_accum_check.cpp` (+ `build_accum_check.bat`) drives the real
|
|
exports and compares against an independent reference. All pass, including the load-bearing one —
|
|
at B=1 the accumulator equals the gradient the *shipped unbatched kernel* forms internally, to
|
|
1.1e-16. That proves the math, **not** that a layer trains in the assembled net; the in-situ check
|
|
is still the per-layer `dW/W` report on a real era.
|
|
- **F6 DONE, and the root cause was not what this report first said.** Confirmed from the deployed
|
|
`.cfg` files that CONV, LSTM and HYBRID were *all* pinned at the 16-unit floor (64 features/bar ×
|
|
20 bars = 1,280 inputs). But the fix is not pruning features: `ComputeFirstLayerWidth()` was
|
|
budgeting against the **raw** input width even on topologies where a conv/LSTM front end has
|
|
already reduced it — an LSTM hands the dense stack 64 values, not 1,280, so it was being charged
|
|
~20x its real fan-in. Now budgeted against the front-end output, and capped at that width so the
|
|
first dense layer can never fan *out* (a shape `FrontEndConfigSummary` already flags as a defect).
|
|
`InitNeuralNetwork` reorders the derivation accordingly (conv filters → LSTM hidden → first-layer
|
|
width → depth). Expected effect on the shipped config: LSTM/HYBRID 16 → 64 units, CONV 16 → 32/64;
|
|
plain MLP unchanged.
|
|
- **N1 DONE, with a correction to this report's premise.** Reading the indicator, `EventCode` is
|
|
±1..7 where the sign is accumulation/distribution and the magnitude is position in the Wyckoff
|
|
schematic (PS→SC→AR→ST→Spring→LPS→SOS) — so it is genuinely *ordinal*, and this report's "event 9
|
|
is three times event 3" framing was wrong (there is no event 9). The real defect is that direction
|
|
and stage are entangled in one scalar across a sign discontinuity — precisely what the base OHLC
|
|
block already fixes by giving direction its own ±1/0 flag. All three signed Wyckoff categoricals
|
|
(`EventCode`, `EventPhase`, `StructuralPhase`) are now split into a direction and a [0,1]-scaled
|
|
magnitude. Information-preserving (the pair reconstructs the original exactly); nothing was dropped
|
|
— including the collinear-but-nonlinear `StructuralPhase`, kept deliberately so this change is
|
|
purely a re-encoding and any effect is attributable to it alone. 13 readings now occupy 16 inputs.
|
|
|
|
## 6b. POST-DEPLOY ADDENDUM — the first run at B=32 regressed, and why
|
|
|
|
The change set above was deployed and run on SP500 H1 (four charts, 12:49–13:30). It regressed every
|
|
topology. The diagnosis found one root cause that predates F4 entirely, plus two ways F4 made it
|
|
visible. Recorded here because the root cause invalidates part of the reasoning above.
|
|
|
|
### What the run showed
|
|
|
|
| reading | before (`ee48381`) | after (`0c01dc2`) |
|
|
|---|---|---|
|
|
| `lstm1` dW/W, eras 2/3/4 | 2.62% / 10.0% / 7.14% | 0.024% / 0.022% / **0.003%** |
|
|
| `conv1` dW/W, early → late | 0.52% → 1.18% @ era 42 | 0.35% → **0.000% @ era 30** |
|
|
| PAI IS error | 0.41–0.42 @ era 1028+ | 0.48 flat, converged @ era 41 |
|
|
|
|
`conv1`/`lstm1` in that report read `getWeightsConv`/`getWeightsLSTM` only, so this is unambiguous:
|
|
the convolution kernel and the LSTM recurrent block stopped training. The `bn*` figures stayed
|
|
healthy, but they are not a counter-example — their norm is dominated by running variance
|
|
(`bn2` norm 263,053 against `var` 2.65e+05) and BN running stats update per *sample* in the forward
|
|
pass, where batching cannot reach them.
|
|
|
|
### R1 — THE ADAM SECOND MOMENT WAS NOT ADAM (all four tiers)
|
|
|
|
Every Adam kernel in the engine stored `v` already square-rooted and then fed that stored value
|
|
**back in as if it were the variance**:
|
|
|
|
```
|
|
v_new = sqrt(b2 * v_old + (1 - b2) * g^2) // v_old is a standard deviation, not a variance
|
|
```
|
|
|
|
That recursion has a fixed point at `v ≈ b2 = 0.999` for *any* gradient below unit scale, because
|
|
`(1-b2)·g²` becomes negligible against `b2·v`. The denominator therefore stops tracking the gradient
|
|
scale, and `delta = lt·m/v` degrades from scale-invariant Adam into plain SGD with `lr = lt`.
|
|
|
|
Measured against the shipped `WarriorCPU.dll`
|
|
([batch_accum_check.cpp](../DirectML/batch_accum_check.cpp), `TestOptimizerScaleInvariance`),
|
|
4000 steps of a constant gradient:
|
|
|
|
| gradient magnitude | stored `v` | displacement | vs magnitude 1 |
|
|
|---|---|---|---|
|
|
| 1e+0 | 1.000000 | 1.019 | 1x |
|
|
| 1e-2 | 0.999000 | 1.050e-2 | 97x less |
|
|
| 1e-5 | 0.999000 | 3.101e-4 | **3285x less** |
|
|
|
|
A scale-invariant optimizer gives the same displacement in every row. After the fix — square `v`
|
|
back before re-entering the recursion, which *is* the textbook second moment, merely carried in
|
|
std-dev form so the stored value can be the denominator directly — all six rows read 1.199 and
|
|
`v` tracks the gradient magnitude exactly.
|
|
|
|
**Why this hit the front end specifically.** Conv and LSTM sit *behind* a batch-norm whose running
|
|
variance is ~2.6e+05, so their gradients arrive divided by ~500 — deep in the degraded regime — while
|
|
the dense stack nearer the loss sees gradients of order 1 and stayed in the working regime. This also
|
|
explains a recurring theme in this engine's history: the front-end stage has *always* been the one
|
|
that "won't train".
|
|
|
|
**It was already known, in one place.** `NeuronBatchNorm.mqh` squares the stored value back for
|
|
gamma/beta and its comment named the kernels as wrong, but deliberately left them alone because
|
|
fixing them "would alter the behaviour of every existing model on all four backends at once". That
|
|
is exactly why gamma/beta kept learning while the stages behind them froze. Now reconciled: OpenCL
|
|
(`Network.cl` x3), `WarriorCPU.cpp` x3, `WarriorDML.cpp` HLSL x3, the host-side batched step in
|
|
`NeuronOCLBase.mqh`, and the legacy scalar tier in `NeuronCPU.mqh`.
|
|
|
|
The persisted `.nnw` needs no migration — `v` keeps its std-dev meaning, only the recursion changed.
|
|
|
|
### R2 — mini-batching was never paid for
|
|
|
|
Two compounding errors in F4, both mine:
|
|
|
|
- **No learning-rate compensation.** B=32 takes 32x fewer optimizer steps per era. The linear rule
|
|
(Goyal et al. 2017) is for SGD; for adaptive methods the rule is `sqrt(B)` (Krizhevsky 2014;
|
|
derived for Adam in Granziol et al. 2022). Now applied once, in `InitialEtaForOptimizer()`.
|
|
- **Patience denominated in eras.** An era is a data pass, so raising B silently made the plateau
|
|
ladder 32x more impatient in the only unit it measures. PAI declared convergence at era 41 on ~49k
|
|
updates where the same config had been finding new bests at era 1028. `TrainPlateauPatienceEras()`
|
|
now stretches it by the same `sqrt(B)`.
|
|
|
|
Both are exact identities at B=1. `TRAIN_BATCH_SIZE` is 32 → **8**: after the `sqrt(B)` LR bump an era
|
|
still makes `sqrt(B)` less progress, so patience must stretch by `sqrt(B)` too — at 32 that is 8 → 45
|
|
eras per ladder stage on a topology already taking 70 s/era, at 8 it is 8 → 23. Noise still falls as
|
|
`1/sqrt(B)`, so 8 keeps ~2.8x of the variance reduction that was the point.
|
|
|
|
Note R2 could not have been fixed in terms of B alone while R1 stood: under the degraded optimizer,
|
|
averaging B samples shrank the front end's gradient by a further `sqrt(B)` *on top of* the B fewer
|
|
steps, and step size was proportional to gradient magnitude.
|
|
|
|
### R3 — the deployability floor let a one-sided model ship
|
|
|
|
PAI reported `Sell:0%` recall in all 41 eras, cleared the floor on Buy alone (36.6% against 34%
|
|
chance) and deployed — the "buy sprayer" on the chart. `tradeableOK` tested coverage and precision
|
|
only; coverage counts directional calls without caring that they are all the same direction. Added
|
|
`DEPLOY_MIN_SIDE_RECALL_PCT` (10%), folded into `tradeableOK` so a one-sided era cannot become the
|
|
best-so-far in the first place, not merely be refused at the deploy gate. Deliberately far below the
|
|
40%-per-class diagnostic, which is unreachable on this data and would block every deployment.
|
|
|
|
CONV, for contrast, was correctly blocked by the existing floor — the gate works, it was just blind
|
|
to one-sidedness.
|
|
|
|
### R4 — the two chart-cleanup bugs
|
|
|
|
- **CONV stranded its arrows** because `OnDeinit` was force-terminated: 13:29:58.202 → "Abnormal
|
|
termination" 13:30:03.002, 4.8 s against ~1.1 s for the three charts that finished, having reached
|
|
none of its cleanup. That is MetaTrader's OnDeinit budget expiring, not a fault. `ExtPanel.Destroy()`
|
|
— an unbounded CAppDialog teardown — sat *ahead* of the arrow purge, the same ordering inversion the
|
|
rule there exists to prevent. Arrow cleanup now runs first, and each step is timed so the log names
|
|
the slow one.
|
|
- **`persisted 10 ... cleared 0`** on LSTM and PAI is unresolved and needs a run to settle: two scans
|
|
microseconds apart disagreed about the same chart, and the WARNING branch built to catch exactly
|
|
that stayed silent because its rescan filtered on `OBJ_ARROW` in the same way the bulk delete did.
|
|
The rescan now walks every object type, and the both-zero case reports the object counts explicitly
|
|
instead of passing silently.
|
|
|
|
### Revised deploy notes
|
|
|
|
### Deploy notes for this change set
|
|
|
|
1. **Both DLLs must be redeployed** with the `.ex5` — they carry new exports, and now also the
|
|
corrected Adam kernels (R1). Verified: both build variants compile 0 errors / 0 warnings. `build_cpu.bat` and
|
|
`build.bat` copy them into every discovered MetaTrader `MQL5\Libraries`.
|
|
2. **N1 forces a retrain** of every Wyckoff-enabled config: the fingerprint gains `|WES:2|` and
|
|
`m_neuronsCount` goes 64 → 67. That is intended — it is also what lets F6's new shape take effect,
|
|
since derived widths are adopted from a model's existing `.cfg` rather than re-derived. A config
|
|
with the Wyckoff events *off* keeps its fingerprint and its model, and will only pick up F6 after a
|
|
manual reset-weights.
|
|
3. **First run to watch, in this order**: the startup `config -` line (first dense layer should no
|
|
longer read 16 on CONV/LSTM/HYBRID); then per-era `dW/W` (every stage must move — a stage stuck at
|
|
~0.000% means the batched path is starving it); then whether IS error falls more smoothly than
|
|
before, which is the whole point of F4.
|
|
4. If anything looks wrong, `TRAIN_BATCH_SIZE 1` in `ExpertSignalAIBase.mqh` restores the exact
|
|
pre-F4 update path without touching anything else — including the `sqrt(B)` LR scale and the
|
|
patience stretch, both of which are identities at 1. It does NOT revert R1; the Adam fix is
|
|
independent of batching and should stay on either way.
|
|
5. **R1 changes training dynamics for every topology, not just the front end.** The single reading
|
|
that confirms it worked is `conv1`/`lstm1` dW/W staying alive (order 0.1–1%) instead of decaying
|
|
toward 0.000% — check that before any accuracy number.
|
|
|
|
A further literature-backed option once the above land: deploy the **shadow EMA** net (already
|
|
maintained at tau 0.01/era, [Training.mqh:1161-1163](../Expert/AIBase/Training.mqh#L1161-L1163))
|
|
as the inference model rather than the raw best checkpoint — that is stochastic weight averaging
|
|
in all but name, and it is the standard cure for "the deployed snapshot happened to be a lucky
|
|
era". The infrastructure exists; only the deploy wiring would change.
|
|
|
|
## 7. Expectation setting — what "linear progression" can mean here
|
|
|
|
Monotone *IS* improvement is achievable (and the fixes above get you close to it). Monotone *OOS*
|
|
improvement up to a stable maximum is achievable only up to the information ceiling of the
|
|
feature/label pair, and the repo's own instruments — the family-wise selection gate
|
|
([Training.mqh:60-78](../Expert/AIBase/Training.mqh#L60-L78)), the MI noise-floor and lag-profile
|
|
verdicts — currently place that ceiling at "no deployable directional edge". The correct reading
|
|
of a post-fix run: if training becomes smooth, plateaus cleanly, and the gate still says
|
|
p_family ≈ 1, then the plateau is *real* and the next lever is the data (features/labels), not
|
|
the optimizer. The one measured positive remains excursion *size* (range 4x its null, p=0.005) —
|
|
a volatility/risk head, not a direction head. Chasing direction with a better optimizer cannot
|
|
overturn a measurement about the data.
|