EnableBatchNorm and BatchNormWindow demoted from inputs to constants. Batch norm is required, not optional: measured on identical MLP_3L topologies it was worth +11.3 points of balanced accuracy (57.0% with, 45.7% without), stable across 150+ and 200+ eras, and the no-BN control converged to ~5% IS and OOS accuracy with no chart signals at all. A user cannot make a good decision here and can easily make a ruinous one, so the choice is not offered. BatchNormWindow goes with it - a running-statistics window in samples has no meaningful setting a trader could reason about, and its only other reachable state (<=1) silently disables the layer. Kept as named constants rather than deleted: the topology builder, the weights fingerprint and the .cfg guard all read them, and a constant keeps those paths - and the ability to flip one for a diagnostic rebuild - intact. Fewer knobs also means a shorter Market description and less room for a buyer to misconfigure. EXPERIMENTS.md records runs 2 and 3, since the MT5 logs are wiped between runs and these measurements are what the design decisions rest on. Run 3 (12h, uncapped tau=1.0) is a write-off: zero eras out of 1,993 across the five batch-norm charts ever called a direction on fewer than half of all bars, at a median precision equal to the ~6.1% base rate. The damage was present at era 1 and never recovered over 292-766 eras. Both builds compile 0 errors, 0 warnings. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
5.2 KiB
Training experiment log
Results from multi-chart comparison runs on SP500 H1. Kept because the raw MT5 logs get wiped between runs and these are the measurements the design decisions rest on.
Run 3 — 2026-07-29 21:06 → 2026-07-30 08:46 (~12h)
Build: logit-adjusted loss at uncapped tau = 1.0, conv pooling removed, BN on all
charts except the control. Verdict: the whole run is a write-off. Kept as the record of
one specific failure.
Six charts, identical config except dense depth and the one EnableBatchNorm = false
control (verified from the startup config lines — all six matched on every other field).
| chart | eras | peak bal | median coverage | median precision | eras calling <50% of bars |
|---|---|---|---|---|---|
CONV_2L [4109] |
766 | 64% | 100% | 6.0% | 0 / 766 |
MLP_3L [39c1] |
485 | 64% | 100% | 6.0% | 0 / 485 |
MLP_4L [d3b2] |
460 | 64% | 100% | 6.2% | 0 / 460 |
HYBRID_2L [4109] |
390 | 64% | 100% | 5.9% | 0 / 390 |
LSTM_2L [4109] |
292 | 64% | 100% | 5.9% | 0 / 292 |
MLP_3L no-BN [271b] |
49 (stopped) | 59% | 83% | 5.4% | 19 / 49 |
Directional base rate is ~6.1%, so median precision ≈ the base rate: no information at all. Not one era in 1,993 across the five batch-norm charts ever called a direction on fewer than half of all bars.
Root cause
tau = 1.0 against a bounded sigmoid head. Offsets are tau·log(prior_c), spread 3.42 on
this 30:1 imbalance, against a head whose widest expressible logit gap is
CLASS_LOGIT_SCALE × 1 = 6. The prior correction consumed 57% of the entire range, so
the network's only available response was to saturate Buy/Sell to 1.0 to overcome a −3.42
training handicap. The offsets are absent at inference, so every bar came out directional.
The damage was present at era 1 and permanent — Neutral recall was 2–7% in the very first era of every batch-norm chart, and 292–766 eras never recovered it. This was not drift; the optimization target itself was wrong from the first gradient step.
Fixed in 0e5f1bb by expressing strength relative to the range actually available.
What the plateau ladder did — correctly
Dozens of PLATEAU stage 3 … nothing safe to deploy cycles. The per-class recall floor
refused every one of these models. Only the no-BN control ever deployed, and only because
it had an early genuinely-balanced checkpoint before it collapsed.
Caveat on the new selection metric
Precision-led selection (a142749) would not have rescued this run. Every one of the
1,993 eras was an over-caller; ranking correctly among uniformly bad options still yields a
bad option (best-precision era per chart was 6.5% at ~100% coverage). Selection prevents
choosing wrongly when good checkpoints exist. The tau cap is what makes good checkpoints
exist at all. They fix different failures and both were needed.
Caveat on batch norm
This run does not support "batch norm is required" — it is confounded. The no-BN control
was the only chart that stayed selective (19 of 49 eras below 50% coverage) precisely
because batch norm preserves logit spread into the head, which under a broken tau made
saturation easier. The real evidence for batch norm is Run 2 below.
Run 2 — 2026-07-29, ~3h
Build: minority replay (28×) + post-hoc prior, batch norm, derived topology, conv pooling still present.
| chart | topology | BN | peak balanced | peak Buy/Sell recall |
|---|---|---|---|---|
042e |
MLP_4L | on | 60.7% — converged, auto-deployed era 58 | 51% / 45% |
6ca2 |
MLP_3L | on | 57.0% | 54% / 26% |
cfcc |
HYBRID_2L | on | 53.0% | 53% / 12% |
cfcc |
LSTM_2L | on | 52.3% | 46% / 16% |
c3a7 |
MLP_3L | off | 45.7% | 1% / 42% |
cfcc |
CONV_2L | on | 45.3% (510 eras) | 49% / 0% |
- Batch norm worth +11.3 points on identical MLP_3L (57.0 vs 45.7), stable across 150+
and 200+ eras. Verified which chart was which by parsing
.nnwlayer records: BN-off = 5 layers, BN-on = 9, MLP_4L = 11. This is the evidence for batch norm. - Depth ordering reversed vs the pre-BN era — deepest now best, closing out the over-parameterization diagnosis.
- MLP_4L deployed at balanced 60.7%, blended OOS 83.9%, calling ~6% of bars at 20–35% win rate against a ~6% base rate — the only tradeable model produced so far.
- Only 1 era in 1,301 cleared the ≥40% per-class recall floor.
- Failure mode was directional mode collapse: each model took one direction to ~50% recall and abandoned the other, direction chosen arbitrarily.
- CONV pinned at ~40% for 510 eras with Sell recall 0% → traced to the pooling stage
reducing across filter channels instead of time (
70cdec2).
Reference frontier (simulated, base rate 6.1%)
Why selection moved off balanced accuracy — measured at fixed signal strength:
| tau | balanced acc | call rate | precision |
|---|---|---|---|
| 0.00 | 34.0% | 0.2% | 27.3% |
| 0.15 | 34.2% | 0.2% | 33.3% |
| 0.35 | 36.3% | 2.0% | 15.5% |
| 1.00 | 53.5% | 49.6% | 6.4% |
Balanced accuracy rises monotonically as the model calls more and is right less. Deploy
order inverts from 1.00 > 0.50 > 0.35 under the old metric to 0.35 > 0.50 > 1.00 under
precision-led selection.