feat(ai): logit-adjusted loss, replacing oversampling and the post-hoc prior
Menon et al. 2021 (ICLR), "Long-tail learning via logit adjustment": add
tau*log(prior_c) to each class logit inside the training gradient. Softmax
CE on adjusted logits is consistent for BALANCED error - the metric
checkpoint selection already ranks on - so the loss and the deploy decision
finally optimize the same thing.
The engine already computed a true softmax + categorical-CE gradient and
wrote it over the per-neuron sigmoid delta, so this is an offset added to
three logits in the two places that gradient is built (backProp scalar path
and backPropOCL). No backend, kernel or DLL change; the forward pass and
every inference path are untouched, which is the point - the network learns
to absorb the offset, so its raw argmax becomes the balanced-optimal
decision with nothing applied at inference.
Replaces rather than stacks. Minority replay is disabled while this is on,
and the post-hoc inference prior is forced off. Stacking is not a
theoretical worry: simulated on the measured 1118/1119/34298 distribution
in the weak-signal regime, plain CE collapses to Neutral (33.4% balanced,
Buy 0%), replay reaches 48.1%, logit adjustment 50.9% with better balance -
and BOTH together score 45.4% with Neutral recall at 0%, worse than either
alone. Buda et al. 2018 predicts exactly that.
Motivation from the six-chart run: every topology took one direction to
~50% recall and abandoned the other, the direction chosen arbitrarily (the
batch-norm control went Buy 1% / Sell 42%, the inverse of the other five).
One era in 1,301 cleared the per-class recall floor.
Fingerprinted conditionally, so the converged 60.7% models on disk keep
their filenames and stay loadable as the fallback.
Both builds compile 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 19:05:14 -04:00
|
|
|
//+------------------------------------------------------------------+
|
2026-08-22 00:30:14 -04:00
|
|
|
//+------------------------------------------------------------------+
|
refactor: split CExpertSignalAIBase implementation by responsibility
ExpertSignalAIBase.mqh was 8216 lines: the class declaration followed by 87
method bodies covering training, labelling, feature extraction, persistence,
chart drawing, online learning, the GA auto-tuner and inference, all in one
file. Train() alone is 1492 lines; a change to arrow drawing meant scrolling
past the era loop.
Moved the bodies into Expert\AIBase\, included at the bottom of the original
after the class declaration:
Training.mqh 1607 era loop, plateau ladder, checkpoint select, deploy
Features.mqh 1093 indicator creation + per-bar input feature vector
ChartUI.mqh 634 arrows, arrow persistence, status panel, cleanup
Persistence.mqh 492 .stats/.cfg sidecars, CPU-inference validation, copy
OnlineLearning.mqh 461 live continual learning, EMA shadow, OOS simulator
Labels.mqh 309 ZigZag pivot labels, async label-cache prebuild
AutoTune.mqh 275 genetic tuner (population, crossover, halving)
Inference.mqh 235 softmax, prior calibration, class priors
ExpertSignalAIBase.mqh 8216 -> 3131 (declaration + topology build only)
This is a pure relocation - verified mechanically, not by eye: HEAD's file
reconstructed from the eight partials plus the surviving remainder is
byte-identical to HEAD, span for span (scratchpad verify_split.py). No
declaration moved, no signature changed, no code rewritten, so behaviour is
unchanged by construction.
Compiles 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 00:42:45 -04:00
|
|
|
//| Warrior_EA |
|
|
|
|
|
//| AnimateDread |
|
|
|
|
|
//| |
|
2026-08-22 00:30:14 -04:00
|
|
|
//| Read-time signal production: softmax, prior calibration, class p |
|
refactor: split CExpertSignalAIBase implementation by responsibility
ExpertSignalAIBase.mqh was 8216 lines: the class declaration followed by 87
method bodies covering training, labelling, feature extraction, persistence,
chart drawing, online learning, the GA auto-tuner and inference, all in one
file. Train() alone is 1492 lines; a change to arrow drawing meant scrolling
past the era loop.
Moved the bodies into Expert\AIBase\, included at the bottom of the original
after the class declaration:
Training.mqh 1607 era loop, plateau ladder, checkpoint select, deploy
Features.mqh 1093 indicator creation + per-bar input feature vector
ChartUI.mqh 634 arrows, arrow persistence, status panel, cleanup
Persistence.mqh 492 .stats/.cfg sidecars, CPU-inference validation, copy
OnlineLearning.mqh 461 live continual learning, EMA shadow, OOS simulator
Labels.mqh 309 ZigZag pivot labels, async label-cache prebuild
AutoTune.mqh 275 genetic tuner (population, crossover, halving)
Inference.mqh 235 softmax, prior calibration, class priors
ExpertSignalAIBase.mqh 8216 -> 3131 (declaration + topology build only)
This is a pure relocation - verified mechanically, not by eye: HEAD's file
reconstructed from the eight partials plus the surviving remainder is
byte-identical to HEAD, span for span (scratchpad verify_split.py). No
declaration moved, no signature changed, no code rewritten, so behaviour is
unchanged by construction.
Compiles 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 00:42:45 -04:00
|
|
|
//+------------------------------------------------------------------+
|
|
|
|
|
#ifndef WARRIOR_AIBASE_INFERENCE_MQH
|
|
|
|
|
#define WARRIOR_AIBASE_INFERENCE_MQH
|
|
|
|
|
//+------------------------------------------------------------------+
|
|
|
|
|
//| Post-convergence "new bar" handler - see ScheduleTrainingIfNeeded()|
|
|
|
|
|
//| for why this exists: once m_trainingComplete is true, a plain new |
|
|
|
|
|
//| bar must NOT re-enter Train()'s full era loop (which resets the |
|
2026-08-20 07:00:09 -04:00
|
|
|
//| best-checkpoint/g_eta-decay tracking and runs real Net.backProp() |
|
refactor: split CExpertSignalAIBase implementation by responsibility
ExpertSignalAIBase.mqh was 8216 lines: the class declaration followed by 87
method bodies covering training, labelling, feature extraction, persistence,
chart drawing, online learning, the GA auto-tuner and inference, all in one
file. Train() alone is 1492 lines; a change to arrow drawing meant scrolling
past the era loop.
Moved the bodies into Expert\AIBase\, included at the bottom of the original
after the class declaration:
Training.mqh 1607 era loop, plateau ladder, checkpoint select, deploy
Features.mqh 1093 indicator creation + per-bar input feature vector
ChartUI.mqh 634 arrows, arrow persistence, status panel, cleanup
Persistence.mqh 492 .stats/.cfg sidecars, CPU-inference validation, copy
OnlineLearning.mqh 461 live continual learning, EMA shadow, OOS simulator
Labels.mqh 309 ZigZag pivot labels, async label-cache prebuild
AutoTune.mqh 275 genetic tuner (population, crossover, halving)
Inference.mqh 235 softmax, prior calibration, class priors
ExpertSignalAIBase.mqh 8216 -> 3131 (declaration + topology build only)
This is a pure relocation - verified mechanically, not by eye: HEAD's file
reconstructed from the eight partials plus the surviving remainder is
byte-identical to HEAD, span for span (scratchpad verify_split.py). No
declaration moved, no signature changed, no code rewritten, so behaviour is
unchanged by construction.
Compiles 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 00:42:45 -04:00
|
|
|
//| passes again, silently perturbing an already-converged model |
|
|
|
|
|
//| forever, once per bar, with no way to ever actually finish). This |
|
|
|
|
|
//| only refreshes the price/indicator buffers and re-runs inference |
|
|
|
|
|
//| for the newest bar so dPrevSignal/the chart arrow stay current - |
|
|
|
|
|
//| identical cost to what Train() does per-bar, minus every bit of |
|
2026-08-01 11:27:28 -04:00
|
|
|
//| training (label caching, backProp, checkpointing). |
|
refactor: split CExpertSignalAIBase implementation by responsibility
ExpertSignalAIBase.mqh was 8216 lines: the class declaration followed by 87
method bodies covering training, labelling, feature extraction, persistence,
chart drawing, online learning, the GA auto-tuner and inference, all in one
file. Train() alone is 1492 lines; a change to arrow drawing meant scrolling
past the era loop.
Moved the bodies into Expert\AIBase\, included at the bottom of the original
after the class declaration:
Training.mqh 1607 era loop, plateau ladder, checkpoint select, deploy
Features.mqh 1093 indicator creation + per-bar input feature vector
ChartUI.mqh 634 arrows, arrow persistence, status panel, cleanup
Persistence.mqh 492 .stats/.cfg sidecars, CPU-inference validation, copy
OnlineLearning.mqh 461 live continual learning, EMA shadow, OOS simulator
Labels.mqh 309 ZigZag pivot labels, async label-cache prebuild
AutoTune.mqh 275 genetic tuner (population, crossover, halving)
Inference.mqh 235 softmax, prior calibration, class priors
ExpertSignalAIBase.mqh 8216 -> 3131 (declaration + topology build only)
This is a pure relocation - verified mechanically, not by eye: HEAD's file
reconstructed from the eight partials plus the surviving remainder is
byte-identical to HEAD, span for span (scratchpad verify_split.py). No
declaration moved, no signature changed, no code rewritten, so behaviour is
unchanged by construction.
Compiles 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 00:42:45 -04:00
|
|
|
//+------------------------------------------------------------------+
|
|
|
|
|
void CExpertSignalAIBase::RefreshConvergedSignal(void)
|
|
|
|
|
{
|
feat(ensemble): per-NN inputs replace the preset selector - the meta head becomes the vote's gate
User design (2026-08-19): 'remove the enum menu that selects neural networks... individual
inputs for every NN just like classic signals... the META NN should be integrated into the
voting decision pipeline when enabled... as a bonus meta labelling is applied to enabled NNs.'
- AI_CHOICE is GONE (tombstoned per the stale-.set doctrine). Use_MLP/Use_CONV/Use_LSTM/
Use_CONVLSTM are ordinary bools like the classic votes; the ensemble arithmetic adapts to
any subset because the consensus divisor is the enabled capable weight. Two or more
enabled = ensemble (|ENS1 token + joint gate, exactly the old AI_HYBRID fingerprints, so
existing weight files keep loading); one = the old solo preset; none = classic-only.
- Use_MetaLabeling un-couples META from the direction NNs (the old selector made them
mutually exclusive). S3 ships: CSignalMETA::LiveMetaGate scores each vote-cleared entry
(shared window at bar 1 + proposal descriptor: side, net vote, live geometry, spread/ATR;
pattern one-hot ZEROED - ranking, not calibrated probability, documented in the body) and
vetoes below the cost-adjusted break-even. Entries only; fail-open everywhere, loudly.
- COEXISTENCE HAZARDS closed: VoteCapableWeight()=0 and ProspectiveVote()=false for the
meta target - solo-only until today, a trained META would otherwise sit in the consensus
divisor as a permanent abstainer and shrink every vote by its module weight.
- CERTIFIED == TRADED: the ensemble era verdict replays the identical veto through the same
g_warriorMetaGate pointer over its OOS fired bars (bar re-resolved from the row's own
time; fail-open counted as fires and reported: 'metaGate: N approved, M vetoed, K
unscored'). The overlay deliberately does NOT replay it (veto-filter-in-replay class,
calendar-cliff precedent) - documented at the sweep site. Solo charts' own gate does not
model the veto - the standing solo-gate caveat, documented at the input.
- DB continuity: the pattern/journal DB fingerprint's first slot was (int)AIType;
DbLegacyAiSlot() maps every legacy-expressible config to its OLD value (new 2-3 member
subsets get 100+bitmask, outside the legacy range) so no existing database re-keys.
filterID becomes the enabled roster via one EnabledNNSummary().
- HUD: the meta line shows the gate (armed/(trn), last P vs BE, ok/veto tally); the
armed/disarmed announcement fires on state change via one latch (MetaGateArmedNow), not
only when an entry happens to be proposed.
NOT COMPILED - user compiles in MetaEditor.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-19 13:01:02 -04:00
|
|
|
//--- Meta target: the meta head scores PROPOSED TRADES, not a bare bar window - a candidate-less
|
2026-08-22 00:25:52 -04:00
|
|
|
//--- forward would also be width-mismatched against its input layer (window + descriptor).
|
feat: S2 meta-labeling head - binary trade-quality model over the classic-candidate corpus
The NN now has a target that is not per-bar direction (closed, best-of-999
p=1.0000): P(win | this journaled candidate, at the EA's own SL/TP, net of
cost). One net for all 52 pattern-sides, AIType=AI_META.
- NetForward.mqh: the host-side softmax+CE gradient generalized total==3 ->
2||3 on both backprop paths; a 2-class softmax IS a logistic head, and no
compute backend changes.
- SignalMETA.mqh (new): corpus loaded read-only from the LARGEST signal DB on
disk (decoupled from the config fingerprint that burned four S1 runs); the
GMT->server offset is measured PER ROW against entryPrice vs bar open
(DST-immune, histogram logged); a window-span regime filter drops the
pre-2017 daily-backfill rows; 31-feature setup descriptor appended at the
input (26 one-hot + side + tanh netVote + SL/TP ATR + spread/ATR).
- Training.mqh: candidate-queued pass 1, binary-target pass 2, per-candidate
calibration (2.5) and OOS (3) walks. Counter mapping win->Buy / loss->Sell
lets checkpoint selection, the edge floor, the plateau ladder and the
family-wise deploy gate run UNCHANGED: precision reads as win rate among
traded candidates, chance as the base win rate, recalls as sensitivity/
specificity. Era-end META line: coverage x (p - break-even) vs the null.
- Labels are the side-conditional triple-barrier win caches - never the DB's
stop-and-reverse outcome. Logit adjustment deliberately skipped (~40% base
rate). Live inference + online learning guarded off until S3.
- Fingerprint: conditional |TGT:META1; State\META\ folder + 2-output filename
slot keep meta models fully separate from direction models.
Compiles clean (0 errors, 0 warnings). S2 run = attach a chart with
AIType=AI_META; S3 wires the votes via the per-side hooks.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-13 06:52:31 -04:00
|
|
|
if(IsMetaTarget())
|
|
|
|
|
return;
|
2026-08-22 00:25:52 -04:00
|
|
|
//--- Size the buffers from what the FEATURE BUILDER actually needs, not from a date delta. The
|
|
|
|
|
//--- result was silent - no error, no short window, just inference computing DIFFERENT features
|
|
|
|
|
//--- from the ones training learned on.
|
feat(ai): triple-barrier labels replace exact-pivot ZigZag targets
The 31:1 class imbalance was self-inflicted by the TARGET, not a property
of the market. Labelling only the exact bar where a ZigZag pivot confirms
gave Buy 1164 / Sell 1164 / Neutral 35841, and every correction mechanism
this codebase accumulated sits downstream of that one choice: the
logit-adjusted loss and its range cap, the prior EMA, the +-3.0 output-bias
seed, balanced-accuracy-then-precision selection with its coverage floor,
the recall floor and its catch-22, the alternation gate, NMS, and the four
oversampling designs that collapsed before them.
The reference this engine is built on (references/neuronetworksbook.pdf
ch. 3.1/3.3) also uses ZigZag, but targets the DIRECTION TO THE NEXT
EXTREMUM on every bar - ~50/50 by construction, with no imbalance to
correct at all. It never had this problem because it never asked "is this
the pivot bar".
Labels are now the triple barrier (Lopez de Prado ch. 3), using the EA's
OWN SL_Mode/TP_Mode: does a trade opened at this bar's close reach its
target before its stop, within a horizon. Buy = long resolves, Sell =
short resolves, Neutral = neither. Consequences:
- dir-precision in the era line stops being a proxy and becomes the win
rate of the strategy under its own exit rules.
- Expected balance ~25/25/50 at the shipped 1:3 (gambler's ruin), i.e.
~2:1 instead of 31:1. Measured and logged at the end of the prebuild.
- Spread is charged on both legs, so it is a NET win rate.
- Intrabar ambiguity resolves to the STOP. OHLC cannot order two touches
inside one bar and the optimistic reading is how a backtested edge
becomes a live loss.
ZigZag stays as input features (EnableSwingContext) and now also supplies
the vertical barrier: the horizon is the median confirmed leg length,
snapped to a coarse ladder. Derived, not configured, and deliberately kept
out of the filename fingerprint - a filename keyed on a measured quantity
orphans a trained model the moment the measurement moves.
Removed, because the premise died with the old target:
- the alternation gate. Correct for pivot labels (a ZigZag cannot emit two
same-type pivots in a row, so a repeat was provably a false fire), and
wrong for barrier labels, which answer each bar independently. It also
took its worst consequence with it: a one-sided model previously got ONE
trade per backtest, a hard blocker on marketplace validation.
- SignalClusterWindow now defaults off - it de-duplicated repeats that are
now real trades. Kept as an opt-in display control.
- LABEL_WINDOW_BARS, the pivot-widening pass, ConfirmedZigZagLabel.
- the era-0 output-bias seed now needs a genuinely dominant class (0.70)
rather than 0.40; at ~50% Neutral a +-3.0 seed is a distortion, not a
correction.
Also fixed, both found while wiring the above:
1. RefreshConvergedSignal sized its buffers from a date delta
(Bars(sym, period, dtStudied, TimeCurrent())). dtStudied is a training
watermark; in the tester it is loaded from a live-chart save AHEAD of
the simulated date, so the interval inverted, Bars() returned ~0, and
the buffer came out at exactly m_historyBars - deep enough for the OHLC
window and far too shallow for the Donchian-50 / 20-bar-return / SMA
extension behind it. Inference silently computed DIFFERENT features
from the ones training learned on, live as well as in the tester. Now
sized from what the feature builder actually needs.
2. The barrier horizon is resolved on the deployed path too. A deployed
model never enters Train(), so it never reached the prebuild, and
OnlineLearnStep reads the horizon as its confirmation delay - left at
the fallback it would have backpropped bars whose barriers had not
resolved. Silent lookahead in the one place that writes to a live model.
SL_Mode/TP_Mode join the weights fingerprint: they define the labels now,
so a model trained at 1:3 must never be silently reused at 1:1. This
re-keys every pre-existing model by design - none were trained on this task.
Inference census extended with the vote gate. LongCondition/ShortCondition
open with a readiness check the refresh counters never see; in the tester it
reduces to "the seeded _optcache.nnw must have LOADED", and if it did not,
every vote is hard-zeroed while the model still answers Buy. The old three
counters would have read that as "the model says Neutral" - false, and a
completely different fix. This is the leading candidate for the
zero-direction backtest and the census can now name it in one run.
Both builds compile 0 errors / 0 warnings. Forces a full retrain.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 20:39:49 -04:00
|
|
|
int need = (int)m_historyBars + SWING_SCAN_CAP_BARS + MathMax(m_barrierHorizonBars, 1) + 2;
|
|
|
|
|
int barsNow = (int)MathMin(need, Bars(m_symbol.Name(), PERIOD_CURRENT));
|
2026-08-22 00:25:52 -04:00
|
|
|
//--- HOLD RATHER THAN TRADE ON A SHORT WINDOW. `need` is the depth the feature builder requires
|
|
|
|
|
//--- for inference features to match the ones training learned on; a shallower buffer does not
|
|
|
|
|
//--- fail, it makes the swing block take its graceful degraded path - which is precisely the
|
|
|
|
|
//--- silent feature-mismatch this whole `need` calculation was introduced to end (see above).
|
fix(depth): route EVERY ResizeBuffers call site through one indicator-depth gate
1dda479 clamped the training sweep. It left five other paths asking the indicators
for a depth they cannot serve, and on a live account the quiet ones are worse than
the stall was - a stalled chart is visible, a chart trading on a degraded feature
window is not.
ServableBars(want, context) is now the single gate, and all six go through it:
training sweep clamp, floored at TRAIN_MIN_CLAMPED_BARS (below that a small
positive BarsCalculated is warm-up, which m_coldSweepTick owns)
label prebuild clamp - labels come from price/ADZigZag and would survive a
capped MA, but ResizeBuffers sizes EVERY buffer and a failed
CopyBuffer leaves m_MA EMPTY for the next reader, so this path
could silently re-break the block Train()'s clamp just fixed
live inference HOLD. Below `need` the swing block takes its degraded path and
inference runs on a different feature distribution than the model
was fitted on. This EA sizes real positions off that output, so
no signal beats a mismatched one
online learning HOLD, same reason and worse - this path WRITES to a live trading
model, so a mismatched (features, label) pair is not a wrong arrow,
it is a wrong weight update that compounds every bar
chart rescan clamp - SIGNAL_RESCAN_LOOKBACK_BARS is 5000 and MT5's smallest
"Max bars in chart" is also 5000, so this one is genuinely
reachable; uncapped it repaints the window all-Neutral
research export clamp before the emptiness test, so a capped symbol exports the
depth it has rather than writing a CSV with a dead feature block -
an artefact that looks complete and is silently wrong
Both HOLDs are insurance, not expected states: `need` tops out near 1,152 bars
(16 + 750 + 384 + 2) against a 5,000 floor on the terminal setting. They exist so
the failure mode is unreachable rather than merely unlikely.
Not changed: a genuinely SHORT price history still takes the old degraded path at
every site. That is pre-existing behaviour and narrowing it would mute charts that
trade today, so it stays a separate decision rather than a side effect of this fix.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 13:33:16 -04:00
|
|
|
int servable = ServableBars(need, "live inference");
|
|
|
|
|
if(servable < need)
|
|
|
|
|
{
|
|
|
|
|
if(!m_inferenceDepthRefusalWarned)
|
|
|
|
|
{
|
|
|
|
|
m_inferenceDepthRefusalWarned = true;
|
|
|
|
|
PrintFormat("%s: LIVE INFERENCE HELD - the feature window needs %d bars and the indicators can"
|
|
|
|
|
" only serve %d. Computing a signal here would silently use the swing block's"
|
|
|
|
|
" degraded path, i.e. different features from the ones this model was trained on,"
|
|
|
|
|
" so no signal is emitted until the depth is available. See the indicator-cap line"
|
|
|
|
|
" above for how to raise it.", ID, need, servable);
|
|
|
|
|
}
|
|
|
|
|
return;
|
|
|
|
|
}
|
|
|
|
|
m_inferenceDepthRefusalWarned = false;
|
refactor: split CExpertSignalAIBase implementation by responsibility
ExpertSignalAIBase.mqh was 8216 lines: the class declaration followed by 87
method bodies covering training, labelling, feature extraction, persistence,
chart drawing, online learning, the GA auto-tuner and inference, all in one
file. Train() alone is 1492 lines; a change to arrow drawing meant scrolling
past the era loop.
Moved the bodies into Expert\AIBase\, included at the bottom of the original
after the class declaration:
Training.mqh 1607 era loop, plateau ladder, checkpoint select, deploy
Features.mqh 1093 indicator creation + per-bar input feature vector
ChartUI.mqh 634 arrows, arrow persistence, status panel, cleanup
Persistence.mqh 492 .stats/.cfg sidecars, CPU-inference validation, copy
OnlineLearning.mqh 461 live continual learning, EMA shadow, OOS simulator
Labels.mqh 309 ZigZag pivot labels, async label-cache prebuild
AutoTune.mqh 275 genetic tuner (population, crossover, halving)
Inference.mqh 235 softmax, prior calibration, class priors
ExpertSignalAIBase.mqh 8216 -> 3131 (declaration + topology build only)
This is a pure relocation - verified mechanically, not by eye: HEAD's file
reconstructed from the eight partials plus the surviving remainder is
byte-identical to HEAD, span for span (scratchpad verify_split.py). No
declaration moved, no signature changed, no code rewritten, so behaviour is
unchanged by construction.
Compiles 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 00:42:45 -04:00
|
|
|
if(!ResizeBuffers(barsNow) || !RefreshData())
|
|
|
|
|
return;
|
2026-08-22 00:25:52 -04:00
|
|
|
//--- INVALIDATE THE NOW-RELATIVE BAR CACHES. Non-obvious and load-bearing: the feature cache is
|
|
|
|
|
//--- keyed by MQL5 series index, and index 0 means "newest bar", so every closed candle shifts
|
|
|
|
|
//--- what every cached row stands for.
|
fix: the sequence models were reading the window backwards
BuildFeatureWindow() replaces eight hand-rolled copies of the same loop
and feeds the window OLDEST BAR FIRST. Every copy fed it newest-first,
because MQL5 timeseries indices run backwards and `r + b` with b ascending
walks into the past.
Harmless for PAI and CONV - a dense layer learns a weight per position
either way, a conv learns time-mirrored kernels. Not harmless for the
recurrent stacks:
- LSTM_SeqStepForward reads `inputs + t*Iw`, so step t is block t.
- It writes output[] only when t == steps-1: the visible output IS the
last hidden state.
- c_t = f*c_{t-1} + i*g decays toward the start of the sequence.
lstm_seq_flowcheck.cpp measured block 0's influence on the output at
1.2e-2 of block T-1's, at the shipped forget bias of 1.0.
So the bar being PREDICTED sat at the far end of the decay and the output
was handed to the OLDEST bar in the window - the exact inverse of what the
window is for. ~80x backwards on LSTM and HYBRID, on all three tiers
(OpenCL kernel, CPU DLL, pure-MQL5 inference), which is why it never
surfaced as a backend discrepancy.
This does not create edge - the MI diagnostics read at the noise floor
(p=0.4975) with a working positive control. It makes the one hypothesis
those diagnostics explicitly do NOT cover testable: they are marginal and
per-bar, and state they "cannot rule out one that only exists in
combination or across time". The sequence model is the instrument for
across-time structure and it has been crippled, so that hypothesis has
never been honestly tested.
Fingerprint gets an unconditional |WIN:2 - the vector keeps its shape and
its features, so a stale .nnw would load cleanly and run a model fitted to
one ordering against the other, silently. Re-keying every config is the
point, not collateral damage. FORCES A FULL RETRAIN.
Also: the now-relative bar caches are re-keyed on the two live paths.
EnsureBarCachesCapacity() was only ever called from training paths, but
once m_trainingComplete is set ScheduleTrainingIfNeeded() routes every bar
to RefreshConvergedSignal() and Train() is never re-entered - so nothing
cleared the feature cache again for the life of the process. A chart that
trained to convergence kept replaying the rows computed for the last
training era's bar grid: the live signal froze at its convergence-time
value, and OnlineLearnStep() backpropped those stale features against
freshly resolved labels. Backtests were never affected (an inference-only
process never allocates the arrays, so every read recomputes).
Compiles clean: 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-07 18:28:44 -04:00
|
|
|
EnsureBarCachesCapacity(barsNow);
|
2026-08-22 00:25:52 -04:00
|
|
|
//--- Same bar grid, same panel. Only as deep as inference actually reads. Asking for the full
|
|
|
|
|
//--- `barsNow` here would rebuild a training-depth panel on EVERY bar, which in the tester means
|
|
|
|
|
//--- one full multi-symbol resample per simulated bar.
|
feat(ai): spread as a volatility-regime feature, and fix a stale-index cache in both new blocks
Adds spread/ATR and the spread change ratio as network inputs (EnableSpreadFeature,
default on). Spread is the one microstructure channel that is both FX-available and
genuinely historical in the Strategy Tester - "during testing, the spread is not modeled
but is taken from historical data" - so unlike swap, signed tick flow or depth of market it
is something a backtest can honestly validate.
What it encodes, stated precisely because the raw measurement overstates it.
research/test_spread.py found spr/atr the strongest single feature in this codebase, on 5
of 8 instrument/geometry cells at 2-4x any volume feature. But the barrier LABEL charges
the spread inside its own barriers, so a wide-spread bar is mechanically likelier to
resolve as a loss and the feature would partly be predicting its own cost model. Relabelling
at zero cost and re-measuring the identical feature showed 20-40% of it WAS that tautology
and the majority was not (XAUUSD retained 97%). What survives is a volatility-regime
reading: spread is near-fixed while ATR is not, so the ratio runs high exactly when
realised volatility is below its own ATR estimate, which genuinely predicts whether
ATR-scaled barriers get reached. It is UNSIGNED - Neutral-vs-directional only, never a side.
Also fixes a stale-index bug I introduced with the cross-asset panel and had just repeated
in the spread series. Both cached on length alone:
if(m_crossAsset.Bars() >= bars) return true;
MQL5 series indices are relative to NOW, so one new closed candle shifts every index by
one. Keyed only on length, the panel keeps serving its index 0 as a bar that is no longer
the newest, and every cross-asset value is read one bar out of step with the price features
sitting beside it in the same vector - silently, with no error and no shape change. This is
the same class of defect as the dtStudied watermark behind the zero-direction backtests.
Both now carry a datetime anchor on m_Time.GetData(0), the same invalidation key the
label/feature bar caches already use.
And a performance fix that fell out of it: with correct invalidation the panel rebuilds on
every new bar, and RefreshConvergedSignal runs per bar - which in the tester would mean one
full multi-symbol resample per simulated bar at training depth. Inference only reads bars
0..m_historyBars-1 plus the panel's own slow window, so it now requests exactly that. The
cache check is >=, so a deeper panel left from training still satisfies it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 17:42:40 -04:00
|
|
|
BuildCrossAssetPanel((int)m_historyBars + CROSSASSET_SLOW_BARS + 2);
|
|
|
|
|
EnsureSpreadSeries(barsNow);
|
feat(ai): triple-barrier labels replace exact-pivot ZigZag targets
The 31:1 class imbalance was self-inflicted by the TARGET, not a property
of the market. Labelling only the exact bar where a ZigZag pivot confirms
gave Buy 1164 / Sell 1164 / Neutral 35841, and every correction mechanism
this codebase accumulated sits downstream of that one choice: the
logit-adjusted loss and its range cap, the prior EMA, the +-3.0 output-bias
seed, balanced-accuracy-then-precision selection with its coverage floor,
the recall floor and its catch-22, the alternation gate, NMS, and the four
oversampling designs that collapsed before them.
The reference this engine is built on (references/neuronetworksbook.pdf
ch. 3.1/3.3) also uses ZigZag, but targets the DIRECTION TO THE NEXT
EXTREMUM on every bar - ~50/50 by construction, with no imbalance to
correct at all. It never had this problem because it never asked "is this
the pivot bar".
Labels are now the triple barrier (Lopez de Prado ch. 3), using the EA's
OWN SL_Mode/TP_Mode: does a trade opened at this bar's close reach its
target before its stop, within a horizon. Buy = long resolves, Sell =
short resolves, Neutral = neither. Consequences:
- dir-precision in the era line stops being a proxy and becomes the win
rate of the strategy under its own exit rules.
- Expected balance ~25/25/50 at the shipped 1:3 (gambler's ruin), i.e.
~2:1 instead of 31:1. Measured and logged at the end of the prebuild.
- Spread is charged on both legs, so it is a NET win rate.
- Intrabar ambiguity resolves to the STOP. OHLC cannot order two touches
inside one bar and the optimistic reading is how a backtested edge
becomes a live loss.
ZigZag stays as input features (EnableSwingContext) and now also supplies
the vertical barrier: the horizon is the median confirmed leg length,
snapped to a coarse ladder. Derived, not configured, and deliberately kept
out of the filename fingerprint - a filename keyed on a measured quantity
orphans a trained model the moment the measurement moves.
Removed, because the premise died with the old target:
- the alternation gate. Correct for pivot labels (a ZigZag cannot emit two
same-type pivots in a row, so a repeat was provably a false fire), and
wrong for barrier labels, which answer each bar independently. It also
took its worst consequence with it: a one-sided model previously got ONE
trade per backtest, a hard blocker on marketplace validation.
- SignalClusterWindow now defaults off - it de-duplicated repeats that are
now real trades. Kept as an opt-in display control.
- LABEL_WINDOW_BARS, the pivot-widening pass, ConfirmedZigZagLabel.
- the era-0 output-bias seed now needs a genuinely dominant class (0.70)
rather than 0.40; at ~50% Neutral a +-3.0 seed is a distortion, not a
correction.
Also fixed, both found while wiring the above:
1. RefreshConvergedSignal sized its buffers from a date delta
(Bars(sym, period, dtStudied, TimeCurrent())). dtStudied is a training
watermark; in the tester it is loaded from a live-chart save AHEAD of
the simulated date, so the interval inverted, Bars() returned ~0, and
the buffer came out at exactly m_historyBars - deep enough for the OHLC
window and far too shallow for the Donchian-50 / 20-bar-return / SMA
extension behind it. Inference silently computed DIFFERENT features
from the ones training learned on, live as well as in the tester. Now
sized from what the feature builder actually needs.
2. The barrier horizon is resolved on the deployed path too. A deployed
model never enters Train(), so it never reached the prebuild, and
OnlineLearnStep reads the horizon as its confirmation delay - left at
the fallback it would have backpropped bars whose barriers had not
resolved. Silent lookahead in the one place that writes to a live model.
SL_Mode/TP_Mode join the weights fingerprint: they define the labels now,
so a model trained at 1:3 must never be silently reused at 1:1. This
re-keys every pre-existing model by design - none were trained on this task.
Inference census extended with the vote gate. LongCondition/ShortCondition
open with a readiness check the refresh counters never see; in the tester it
reduces to "the seeded _optcache.nnw must have LOADED", and if it did not,
every vote is hard-zeroed while the model still answers Buy. The old three
counters would have read that as "the model says Neutral" - false, and a
completely different fix. This is the leading candidate for the
zero-direction backtest and the census can now name it in one run.
Both builds compile 0 errors / 0 warnings. Forces a full retrain.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 20:39:49 -04:00
|
|
|
//--- A deployed model never enters Train(), so this is the only place its barrier horizon gets
|
|
|
|
|
//--- measured - and OnlineLearnStep() below depends on it being right. First call sizes buffers
|
|
|
|
|
//--- against the fallback, which is harmless: `need` is dominated by SWING_SCAN_CAP_BARS either way.
|
|
|
|
|
EnsureBarrierHorizon(barsNow);
|
fix: live inference queried the 1-tick forming bar - a window training never built
RefreshLatestSignal ran at the first tick after a bar opens and built its
window at r=0: series index 0 at that instant is a candle with one tick of
data - (close-open)/atr ~ 0, high ~ low, degenerate volume, indicators on a
1-tick bar. Training never produces such a window (every labeled bar is fully
closed, entry at that bar's CLOSE), so the deployed model's final timestep -
the one the LSTM/HYBRID output is keyed to - was out-of-distribution on every
live decision, and pass 3's deploy-gate OOS scores measured a different query
than live executed. The parity index is r=1: the newest CLOSED bar, whose
close IS the current price - the exact instant the label's hypothetical entry
happens. Single backtests shared the old skew (same r=0), which is why the
tester agreed with live while both disagreed with training.
Bookkeeping split that the index change forces: m_lastBarTime/dtStudied stay
anchored to the FORMING bar's open (they gate against SERIES_LASTBAR_DATE;
anchoring at bar 1 would re-fire the refresh every tick), while bt - the
arrow, its High/Low placement, and NMS declustering - anchors to the decision
bar, now matching the rescan path's convention.
Also: a failed refresh no longer trades the previous bar's signal for the
whole bar. RefreshLatestSignal returns success, zeroes dPrevSignal on failure
(no opinion beats a stale one), and RefreshConvergedSignal advances dtStudied
only on success so the next tick retries - the tester path (m_lastBarTime)
already worked this way; this is the live path catching up.
FORCES RE-VALIDATION of deployed models: the effective live query distribution
changes. Bundled with the backprop transpose fix's retrain.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-11 18:10:23 -04:00
|
|
|
bool refreshed = RefreshLatestSignal();
|
2026-08-22 00:25:52 -04:00
|
|
|
//--- Continual learning: on a LIVE chart (never the tester/optimizer - OnlineLearnStep() self-
|
|
|
|
|
//--- guards on m_inferenceOnly) a deployed model keeps adapting to newly-confirmed structure.
|
refactor: split CExpertSignalAIBase implementation by responsibility
ExpertSignalAIBase.mqh was 8216 lines: the class declaration followed by 87
method bodies covering training, labelling, feature extraction, persistence,
chart drawing, online learning, the GA auto-tuner and inference, all in one
file. Train() alone is 1492 lines; a change to arrow drawing meant scrolling
past the era loop.
Moved the bodies into Expert\AIBase\, included at the bottom of the original
after the class declaration:
Training.mqh 1607 era loop, plateau ladder, checkpoint select, deploy
Features.mqh 1093 indicator creation + per-bar input feature vector
ChartUI.mqh 634 arrows, arrow persistence, status panel, cleanup
Persistence.mqh 492 .stats/.cfg sidecars, CPU-inference validation, copy
OnlineLearning.mqh 461 live continual learning, EMA shadow, OOS simulator
Labels.mqh 309 ZigZag pivot labels, async label-cache prebuild
AutoTune.mqh 275 genetic tuner (population, crossover, halving)
Inference.mqh 235 softmax, prior calibration, class priors
ExpertSignalAIBase.mqh 8216 -> 3131 (declaration + topology build only)
This is a pure relocation - verified mechanically, not by eye: HEAD's file
reconstructed from the eight partials plus the surviving remainder is
byte-identical to HEAD, span for span (scratchpad verify_split.py). No
declaration moved, no signature changed, no code rewritten, so behaviour is
unchanged by construction.
Compiles 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 00:42:45 -04:00
|
|
|
OnlineLearnStep();
|
2026-08-22 00:25:52 -04:00
|
|
|
//--- Advance the live new-bar watermark ONLY on success. On failure the gate stays open, so the
|
|
|
|
|
//--- next tick retries.
|
fix: live inference queried the 1-tick forming bar - a window training never built
RefreshLatestSignal ran at the first tick after a bar opens and built its
window at r=0: series index 0 at that instant is a candle with one tick of
data - (close-open)/atr ~ 0, high ~ low, degenerate volume, indicators on a
1-tick bar. Training never produces such a window (every labeled bar is fully
closed, entry at that bar's CLOSE), so the deployed model's final timestep -
the one the LSTM/HYBRID output is keyed to - was out-of-distribution on every
live decision, and pass 3's deploy-gate OOS scores measured a different query
than live executed. The parity index is r=1: the newest CLOSED bar, whose
close IS the current price - the exact instant the label's hypothetical entry
happens. Single backtests shared the old skew (same r=0), which is why the
tester agreed with live while both disagreed with training.
Bookkeeping split that the index change forces: m_lastBarTime/dtStudied stay
anchored to the FORMING bar's open (they gate against SERIES_LASTBAR_DATE;
anchoring at bar 1 would re-fire the refresh every tick), while bt - the
arrow, its High/Low placement, and NMS declustering - anchors to the decision
bar, now matching the rescan path's convention.
Also: a failed refresh no longer trades the previous bar's signal for the
whole bar. RefreshLatestSignal returns success, zeroes dPrevSignal on failure
(no opinion beats a stale one), and RefreshConvergedSignal advances dtStudied
only on success so the next tick retries - the tester path (m_lastBarTime)
already worked this way; this is the live path catching up.
FORCES RE-VALIDATION of deployed models: the effective live query distribution
changes. Bundled with the backprop transpose fix's retrain.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-11 18:10:23 -04:00
|
|
|
if(refreshed)
|
|
|
|
|
dtStudied = m_Time.GetData(0);
|
refactor: split CExpertSignalAIBase implementation by responsibility
ExpertSignalAIBase.mqh was 8216 lines: the class declaration followed by 87
method bodies covering training, labelling, feature extraction, persistence,
chart drawing, online learning, the GA auto-tuner and inference, all in one
file. Train() alone is 1492 lines; a change to arrow drawing meant scrolling
past the era loop.
Moved the bodies into Expert\AIBase\, included at the bottom of the original
after the class declaration:
Training.mqh 1607 era loop, plateau ladder, checkpoint select, deploy
Features.mqh 1093 indicator creation + per-bar input feature vector
ChartUI.mqh 634 arrows, arrow persistence, status panel, cleanup
Persistence.mqh 492 .stats/.cfg sidecars, CPU-inference validation, copy
OnlineLearning.mqh 461 live continual learning, EMA shadow, OOS simulator
Labels.mqh 309 ZigZag pivot labels, async label-cache prebuild
AutoTune.mqh 275 genetic tuner (population, crossover, halving)
Inference.mqh 235 softmax, prior calibration, class priors
ExpertSignalAIBase.mqh 8216 -> 3131 (declaration + topology build only)
This is a pure relocation - verified mechanically, not by eye: HEAD's file
reconstructed from the eight partials plus the surviving remainder is
byte-identical to HEAD, span for span (scratchpad verify_split.py). No
declaration moved, no signature changed, no code rewritten, so behaviour is
unchanged by construction.
Compiles 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 00:42:45 -04:00
|
|
|
}
|
|
|
|
|
//+------------------------------------------------------------------+
|
|
|
|
|
//| |
|
|
|
|
|
//+------------------------------------------------------------------+
|
fix: live inference queried the 1-tick forming bar - a window training never built
RefreshLatestSignal ran at the first tick after a bar opens and built its
window at r=0: series index 0 at that instant is a candle with one tick of
data - (close-open)/atr ~ 0, high ~ low, degenerate volume, indicators on a
1-tick bar. Training never produces such a window (every labeled bar is fully
closed, entry at that bar's CLOSE), so the deployed model's final timestep -
the one the LSTM/HYBRID output is keyed to - was out-of-distribution on every
live decision, and pass 3's deploy-gate OOS scores measured a different query
than live executed. The parity index is r=1: the newest CLOSED bar, whose
close IS the current price - the exact instant the label's hypothetical entry
happens. Single backtests shared the old skew (same r=0), which is why the
tester agreed with live while both disagreed with training.
Bookkeeping split that the index change forces: m_lastBarTime/dtStudied stay
anchored to the FORMING bar's open (they gate against SERIES_LASTBAR_DATE;
anchoring at bar 1 would re-fire the refresh every tick), while bt - the
arrow, its High/Low placement, and NMS declustering - anchors to the decision
bar, now matching the rescan path's convention.
Also: a failed refresh no longer trades the previous bar's signal for the
whole bar. RefreshLatestSignal returns success, zeroes dPrevSignal on failure
(no opinion beats a stale one), and RefreshConvergedSignal advances dtStudied
only on success so the next tick retries - the tester path (m_lastBarTime)
already worked this way; this is the live path catching up.
FORCES RE-VALIDATION of deployed models: the effective live query distribution
changes. Bundled with the backprop transpose fix's retrain.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-11 18:10:23 -04:00
|
|
|
bool CExpertSignalAIBase::RefreshLatestSignal(void)
|
refactor: split CExpertSignalAIBase implementation by responsibility
ExpertSignalAIBase.mqh was 8216 lines: the class declaration followed by 87
method bodies covering training, labelling, feature extraction, persistence,
chart drawing, online learning, the GA auto-tuner and inference, all in one
file. Train() alone is 1492 lines; a change to arrow drawing meant scrolling
past the era loop.
Moved the bodies into Expert\AIBase\, included at the bottom of the original
after the class declaration:
Training.mqh 1607 era loop, plateau ladder, checkpoint select, deploy
Features.mqh 1093 indicator creation + per-bar input feature vector
ChartUI.mqh 634 arrows, arrow persistence, status panel, cleanup
Persistence.mqh 492 .stats/.cfg sidecars, CPU-inference validation, copy
OnlineLearning.mqh 461 live continual learning, EMA shadow, OOS simulator
Labels.mqh 309 ZigZag pivot labels, async label-cache prebuild
AutoTune.mqh 275 genetic tuner (population, crossover, halving)
Inference.mqh 235 softmax, prior calibration, class priors
ExpertSignalAIBase.mqh 8216 -> 3131 (declaration + topology build only)
This is a pure relocation - verified mechanically, not by eye: HEAD's file
reconstructed from the eight partials plus the surviving remainder is
byte-identical to HEAD, span for span (scratchpad verify_split.py). No
declaration moved, no signature changed, no code rewritten, so behaviour is
unchanged by construction.
Compiles 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 00:42:45 -04:00
|
|
|
{
|
refactor(meta): the veto is a gate, not a virtual every signal carries
Since S3 (f64e0f8) the meta head casts no vote - it scores an entry the
consensus already cleared and vetoes the ones under the cost-adjusted
break-even. The code still said otherwise. LiveMetaGate() was a virtual on
CExpertSignalCustom, so MA, RSI, MACD, Ichimoku, the four direction nets,
the session and news filters and the risk guard each carried a meta-gate
method they had no business having; one class implemented it and a dozen
inherited it. The trading pipeline held the gate as a CExpertSignalCustom*
- a signal pointer, with a signal's two hundred other methods reachable
from the entry path.
Expert\Trading\MetaGate.mqh now owns the abstraction:
CMetaGate one pure virtual, Evaluate(), and the two static
readings of a verdict (Blocks / Scored)
META_GATE_* names for the four codes the three call sites used
to spell as bare 0/1/2 and test three different ways
(`< 0` here, `== 2` there, `else` for the rest).
Codes unchanged; only ONE of them blocks, and that
asymmetry is now stated where it lives.
SMetaGateTelemetry the five m_metaGate* members that were on the AI
signal base - inherited by every direction model,
meaningful for none of them. One lifetime, one
writer, one object; the arm latch and the two
counters are a set that clears together.
g_warriorMetaGate is a CMetaGate*. MQL5's single inheritance means the head
cannot also BE one (it already extends the AI base for the net, the era
loop, the feature windows, the label caches and persistence), so it owns a
bound CMetaGateAdapter and hands that out - the same shape CTrainingDataView
uses for the same reason. LiveMetaGate() is gone from the signal base.
Behaviour unchanged: same codes, same thresholds, same fail-open doctrine,
same live-only telemetry rule. The adapter fails open when unbound, on that
same doctrine.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-23 15:33:47 -04:00
|
|
|
//--- Meta target: the live path is the S3 gate (CMetaGate), not the per-bar vote - see
|
feat(ensemble): per-NN inputs replace the preset selector - the meta head becomes the vote's gate
User design (2026-08-19): 'remove the enum menu that selects neural networks... individual
inputs for every NN just like classic signals... the META NN should be integrated into the
voting decision pipeline when enabled... as a bonus meta labelling is applied to enabled NNs.'
- AI_CHOICE is GONE (tombstoned per the stale-.set doctrine). Use_MLP/Use_CONV/Use_LSTM/
Use_CONVLSTM are ordinary bools like the classic votes; the ensemble arithmetic adapts to
any subset because the consensus divisor is the enabled capable weight. Two or more
enabled = ensemble (|ENS1 token + joint gate, exactly the old AI_HYBRID fingerprints, so
existing weight files keep loading); one = the old solo preset; none = classic-only.
- Use_MetaLabeling un-couples META from the direction NNs (the old selector made them
mutually exclusive). S3 ships: CSignalMETA::LiveMetaGate scores each vote-cleared entry
(shared window at bar 1 + proposal descriptor: side, net vote, live geometry, spread/ATR;
pattern one-hot ZEROED - ranking, not calibrated probability, documented in the body) and
vetoes below the cost-adjusted break-even. Entries only; fail-open everywhere, loudly.
- COEXISTENCE HAZARDS closed: VoteCapableWeight()=0 and ProspectiveVote()=false for the
meta target - solo-only until today, a trained META would otherwise sit in the consensus
divisor as a permanent abstainer and shrink every vote by its module weight.
- CERTIFIED == TRADED: the ensemble era verdict replays the identical veto through the same
g_warriorMetaGate pointer over its OOS fired bars (bar re-resolved from the row's own
time; fail-open counted as fires and reported: 'metaGate: N approved, M vetoed, K
unscored'). The overlay deliberately does NOT replay it (veto-filter-in-replay class,
calendar-cliff precedent) - documented at the sweep site. Solo charts' own gate does not
model the veto - the standing solo-gate caveat, documented at the input.
- DB continuity: the pattern/journal DB fingerprint's first slot was (int)AIType;
DbLegacyAiSlot() maps every legacy-expressible config to its OLD value (new 2-3 member
subsets get 100+bitmask, outside the legacy range) so no existing database re-keys.
filterID becomes the enabled roster via one EnabledNNSummary().
- HUD: the meta line shows the gate (armed/(trn), last P vs BE, ok/veto tally); the
armed/disarmed announcement fires on state change via one latch (MetaGateArmedNow), not
only when an entry happens to be proposed.
NOT COMPILED - user compiles in MetaEditor.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-19 13:01:02 -04:00
|
|
|
//--- RefreshConvergedSignal's meta guard.
|
feat: S2 meta-labeling head - binary trade-quality model over the classic-candidate corpus
The NN now has a target that is not per-bar direction (closed, best-of-999
p=1.0000): P(win | this journaled candidate, at the EA's own SL/TP, net of
cost). One net for all 52 pattern-sides, AIType=AI_META.
- NetForward.mqh: the host-side softmax+CE gradient generalized total==3 ->
2||3 on both backprop paths; a 2-class softmax IS a logistic head, and no
compute backend changes.
- SignalMETA.mqh (new): corpus loaded read-only from the LARGEST signal DB on
disk (decoupled from the config fingerprint that burned four S1 runs); the
GMT->server offset is measured PER ROW against entryPrice vs bar open
(DST-immune, histogram logged); a window-span regime filter drops the
pre-2017 daily-backfill rows; 31-feature setup descriptor appended at the
input (26 one-hot + side + tanh netVote + SL/TP ATR + spread/ATR).
- Training.mqh: candidate-queued pass 1, binary-target pass 2, per-candidate
calibration (2.5) and OOS (3) walks. Counter mapping win->Buy / loss->Sell
lets checkpoint selection, the edge floor, the plateau ladder and the
family-wise deploy gate run UNCHANGED: precision reads as win rate among
traded candidates, chance as the base win rate, recalls as sensitivity/
specificity. Era-end META line: coverage x (p - break-even) vs the null.
- Labels are the side-conditional triple-barrier win caches - never the DB's
stop-and-reverse outcome. Logit adjustment deliberately skipped (~40% base
rate). Live inference + online learning guarded off until S3.
- Fingerprint: conditional |TGT:META1; State\META\ folder + 2-output filename
slot keep meta models fully separate from direction models.
Compiles clean (0 errors, 0 warnings). S2 run = attach a chart with
AIType=AI_META; S3 wires the votes via the per-side hooks.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-13 06:52:31 -04:00
|
|
|
if(IsMetaTarget())
|
|
|
|
|
return false;
|
2026-08-22 00:25:52 -04:00
|
|
|
//--- Bar 1: the newest CLOSED bar, NOT the forming bar. Train() never produces such a window -
|
|
|
|
|
//--- every labeled bar is fully closed, and its label assumes entry at that bar's CLOSE (see
|
|
|
|
|
//--- TripleBarrierLabel's header).
|
fix: live inference queried the 1-tick forming bar - a window training never built
RefreshLatestSignal ran at the first tick after a bar opens and built its
window at r=0: series index 0 at that instant is a candle with one tick of
data - (close-open)/atr ~ 0, high ~ low, degenerate volume, indicators on a
1-tick bar. Training never produces such a window (every labeled bar is fully
closed, entry at that bar's CLOSE), so the deployed model's final timestep -
the one the LSTM/HYBRID output is keyed to - was out-of-distribution on every
live decision, and pass 3's deploy-gate OOS scores measured a different query
than live executed. The parity index is r=1: the newest CLOSED bar, whose
close IS the current price - the exact instant the label's hypothetical entry
happens. Single backtests shared the old skew (same r=0), which is why the
tester agreed with live while both disagreed with training.
Bookkeeping split that the index change forces: m_lastBarTime/dtStudied stay
anchored to the FORMING bar's open (they gate against SERIES_LASTBAR_DATE;
anchoring at bar 1 would re-fire the refresh every tick), while bt - the
arrow, its High/Low placement, and NMS declustering - anchors to the decision
bar, now matching the rescan path's convention.
Also: a failed refresh no longer trades the previous bar's signal for the
whole bar. RefreshLatestSignal returns success, zeroes dPrevSignal on failure
(no opinion beats a stale one), and RefreshConvergedSignal advances dtStudied
only on success so the next tick retries - the tester path (m_lastBarTime)
already worked this way; this is the live path catching up.
FORCES RE-VALIDATION of deployed models: the effective live query distribution
changes. Bundled with the backprop transpose fix's retrain.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-11 18:10:23 -04:00
|
|
|
int i = 1;
|
refactor: split CExpertSignalAIBase implementation by responsibility
ExpertSignalAIBase.mqh was 8216 lines: the class declaration followed by 87
method bodies covering training, labelling, feature extraction, persistence,
chart drawing, online learning, the GA auto-tuner and inference, all in one
file. Train() alone is 1492 lines; a change to arrow drawing meant scrolling
past the era loop.
Moved the bodies into Expert\AIBase\, included at the bottom of the original
after the class declaration:
Training.mqh 1607 era loop, plateau ladder, checkpoint select, deploy
Features.mqh 1093 indicator creation + per-bar input feature vector
ChartUI.mqh 634 arrows, arrow persistence, status panel, cleanup
Persistence.mqh 492 .stats/.cfg sidecars, CPU-inference validation, copy
OnlineLearning.mqh 461 live continual learning, EMA shadow, OOS simulator
Labels.mqh 309 ZigZag pivot labels, async label-cache prebuild
AutoTune.mqh 275 genetic tuner (population, crossover, halving)
Inference.mqh 235 softmax, prior calibration, class priors
ExpertSignalAIBase.mqh 8216 -> 3131 (declaration + topology build only)
This is a pure relocation - verified mechanically, not by eye: HEAD's file
reconstructed from the eight partials plus the surviving remainder is
byte-identical to HEAD, span for span (scratchpad verify_split.py). No
declaration moved, no signature changed, no code rewritten, so behaviour is
unchanged by construction.
Compiles 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 00:42:45 -04:00
|
|
|
int r = i;
|
fix: the sequence models were reading the window backwards
BuildFeatureWindow() replaces eight hand-rolled copies of the same loop
and feeds the window OLDEST BAR FIRST. Every copy fed it newest-first,
because MQL5 timeseries indices run backwards and `r + b` with b ascending
walks into the past.
Harmless for PAI and CONV - a dense layer learns a weight per position
either way, a conv learns time-mirrored kernels. Not harmless for the
recurrent stacks:
- LSTM_SeqStepForward reads `inputs + t*Iw`, so step t is block t.
- It writes output[] only when t == steps-1: the visible output IS the
last hidden state.
- c_t = f*c_{t-1} + i*g decays toward the start of the sequence.
lstm_seq_flowcheck.cpp measured block 0's influence on the output at
1.2e-2 of block T-1's, at the shipped forget bias of 1.0.
So the bar being PREDICTED sat at the far end of the decay and the output
was handed to the OLDEST bar in the window - the exact inverse of what the
window is for. ~80x backwards on LSTM and HYBRID, on all three tiers
(OpenCL kernel, CPU DLL, pure-MQL5 inference), which is why it never
surfaced as a backend discrepancy.
This does not create edge - the MI diagnostics read at the noise floor
(p=0.4975) with a working positive control. It makes the one hypothesis
those diagnostics explicitly do NOT cover testable: they are marginal and
per-bar, and state they "cannot rule out one that only exists in
combination or across time". The sequence model is the instrument for
across-time structure and it has been crippled, so that hypothesis has
never been honestly tested.
Fingerprint gets an unconditional |WIN:2 - the vector keeps its shape and
its features, so a stale .nnw would load cleanly and run a model fitted to
one ordering against the other, silently. Re-keying every config is the
point, not collateral damage. FORCES A FULL RETRAIN.
Also: the now-relative bar caches are re-keyed on the two live paths.
EnsureBarCachesCapacity() was only ever called from training paths, but
once m_trainingComplete is set ScheduleTrainingIfNeeded() routes every bar
to RefreshConvergedSignal() and Train() is never re-entered - so nothing
cleared the feature cache again for the life of the process. A chart that
trained to convergence kept replaying the rows computed for the last
training era's bar grid: the live signal froze at its convergence-time
value, and OnlineLearnStep() backpropped those stale features against
freshly resolved labels. Backtests were never affected (an inference-only
process never allocates the arrays, so every read recomputes).
Compiles clean: 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-07 18:28:44 -04:00
|
|
|
if(!BuildFeatureWindow(r))
|
refactor: split CExpertSignalAIBase implementation by responsibility
ExpertSignalAIBase.mqh was 8216 lines: the class declaration followed by 87
method bodies covering training, labelling, feature extraction, persistence,
chart drawing, online learning, the GA auto-tuner and inference, all in one
file. Train() alone is 1492 lines; a change to arrow drawing meant scrolling
past the era loop.
Moved the bodies into Expert\AIBase\, included at the bottom of the original
after the class declaration:
Training.mqh 1607 era loop, plateau ladder, checkpoint select, deploy
Features.mqh 1093 indicator creation + per-bar input feature vector
ChartUI.mqh 634 arrows, arrow persistence, status panel, cleanup
Persistence.mqh 492 .stats/.cfg sidecars, CPU-inference validation, copy
OnlineLearning.mqh 461 live continual learning, EMA shadow, OOS simulator
Labels.mqh 309 ZigZag pivot labels, async label-cache prebuild
AutoTune.mqh 275 genetic tuner (population, crossover, halving)
Inference.mqh 235 softmax, prior calibration, class priors
ExpertSignalAIBase.mqh 8216 -> 3131 (declaration + topology build only)
This is a pure relocation - verified mechanically, not by eye: HEAD's file
reconstructed from the eight partials plus the surviving remainder is
byte-identical to HEAD, span for span (scratchpad verify_split.py). No
declaration moved, no signature changed, no code rewritten, so behaviour is
unchanged by construction.
Compiles 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 00:42:45 -04:00
|
|
|
{
|
fix: the sequence models were reading the window backwards
BuildFeatureWindow() replaces eight hand-rolled copies of the same loop
and feeds the window OLDEST BAR FIRST. Every copy fed it newest-first,
because MQL5 timeseries indices run backwards and `r + b` with b ascending
walks into the past.
Harmless for PAI and CONV - a dense layer learns a weight per position
either way, a conv learns time-mirrored kernels. Not harmless for the
recurrent stacks:
- LSTM_SeqStepForward reads `inputs + t*Iw`, so step t is block t.
- It writes output[] only when t == steps-1: the visible output IS the
last hidden state.
- c_t = f*c_{t-1} + i*g decays toward the start of the sequence.
lstm_seq_flowcheck.cpp measured block 0's influence on the output at
1.2e-2 of block T-1's, at the shipped forget bias of 1.0.
So the bar being PREDICTED sat at the far end of the decay and the output
was handed to the OLDEST bar in the window - the exact inverse of what the
window is for. ~80x backwards on LSTM and HYBRID, on all three tiers
(OpenCL kernel, CPU DLL, pure-MQL5 inference), which is why it never
surfaced as a backend discrepancy.
This does not create edge - the MI diagnostics read at the noise floor
(p=0.4975) with a working positive control. It makes the one hypothesis
those diagnostics explicitly do NOT cover testable: they are marginal and
per-bar, and state they "cannot rule out one that only exists in
combination or across time". The sequence model is the instrument for
across-time structure and it has been crippled, so that hypothesis has
never been honestly tested.
Fingerprint gets an unconditional |WIN:2 - the vector keeps its shape and
its features, so a stale .nnw would load cleanly and run a model fitted to
one ordering against the other, silently. Re-keying every config is the
point, not collateral damage. FORCES A FULL RETRAIN.
Also: the now-relative bar caches are re-keyed on the two live paths.
EnsureBarCachesCapacity() was only ever called from training paths, but
once m_trainingComplete is set ScheduleTrainingIfNeeded() routes every bar
to RefreshConvergedSignal() and Train() is never re-entered - so nothing
cleared the feature cache again for the life of the process. A chart that
trained to convergence kept replaying the rows computed for the last
training era's bar grid: the live signal froze at its convergence-time
value, and OnlineLearnStep() backpropped those stale features against
freshly resolved labels. Backtests were never affected (an inference-only
process never allocates the arrays, so every read recomputes).
Compiles clean: 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-07 18:28:44 -04:00
|
|
|
//--- One combined failure now (partial window OR short total) where there used to be two counters.
|
|
|
|
|
//--- Kept distinct in the tally by testing what actually landed: a window that built every bar but
|
|
|
|
|
//--- came up short is the "short" case, anything else is a feature-build failure.
|
|
|
|
|
if(TempData.Total() > 0 && TempData.Total() < (int)m_historyBars * m_neuronsCount)
|
|
|
|
|
m_refreshFailShort++;
|
|
|
|
|
else
|
diag: inference-path census, to explain zero-trade backtests
A backtest of the CONVERGED CONV model produced "Final directional result:
0.00000000" on every one of 1744 bars and therefore zero trades. Nothing in
the log could separate the three candidate causes, and each needs a
different fix:
1. RefreshLatestSignal never called (new-bar gate never fires)
2. called, but bailing at one of its two early returns
3. running fine, and the model genuinely answers Neutral every bar
Counts all three plus the Buy/Sell/Neutral split, printed once at shutdown
via StopTraining (which the tester reaches through OnDeinit). Three
increments per bar against a full feedForward - not worth gating.
Ruled out while writing this, so the next session does not re-derive it:
- the alternation gate (m_lastNonNeutralSignal) is NOT the cause. It starts
at Neutral, so a first Buy would still fire and show up as one non-zero
direction. We saw zero. It IS still a live hazard for a one-sided model -
CONV currently calls Buy:17% Sell:0%, and after the first Buy every later
Buy is suppressed until a Sell that never comes - but it cannot explain
an all-zero run.
- shallow buffers do not hard-fail the feature builder: the swing-context
Donchian loop breaks gracefully when it runs off loaded history. It does
mean converged-path inference computes Donchian/return/SMA features over
a TRUNCATED window versus training, which is a real train/inference skew
worth its own fix, but it degrades features rather than zeroing them.
Both builds 0/0. Diagnostic only.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 18:24:32 -04:00
|
|
|
m_refreshFailFeatures++; // see PrintInferenceTally()
|
fix: live inference queried the 1-tick forming bar - a window training never built
RefreshLatestSignal ran at the first tick after a bar opens and built its
window at r=0: series index 0 at that instant is a candle with one tick of
data - (close-open)/atr ~ 0, high ~ low, degenerate volume, indicators on a
1-tick bar. Training never produces such a window (every labeled bar is fully
closed, entry at that bar's CLOSE), so the deployed model's final timestep -
the one the LSTM/HYBRID output is keyed to - was out-of-distribution on every
live decision, and pass 3's deploy-gate OOS scores measured a different query
than live executed. The parity index is r=1: the newest CLOSED bar, whose
close IS the current price - the exact instant the label's hypothetical entry
happens. Single backtests shared the old skew (same r=0), which is why the
tester agreed with live while both disagreed with training.
Bookkeeping split that the index change forces: m_lastBarTime/dtStudied stay
anchored to the FORMING bar's open (they gate against SERIES_LASTBAR_DATE;
anchoring at bar 1 would re-fire the refresh every tick), while bt - the
arrow, its High/Low placement, and NMS declustering - anchors to the decision
bar, now matching the rescan path's convention.
Also: a failed refresh no longer trades the previous bar's signal for the
whole bar. RefreshLatestSignal returns success, zeroes dPrevSignal on failure
(no opinion beats a stale one), and RefreshConvergedSignal advances dtStudied
only on success so the next tick retries - the tester path (m_lastBarTime)
already worked this way; this is the live path catching up.
FORCES RE-VALIDATION of deployed models: the effective live query distribution
changes. Bundled with the backprop transpose fix's retrain.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-11 18:10:23 -04:00
|
|
|
//--- No opinion this bar rather than a stale one: dPrevSignal still holds the PREVIOUS bar's
|
2026-08-22 00:25:52 -04:00
|
|
|
//--- decision, and LongCondition()/ShortCondition() would keep voting that stale direction
|
|
|
|
|
//--- all bar.
|
fix: live inference queried the 1-tick forming bar - a window training never built
RefreshLatestSignal ran at the first tick after a bar opens and built its
window at r=0: series index 0 at that instant is a candle with one tick of
data - (close-open)/atr ~ 0, high ~ low, degenerate volume, indicators on a
1-tick bar. Training never produces such a window (every labeled bar is fully
closed, entry at that bar's CLOSE), so the deployed model's final timestep -
the one the LSTM/HYBRID output is keyed to - was out-of-distribution on every
live decision, and pass 3's deploy-gate OOS scores measured a different query
than live executed. The parity index is r=1: the newest CLOSED bar, whose
close IS the current price - the exact instant the label's hypothetical entry
happens. Single backtests shared the old skew (same r=0), which is why the
tester agreed with live while both disagreed with training.
Bookkeeping split that the index change forces: m_lastBarTime/dtStudied stay
anchored to the FORMING bar's open (they gate against SERIES_LASTBAR_DATE;
anchoring at bar 1 would re-fire the refresh every tick), while bt - the
arrow, its High/Low placement, and NMS declustering - anchors to the decision
bar, now matching the rescan path's convention.
Also: a failed refresh no longer trades the previous bar's signal for the
whole bar. RefreshLatestSignal returns success, zeroes dPrevSignal on failure
(no opinion beats a stale one), and RefreshConvergedSignal advances dtStudied
only on success so the next tick retries - the tester path (m_lastBarTime)
already worked this way; this is the live path catching up.
FORCES RE-VALIDATION of deployed models: the effective live query distribution
changes. Bundled with the backprop transpose fix's retrain.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-11 18:10:23 -04:00
|
|
|
dPrevSignal = 0.0;
|
|
|
|
|
return false;
|
diag: inference-path census, to explain zero-trade backtests
A backtest of the CONVERGED CONV model produced "Final directional result:
0.00000000" on every one of 1744 bars and therefore zero trades. Nothing in
the log could separate the three candidate causes, and each needs a
different fix:
1. RefreshLatestSignal never called (new-bar gate never fires)
2. called, but bailing at one of its two early returns
3. running fine, and the model genuinely answers Neutral every bar
Counts all three plus the Buy/Sell/Neutral split, printed once at shutdown
via StopTraining (which the tester reaches through OnDeinit). Three
increments per bar against a full feedForward - not worth gating.
Ruled out while writing this, so the next session does not re-derive it:
- the alternation gate (m_lastNonNeutralSignal) is NOT the cause. It starts
at Neutral, so a first Buy would still fire and show up as one non-zero
direction. We saw zero. It IS still a live hazard for a one-sided model -
CONV currently calls Buy:17% Sell:0%, and after the first Buy every later
Buy is suppressed until a Sell that never comes - but it cannot explain
an all-zero run.
- shallow buffers do not hard-fail the feature builder: the swing-context
Donchian loop breaks gracefully when it runs off loaded history. It does
mean converged-path inference computes Donchian/return/SMA features over
a TRUNCATED window versus training, which is a real train/inference skew
worth its own fix, but it degrades features rather than zeroing them.
Both builds 0/0. Diagnostic only.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 18:24:32 -04:00
|
|
|
}
|
refactor: split CExpertSignalAIBase implementation by responsibility
ExpertSignalAIBase.mqh was 8216 lines: the class declaration followed by 87
method bodies covering training, labelling, feature extraction, persistence,
chart drawing, online learning, the GA auto-tuner and inference, all in one
file. Train() alone is 1492 lines; a change to arrow drawing meant scrolling
past the era loop.
Moved the bodies into Expert\AIBase\, included at the bottom of the original
after the class declaration:
Training.mqh 1607 era loop, plateau ladder, checkpoint select, deploy
Features.mqh 1093 indicator creation + per-bar input feature vector
ChartUI.mqh 634 arrows, arrow persistence, status panel, cleanup
Persistence.mqh 492 .stats/.cfg sidecars, CPU-inference validation, copy
OnlineLearning.mqh 461 live continual learning, EMA shadow, OOS simulator
Labels.mqh 309 ZigZag pivot labels, async label-cache prebuild
AutoTune.mqh 275 genetic tuner (population, crossover, halving)
Inference.mqh 235 softmax, prior calibration, class priors
ExpertSignalAIBase.mqh 8216 -> 3131 (declaration + topology build only)
This is a pure relocation - verified mechanically, not by eye: HEAD's file
reconstructed from the eight partials plus the surviving remainder is
byte-identical to HEAD, span for span (scratchpad verify_split.py). No
declaration moved, no signature changed, no code rewritten, so behaviour is
unchanged by construction.
Compiles 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 00:42:45 -04:00
|
|
|
//--- Live trading/inference reads from the EMA shadow net, not Net directly - see m_shadowNet's
|
|
|
|
|
//--- declaration comment. Falls back to Net if the shadow isn't bootstrapped yet (should only be
|
|
|
|
|
//--- momentarily, on a genuinely fresh start before EnsureShadowNet() has run).
|
|
|
|
|
EnsureShadowNet();
|
|
|
|
|
CNet *deployNet = (CheckPointer(m_shadowNet) != POINTER_INVALID) ? m_shadowNet : Net;
|
|
|
|
|
deployNet.feedForward(TempData);
|
|
|
|
|
deployNet.getResults(TempData);
|
|
|
|
|
if(m_outputNeuronsCount == 1)
|
|
|
|
|
dPrevSignal = TempData[0];
|
|
|
|
|
else
|
|
|
|
|
if(m_outputNeuronsCount == 3)
|
|
|
|
|
{
|
2026-08-22 00:25:52 -04:00
|
|
|
//--- Live decision. The call is kept because a dozen sites name it as "the live decision
|
|
|
|
|
//--- rule" and that is still exactly what it is - the correction now lives in the trained
|
|
|
|
|
//--- weights instead of here.
|
refactor(ai): nine class-imbalance inputs down to two
The imbalance section offered nine controls for one job. Audited against the
code, five of them did not do what their names said at the shipped defaults:
AILogitPriorStrength DEAD - Inference.mqh's post-hoc prior early-returns
whenever the adjusted loss is on, which is default.
OversampleParity DEAD in training - Training.mqh gated the replay loop
on !useLogitAdjustedLoss (correctly, citing Buda et
al. 2018). Live only in the online-learning path.
EnableMinorityReplay DEAD as replay. It survived ONLY as a focal-gamma
damper - "replay minority bars through pass-2
oversampling" was a focal-loss switch.
ConstrainReplay DEAD as a cap; it only chose damper 0.125 vs 0.25.
UseStaticPrior An exact duplicate of FreezePriorCalibration - the two
were OR'd together in the single place either is read.
So they were not five mechanisms fighting; they were one mechanism plus eight
knobs that mostly described machinery that no longer ran. That is worse than
a real conflict, because the log agreed with the names: the label-cache line
printed "reps up to 28x (90% parity) (seeding era 0's class-balance
oversampling)" on every run, describing an oversampling pass that had been
switched off. It is fixed here too - it cost this session a wrong diagnosis.
The one genuine redundancy was focal loss, running at gamma*0.125 alongside
the adjusted loss: two corrections on the same axis, the exact stacking
failure this file already cited Buda et al. for in two other places, damped
by a replay flag whose replay path was itself dead. Removed rather than
re-tuned. The plateau ladder is unaffected - its escape is the learning-rate
warm restart; the gamma anneal beside it only ever stepped toward zero.
WHAT REMAINS is logit-adjusted loss (Menon et al. 2021) plus a prior freeze:
LogitAdjustTau 0 = off; replaces the separate EnableLogitAdjusted-
Loss boolean, since a strength dial where 0 already
means off does not need an on/off switch beside it.
FreezePriorCalibration unchanged.
It is the only one of the six corrections with a consistency guarantee, and
it is consistent for exactly the balanced-error metric checkpoint selection
already ranks on - so the loss and the deploy decision optimize one thing.
The online continual-learning path keeps its own alpha-balanced focal weight,
now as constants pinned to the removed inputs' shipped defaults, so its
behaviour is unchanged. It legitimately needs its own correction:
ApplyLogitAdjustment() only runs inside a training run, so a deployed model
that was reloaded carries no logit offsets and would otherwise stream 31:1
data into itself uncorrected.
The weights-filename fingerprint is BYTE-IDENTICAL. The focal slot was a
double fed to a %d conversion and had always emitted a literal 0; the |MR:
segment is written as the constant its shipped defaults produced. Dropping
either would have re-keyed every model and forced a from-scratch retrain of
the one topology currently converged and trading.
Also removed as orphans: FOCAL_GAMMA_PRESET, MAX_OVERSAMPLE_REPLICAS,
OVERSAMPLE_PARITY_FRACTION, PLATEAU_GAMMA_STEP, and the now-unreachable
"neutralized by prior correction" diagnostic.
Both builds compile 0 errors, 0 warnings. No retrain forced.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 11:46:57 -04:00
|
|
|
ApplyClassificationSoftmax();
|
refactor: split CExpertSignalAIBase implementation by responsibility
ExpertSignalAIBase.mqh was 8216 lines: the class declaration followed by 87
method bodies covering training, labelling, feature extraction, persistence,
chart drawing, online learning, the GA auto-tuner and inference, all in one
file. Train() alone is 1492 lines; a change to arrow drawing meant scrolling
past the era loop.
Moved the bodies into Expert\AIBase\, included at the bottom of the original
after the class declaration:
Training.mqh 1607 era loop, plateau ladder, checkpoint select, deploy
Features.mqh 1093 indicator creation + per-bar input feature vector
ChartUI.mqh 634 arrows, arrow persistence, status panel, cleanup
Persistence.mqh 492 .stats/.cfg sidecars, CPU-inference validation, copy
OnlineLearning.mqh 461 live continual learning, EMA shadow, OOS simulator
Labels.mqh 309 ZigZag pivot labels, async label-cache prebuild
AutoTune.mqh 275 genetic tuner (population, crossover, halving)
Inference.mqh 235 softmax, prior calibration, class priors
ExpertSignalAIBase.mqh 8216 -> 3131 (declaration + topology build only)
This is a pure relocation - verified mechanically, not by eye: HEAD's file
reconstructed from the eight partials plus the surviving remainder is
byte-identical to HEAD, span for span (scratchpad verify_split.py). No
declaration moved, no signature changed, no code rewritten, so behaviour is
unchanged by construction.
Compiles 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 00:42:45 -04:00
|
|
|
dPrevSignal = AdjustedSignalFromSoftmax();
|
|
|
|
|
}
|
diag: inference-path census, to explain zero-trade backtests
A backtest of the CONVERGED CONV model produced "Final directional result:
0.00000000" on every one of 1744 bars and therefore zero trades. Nothing in
the log could separate the three candidate causes, and each needs a
different fix:
1. RefreshLatestSignal never called (new-bar gate never fires)
2. called, but bailing at one of its two early returns
3. running fine, and the model genuinely answers Neutral every bar
Counts all three plus the Buy/Sell/Neutral split, printed once at shutdown
via StopTraining (which the tester reaches through OnDeinit). Three
increments per bar against a full feedForward - not worth gating.
Ruled out while writing this, so the next session does not re-derive it:
- the alternation gate (m_lastNonNeutralSignal) is NOT the cause. It starts
at Neutral, so a first Buy would still fire and show up as one non-zero
direction. We saw zero. It IS still a live hazard for a one-sided model -
CONV currently calls Buy:17% Sell:0%, and after the first Buy every later
Buy is suppressed until a Sell that never comes - but it cannot explain
an all-zero run.
- shallow buffers do not hard-fail the feature builder: the swing-context
Donchian loop breaks gracefully when it runs off loaded history. It does
mean converged-path inference computes Donchian/return/SMA features over
a TRUNCATED window versus training, which is a real train/inference skew
worth its own fix, but it degrades features rather than zeroing them.
Both builds 0/0. Diagnostic only.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 18:24:32 -04:00
|
|
|
m_refreshOk++;
|
fix: live inference queried the 1-tick forming bar - a window training never built
RefreshLatestSignal ran at the first tick after a bar opens and built its
window at r=0: series index 0 at that instant is a candle with one tick of
data - (close-open)/atr ~ 0, high ~ low, degenerate volume, indicators on a
1-tick bar. Training never produces such a window (every labeled bar is fully
closed, entry at that bar's CLOSE), so the deployed model's final timestep -
the one the LSTM/HYBRID output is keyed to - was out-of-distribution on every
live decision, and pass 3's deploy-gate OOS scores measured a different query
than live executed. The parity index is r=1: the newest CLOSED bar, whose
close IS the current price - the exact instant the label's hypothetical entry
happens. Single backtests shared the old skew (same r=0), which is why the
tester agreed with live while both disagreed with training.
Bookkeeping split that the index change forces: m_lastBarTime/dtStudied stay
anchored to the FORMING bar's open (they gate against SERIES_LASTBAR_DATE;
anchoring at bar 1 would re-fire the refresh every tick), while bt - the
arrow, its High/Low placement, and NMS declustering - anchors to the decision
bar, now matching the rescan path's convention.
Also: a failed refresh no longer trades the previous bar's signal for the
whole bar. RefreshLatestSignal returns success, zeroes dPrevSignal on failure
(no opinion beats a stale one), and RefreshConvergedSignal advances dtStudied
only on success so the next tick retries - the tester path (m_lastBarTime)
already worked this way; this is the live path catching up.
FORCES RE-VALIDATION of deployed models: the effective live query distribution
changes. Bundled with the backprop transpose fix's retrain.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-11 18:10:23 -04:00
|
|
|
//--- bt anchors the DECISION bar (bar 1, the closed bar the window ends on) - it keys the arrow,
|
|
|
|
|
//--- its High/Low placement and NMS declustering, and now matches the rescan path, which draws
|
|
|
|
|
//--- each historical arrow at the bar its window ends on.
|
fix: NMS gates the TRADE, not just the arrow - one arrow is now one trade
NmsLiveAccept() appeared in exactly one place: wrapped around DrawObject().
It never touched dPrevSignal, and dPrevSignal is what LongCondition() /
ShortCondition() / SignedAIConfidence() read. So a declustered bar lost its
arrow and still opened a position.
Measured on SP500 H1 2026-08-09: CONV called a direction on 64% of bars,
so the ~500 bars visible on screen held ~320 decisions - and ~40 arrows
were drawn. Roughly one arrow per eight positions the EA would take.
And the survivors are not a random eighth. Rule 2 of the declustering
keeps the HIGHER-CONFIDENCE side of a cluster, so the visible set is
systematically the best member of each run. A chart showing the best of
every eight decisions and hiding the rest reads far better than the model
is - the same best-of-N selection error already corrected in the geometry
scan, the indicator tuner, the lag profile and the deploy gate, this time
on the display layer, where it is most likely to mislead the person
deciding whether to trade.
Fixed by neutralising dPrevSignal when NMS rejects, rather than adding a
"may trade" flag consulted at each read site: that leaves exactly ONE
definition of what the model decided this bar, so the arrow, the panel's
"Current signal", the confidence feeding sizing/SL/TP/trailing, the
refresh tally and the order itself cannot drift apart again.
Also reports the consequence instead of hiding it. Every OOS counter on
the era line still scores every directional call - a population ~8x larger
than what now trades - so the line carries a second figure:
| TRADED (declustered) NN% on N calls (edge +Npp)
replaying the identical rule over pass 3 (which walks OOS bars oldest to
newest, the same order the live sweep sees). Its cursors are separate
members from the live ones so a training pass can never disturb the live
chart's declustering.
Deliberately NOT switched into selectionScore yet. Declustering cuts
coverage from ~64% of bars to ~8%, well under
MIN_COVERAGE_FRACTION_OF_BASE_RATE, which would make every checkpoint
undeployable overnight - the minRR collision and the recall-floor catch-22
twice over. The floor gets re-derived from these measurements first.
Compiles clean: 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-09 10:22:31 -04:00
|
|
|
datetime bt = m_Time.GetData(i);
|
2026-08-22 00:25:52 -04:00
|
|
|
//--- Keep a pure inference-side watermark of the newest bar FRAME this model has already
|
|
|
|
|
//--- evaluated.
|
fix: live inference queried the 1-tick forming bar - a window training never built
RefreshLatestSignal ran at the first tick after a bar opens and built its
window at r=0: series index 0 at that instant is a candle with one tick of
data - (close-open)/atr ~ 0, high ~ low, degenerate volume, indicators on a
1-tick bar. Training never produces such a window (every labeled bar is fully
closed, entry at that bar's CLOSE), so the deployed model's final timestep -
the one the LSTM/HYBRID output is keyed to - was out-of-distribution on every
live decision, and pass 3's deploy-gate OOS scores measured a different query
than live executed. The parity index is r=1: the newest CLOSED bar, whose
close IS the current price - the exact instant the label's hypothetical entry
happens. Single backtests shared the old skew (same r=0), which is why the
tester agreed with live while both disagreed with training.
Bookkeeping split that the index change forces: m_lastBarTime/dtStudied stay
anchored to the FORMING bar's open (they gate against SERIES_LASTBAR_DATE;
anchoring at bar 1 would re-fire the refresh every tick), while bt - the
arrow, its High/Low placement, and NMS declustering - anchors to the decision
bar, now matching the rescan path's convention.
Also: a failed refresh no longer trades the previous bar's signal for the
whole bar. RefreshLatestSignal returns success, zeroes dPrevSignal on failure
(no opinion beats a stale one), and RefreshConvergedSignal advances dtStudied
only on success so the next tick retries - the tester path (m_lastBarTime)
already worked this way; this is the live path catching up.
FORCES RE-VALIDATION of deployed models: the effective live query distribution
changes. Bundled with the backprop transpose fix's retrain.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-11 18:10:23 -04:00
|
|
|
m_lastBarTime = m_Time.GetData(0);
|
fix: NMS gates the TRADE, not just the arrow - one arrow is now one trade
NmsLiveAccept() appeared in exactly one place: wrapped around DrawObject().
It never touched dPrevSignal, and dPrevSignal is what LongCondition() /
ShortCondition() / SignedAIConfidence() read. So a declustered bar lost its
arrow and still opened a position.
Measured on SP500 H1 2026-08-09: CONV called a direction on 64% of bars,
so the ~500 bars visible on screen held ~320 decisions - and ~40 arrows
were drawn. Roughly one arrow per eight positions the EA would take.
And the survivors are not a random eighth. Rule 2 of the declustering
keeps the HIGHER-CONFIDENCE side of a cluster, so the visible set is
systematically the best member of each run. A chart showing the best of
every eight decisions and hiding the rest reads far better than the model
is - the same best-of-N selection error already corrected in the geometry
scan, the indicator tuner, the lag profile and the deploy gate, this time
on the display layer, where it is most likely to mislead the person
deciding whether to trade.
Fixed by neutralising dPrevSignal when NMS rejects, rather than adding a
"may trade" flag consulted at each read site: that leaves exactly ONE
definition of what the model decided this bar, so the arrow, the panel's
"Current signal", the confidence feeding sizing/SL/TP/trailing, the
refresh tally and the order itself cannot drift apart again.
Also reports the consequence instead of hiding it. Every OOS counter on
the era line still scores every directional call - a population ~8x larger
than what now trades - so the line carries a second figure:
| TRADED (declustered) NN% on N calls (edge +Npp)
replaying the identical rule over pass 3 (which walks OOS bars oldest to
newest, the same order the live sweep sees). Its cursors are separate
members from the live ones so a training pass can never disturb the live
chart's declustering.
Deliberately NOT switched into selectionScore yet. Declustering cuts
coverage from ~64% of bars to ~8%, well under
MIN_COVERAGE_FRACTION_OF_BASE_RATE, which would make every checkpoint
undeployable overnight - the minRR collision and the recall-floor catch-22
twice over. The floor gets re-derived from these measurements first.
Compiles clean: 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-09 10:22:31 -04:00
|
|
|
//--- LIVE NMS, AND IT NOW GATES THE TRADE, NOT JUST THE ARROW.
|
|
|
|
|
ENUM_SIGNAL lsig = DoubleToSignal(dPrevSignal);
|
|
|
|
|
bool nmsAccept = (lsig != Neutral) && NmsLiveAccept(bt, lsig, MathAbs(dPrevSignal));
|
|
|
|
|
if(lsig != Neutral && !nmsAccept)
|
|
|
|
|
dPrevSignal = 0.0; // declustered away: no arrow, no vote, no position
|
diag: inference-path census, to explain zero-trade backtests
A backtest of the CONVERGED CONV model produced "Final directional result:
0.00000000" on every one of 1744 bars and therefore zero trades. Nothing in
the log could separate the three candidate causes, and each needs a
different fix:
1. RefreshLatestSignal never called (new-bar gate never fires)
2. called, but bailing at one of its two early returns
3. running fine, and the model genuinely answers Neutral every bar
Counts all three plus the Buy/Sell/Neutral split, printed once at shutdown
via StopTraining (which the tester reaches through OnDeinit). Three
increments per bar against a full feedForward - not worth gating.
Ruled out while writing this, so the next session does not re-derive it:
- the alternation gate (m_lastNonNeutralSignal) is NOT the cause. It starts
at Neutral, so a first Buy would still fire and show up as one non-zero
direction. We saw zero. It IS still a live hazard for a one-sided model -
CONV currently calls Buy:17% Sell:0%, and after the first Buy every later
Buy is suppressed until a Sell that never comes - but it cannot explain
an all-zero run.
- shallow buffers do not hard-fail the feature builder: the swing-context
Donchian loop breaks gracefully when it runs off loaded history. It does
mean converged-path inference computes Donchian/return/SMA features over
a TRUNCATED window versus training, which is a real train/inference skew
worth its own fix, but it degrades features rather than zeroing them.
Both builds 0/0. Diagnostic only.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 18:24:32 -04:00
|
|
|
switch(DoubleToSignal(dPrevSignal))
|
|
|
|
|
{
|
|
|
|
|
case Buy:
|
|
|
|
|
m_refreshBuy++;
|
|
|
|
|
break;
|
|
|
|
|
case Sell:
|
|
|
|
|
m_refreshSell++;
|
|
|
|
|
break;
|
|
|
|
|
default:
|
|
|
|
|
m_refreshNeutral++;
|
|
|
|
|
break;
|
|
|
|
|
}
|
fix: NMS gates the TRADE, not just the arrow - one arrow is now one trade
NmsLiveAccept() appeared in exactly one place: wrapped around DrawObject().
It never touched dPrevSignal, and dPrevSignal is what LongCondition() /
ShortCondition() / SignedAIConfidence() read. So a declustered bar lost its
arrow and still opened a position.
Measured on SP500 H1 2026-08-09: CONV called a direction on 64% of bars,
so the ~500 bars visible on screen held ~320 decisions - and ~40 arrows
were drawn. Roughly one arrow per eight positions the EA would take.
And the survivors are not a random eighth. Rule 2 of the declustering
keeps the HIGHER-CONFIDENCE side of a cluster, so the visible set is
systematically the best member of each run. A chart showing the best of
every eight decisions and hiding the rest reads far better than the model
is - the same best-of-N selection error already corrected in the geometry
scan, the indicator tuner, the lag profile and the deploy gate, this time
on the display layer, where it is most likely to mislead the person
deciding whether to trade.
Fixed by neutralising dPrevSignal when NMS rejects, rather than adding a
"may trade" flag consulted at each read site: that leaves exactly ONE
definition of what the model decided this bar, so the arrow, the panel's
"Current signal", the confidence feeding sizing/SL/TP/trailing, the
refresh tally and the order itself cannot drift apart again.
Also reports the consequence instead of hiding it. Every OOS counter on
the era line still scores every directional call - a population ~8x larger
than what now trades - so the line carries a second figure:
| TRADED (declustered) NN% on N calls (edge +Npp)
replaying the identical rule over pass 3 (which walks OOS bars oldest to
newest, the same order the live sweep sees). Its cursors are separate
members from the live ones so a training pass can never disturb the live
chart's declustering.
Deliberately NOT switched into selectionScore yet. Declustering cuts
coverage from ~64% of bars to ~8%, well under
MIN_COVERAGE_FRACTION_OF_BASE_RATE, which would make every checkpoint
undeployable overnight - the minRR collision and the recall-floor catch-22
twice over. The floor gets re-derived from these measurements first.
Compiles clean: 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-09 10:22:31 -04:00
|
|
|
if(nmsAccept)
|
feat(chart): signal marks become price LEVELS at the trigger, not arrows beside the candle
User request: 'move from arrows on lows and highs to small horizontal lines at the actual
prices the entry/exit would trigger, just a bit larger than the candles. dark green for
buy, dark red for sell.'
Every mark is now an OBJ_TREND segment with both anchors at one price and both rays off,
spanning 1.3 bar widths, drawn at the bar's CLOSE - the price a market order actually
fires at, and the exact entry TripleBarrierLabel assumes. It used to sit on the candle's
LOW for a Buy and its HIGH for a Sell: prices the trade never touches, picked so an arrow
glyph would clear the candle. The tooltip now carries that price too.
COLOUR NOW MEANS DIRECTION AND ONLY DIRECTION on every layer (dark green / dark red).
Layer moves to width+style - the traded vote is solid and thick and drawn in front, a
single model's raw opinion is thin, dotted and behind the candles - which keeps the
distinction the old palette existed to draw (a model's opinion must never read as a trade)
while freeing colour to say one thing consistently.
Consequences handled, all of them the same 'a typed scan went blind' failure:
- SaveChartSignals filtered OBJPROP_TYPE == OBJ_ARROW and read OBJPROP_ARROWCODE. It now
filters OBJ_TREND and recovers direction from the colour. The sidecar keeps the old
217/218 numbers as its buy/sell token deliberately, so existing .arrows files still load.
- AdvanceChartSignalRestore now rebuilds through the SAME creation point the live path
uses, so a restored mark and a fresh one are identical objects.
- The rescan-scoped delete enumerated ObjectsTotal(OBJ_ARROW) - retyped, or it silently
deletes nothing.
- ApplySignalsVisibility enumerated OBJ_ARROW with NO prefix filter. Under the new type
that would have hidden and shown THE USER'S OWN trend lines on every Hide/Show click;
it is now prefix-scoped. The old type was uncommon enough on a real chart to mask the
missing check - trend lines are the most hand-drawn object there is.
- DrawObject's high/low parameters are gone (6 call sites pass m_Close instead), so no
caller can hand it a price it no longer draws at.
- Fixed a pre-existing stale comment that still described the purge sweep as OBJ_ARROW-only
three lines above the note explaining it had been widened to every type.
NOT COMPILED - user compiles in MetaEditor.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-19 14:11:49 -04:00
|
|
|
DrawObject(bt, dPrevSignal, m_Close.GetData(i));
|
refactor: split CExpertSignalAIBase implementation by responsibility
ExpertSignalAIBase.mqh was 8216 lines: the class declaration followed by 87
method bodies covering training, labelling, feature extraction, persistence,
chart drawing, online learning, the GA auto-tuner and inference, all in one
file. Train() alone is 1492 lines; a change to arrow drawing meant scrolling
past the era loop.
Moved the bodies into Expert\AIBase\, included at the bottom of the original
after the class declaration:
Training.mqh 1607 era loop, plateau ladder, checkpoint select, deploy
Features.mqh 1093 indicator creation + per-bar input feature vector
ChartUI.mqh 634 arrows, arrow persistence, status panel, cleanup
Persistence.mqh 492 .stats/.cfg sidecars, CPU-inference validation, copy
OnlineLearning.mqh 461 live continual learning, EMA shadow, OOS simulator
Labels.mqh 309 ZigZag pivot labels, async label-cache prebuild
AutoTune.mqh 275 genetic tuner (population, crossover, halving)
Inference.mqh 235 softmax, prior calibration, class priors
ExpertSignalAIBase.mqh 8216 -> 3131 (declaration + topology build only)
This is a pure relocation - verified mechanically, not by eye: HEAD's file
reconstructed from the eight partials plus the surviving remainder is
byte-identical to HEAD, span for span (scratchpad verify_split.py). No
declaration moved, no signature changed, no code rewritten, so behaviour is
unchanged by construction.
Compiles 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 00:42:45 -04:00
|
|
|
else
|
|
|
|
|
DeleteObject(bt);
|
fix: live inference queried the 1-tick forming bar - a window training never built
RefreshLatestSignal ran at the first tick after a bar opens and built its
window at r=0: series index 0 at that instant is a candle with one tick of
data - (close-open)/atr ~ 0, high ~ low, degenerate volume, indicators on a
1-tick bar. Training never produces such a window (every labeled bar is fully
closed, entry at that bar's CLOSE), so the deployed model's final timestep -
the one the LSTM/HYBRID output is keyed to - was out-of-distribution on every
live decision, and pass 3's deploy-gate OOS scores measured a different query
than live executed. The parity index is r=1: the newest CLOSED bar, whose
close IS the current price - the exact instant the label's hypothetical entry
happens. Single backtests shared the old skew (same r=0), which is why the
tester agreed with live while both disagreed with training.
Bookkeeping split that the index change forces: m_lastBarTime/dtStudied stay
anchored to the FORMING bar's open (they gate against SERIES_LASTBAR_DATE;
anchoring at bar 1 would re-fire the refresh every tick), while bt - the
arrow, its High/Low placement, and NMS declustering - anchors to the decision
bar, now matching the rescan path's convention.
Also: a failed refresh no longer trades the previous bar's signal for the
whole bar. RefreshLatestSignal returns success, zeroes dPrevSignal on failure
(no opinion beats a stale one), and RefreshConvergedSignal advances dtStudied
only on success so the next tick retries - the tester path (m_lastBarTime)
already worked this way; this is the live path catching up.
FORCES RE-VALIDATION of deployed models: the effective live query distribution
changes. Bundled with the backprop transpose fix's retrain.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-11 18:10:23 -04:00
|
|
|
return true;
|
refactor: split CExpertSignalAIBase implementation by responsibility
ExpertSignalAIBase.mqh was 8216 lines: the class declaration followed by 87
method bodies covering training, labelling, feature extraction, persistence,
chart drawing, online learning, the GA auto-tuner and inference, all in one
file. Train() alone is 1492 lines; a change to arrow drawing meant scrolling
past the era loop.
Moved the bodies into Expert\AIBase\, included at the bottom of the original
after the class declaration:
Training.mqh 1607 era loop, plateau ladder, checkpoint select, deploy
Features.mqh 1093 indicator creation + per-bar input feature vector
ChartUI.mqh 634 arrows, arrow persistence, status panel, cleanup
Persistence.mqh 492 .stats/.cfg sidecars, CPU-inference validation, copy
OnlineLearning.mqh 461 live continual learning, EMA shadow, OOS simulator
Labels.mqh 309 ZigZag pivot labels, async label-cache prebuild
AutoTune.mqh 275 genetic tuner (population, crossover, halving)
Inference.mqh 235 softmax, prior calibration, class priors
ExpertSignalAIBase.mqh 8216 -> 3131 (declaration + topology build only)
This is a pure relocation - verified mechanically, not by eye: HEAD's file
reconstructed from the eight partials plus the surviving remainder is
byte-identical to HEAD, span for span (scratchpad verify_split.py). No
declaration moved, no signature changed, no code rewritten, so behaviour is
unchanged by construction.
Compiles 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 00:42:45 -04:00
|
|
|
}
|
|
|
|
|
//+------------------------------------------------------------------+
|
|
|
|
|
//| |
|
|
|
|
|
//+------------------------------------------------------------------+
|
|
|
|
|
double CExpertSignalAIBase::ApplyClassificationSoftmax(void)
|
|
|
|
|
{
|
2026-08-22 00:25:52 -04:00
|
|
|
//--- A non-finite logit poisons everything downstream: maxLogit, every exp(), the sum, and all
|
|
|
|
|
//--- three probabilities become NaN, and since NaN fails every comparison the two directional
|
|
|
|
|
//--- tests below are both false - so a NaN'd net returns Neutral on every bar forever and looks
|
|
|
|
|
//--- EXACTLY like a model that has simply gone quiet.
|
2026-08-01 11:27:28 -04:00
|
|
|
if(!MathIsValidNumber(TempData.At(0)) || !MathIsValidNumber(TempData.At(1)) || !MathIsValidNumber(TempData.At(2)))
|
|
|
|
|
{
|
|
|
|
|
static int nanLogitReports = 0;
|
|
|
|
|
// Bounded: this cannot heal on its own (the weights are already corrupt), so unlimited logging
|
|
|
|
|
// would fill the journal for as long as the chart stays attached. Three is enough to prove it.
|
|
|
|
|
if(nanLogitReports < 3)
|
|
|
|
|
{
|
|
|
|
|
nanLogitReports++;
|
|
|
|
|
PrintFormat("%s: %s NON-FINITE network output (%g / %g / %g) - forcing Neutral. The weights are "
|
|
|
|
|
"corrupt; reload the last good .nnw or reset and retrain. Report %d of 3.",
|
|
|
|
|
__FUNCTION__, ID, TempData.At(0), TempData.At(1), TempData.At(2), nanLogitReports);
|
|
|
|
|
}
|
|
|
|
|
return 0;
|
|
|
|
|
}
|
refactor: split CExpertSignalAIBase implementation by responsibility
ExpertSignalAIBase.mqh was 8216 lines: the class declaration followed by 87
method bodies covering training, labelling, feature extraction, persistence,
chart drawing, online learning, the GA auto-tuner and inference, all in one
file. Train() alone is 1492 lines; a change to arrow drawing meant scrolling
past the era loop.
Moved the bodies into Expert\AIBase\, included at the bottom of the original
after the class declaration:
Training.mqh 1607 era loop, plateau ladder, checkpoint select, deploy
Features.mqh 1093 indicator creation + per-bar input feature vector
ChartUI.mqh 634 arrows, arrow persistence, status panel, cleanup
Persistence.mqh 492 .stats/.cfg sidecars, CPU-inference validation, copy
OnlineLearning.mqh 461 live continual learning, EMA shadow, OOS simulator
Labels.mqh 309 ZigZag pivot labels, async label-cache prebuild
AutoTune.mqh 275 genetic tuner (population, crossover, halving)
Inference.mqh 235 softmax, prior calibration, class priors
ExpertSignalAIBase.mqh 8216 -> 3131 (declaration + topology build only)
This is a pure relocation - verified mechanically, not by eye: HEAD's file
reconstructed from the eight partials plus the surviving remainder is
byte-identical to HEAD, span for span (scratchpad verify_split.py). No
declaration moved, no signature changed, no code rewritten, so behaviour is
unchanged by construction.
Compiles 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 00:42:45 -04:00
|
|
|
// CLASS_LOGIT_SCALE (AI\Network.mqh) must match the training-gradient softmax in
|
|
|
|
|
// backProp/backPropOCL exactly - this is the same normalization the loss was trained against.
|
|
|
|
|
double maxLogit = CLASS_LOGIT_SCALE * MathMax(TempData.At(0), MathMax(TempData.At(1), TempData.At(2)));
|
|
|
|
|
double sum = 0;
|
|
|
|
|
for(int res = 0; res < 3; res++)
|
|
|
|
|
{
|
|
|
|
|
double temp = exp(CLASS_LOGIT_SCALE * TempData.At(res) - maxLogit);
|
|
|
|
|
sum += temp;
|
|
|
|
|
TempData.Update(res, temp);
|
|
|
|
|
}
|
|
|
|
|
for(int res = 0; res < 3; res++)
|
|
|
|
|
TempData.Update(res, TempData.At(res) / sum);
|
|
|
|
|
double pBuy = TempData.At(0);
|
|
|
|
|
double pSell = TempData.At(1);
|
|
|
|
|
double pNeutral = TempData.At(2);
|
2026-08-22 00:25:52 -04:00
|
|
|
//--- TempData.Maximum(0,3) scans left-to-right and keeps the FIRST index on a tie, so any tie
|
|
|
|
|
//--- (including the degenerate all-equal 0.3333/0.3333/0.3333 case from a collapsed/untrained
|
|
|
|
|
//--- net) always resolved to Buy (index 0) - silently turning "the model has no idea" into a
|
|
|
|
|
//--- directional trade.
|
refactor: split CExpertSignalAIBase implementation by responsibility
ExpertSignalAIBase.mqh was 8216 lines: the class declaration followed by 87
method bodies covering training, labelling, feature extraction, persistence,
chart drawing, online learning, the GA auto-tuner and inference, all in one
file. Train() alone is 1492 lines; a change to arrow drawing meant scrolling
past the era loop.
Moved the bodies into Expert\AIBase\, included at the bottom of the original
after the class declaration:
Training.mqh 1607 era loop, plateau ladder, checkpoint select, deploy
Features.mqh 1093 indicator creation + per-bar input feature vector
ChartUI.mqh 634 arrows, arrow persistence, status panel, cleanup
Persistence.mqh 492 .stats/.cfg sidecars, CPU-inference validation, copy
OnlineLearning.mqh 461 live continual learning, EMA shadow, OOS simulator
Labels.mqh 309 ZigZag pivot labels, async label-cache prebuild
AutoTune.mqh 275 genetic tuner (population, crossover, halving)
Inference.mqh 235 softmax, prior calibration, class priors
ExpertSignalAIBase.mqh 8216 -> 3131 (declaration + topology build only)
This is a pure relocation - verified mechanically, not by eye: HEAD's file
reconstructed from the eight partials plus the surviving remainder is
byte-identical to HEAD, span for span (scratchpad verify_split.py). No
declaration moved, no signature changed, no code rewritten, so behaviour is
unchanged by construction.
Compiles 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 00:42:45 -04:00
|
|
|
if(pBuy > pSell && pBuy > pNeutral)
|
|
|
|
|
return pBuy; // Buy signal
|
|
|
|
|
if(pSell > pBuy && pSell > pNeutral)
|
|
|
|
|
return -pSell; // Sell signal
|
|
|
|
|
return 0; // Neutral signal (also the fallback on any tie)
|
|
|
|
|
}
|
|
|
|
|
//+------------------------------------------------------------------+
|
|
|
|
|
//| Post-hoc logit adjustment (prior correction) of the 3-class |
|
2026-08-22 00:30:14 -04:00
|
|
|
//| decision. |
|
refactor: split CExpertSignalAIBase implementation by responsibility
ExpertSignalAIBase.mqh was 8216 lines: the class declaration followed by 87
method bodies covering training, labelling, feature extraction, persistence,
chart drawing, online learning, the GA auto-tuner and inference, all in one
file. Train() alone is 1492 lines; a change to arrow drawing meant scrolling
past the era loop.
Moved the bodies into Expert\AIBase\, included at the bottom of the original
after the class declaration:
Training.mqh 1607 era loop, plateau ladder, checkpoint select, deploy
Features.mqh 1093 indicator creation + per-bar input feature vector
ChartUI.mqh 634 arrows, arrow persistence, status panel, cleanup
Persistence.mqh 492 .stats/.cfg sidecars, CPU-inference validation, copy
OnlineLearning.mqh 461 live continual learning, EMA shadow, OOS simulator
Labels.mqh 309 ZigZag pivot labels, async label-cache prebuild
AutoTune.mqh 275 genetic tuner (population, crossover, halving)
Inference.mqh 235 softmax, prior calibration, class priors
ExpertSignalAIBase.mqh 8216 -> 3131 (declaration + topology build only)
This is a pure relocation - verified mechanically, not by eye: HEAD's file
reconstructed from the eight partials plus the surviving remainder is
byte-identical to HEAD, span for span (scratchpad verify_split.py). No
declaration moved, no signature changed, no code rewritten, so behaviour is
unchanged by construction.
Compiles 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 00:42:45 -04:00
|
|
|
//+------------------------------------------------------------------+
|
|
|
|
|
double CExpertSignalAIBase::AdjustedSignalFromSoftmax(void)
|
|
|
|
|
{
|
|
|
|
|
if(TempData.Total() < 3)
|
|
|
|
|
return 0.0;
|
|
|
|
|
double pBuy = TempData.At(0), pSell = TempData.At(1), pNeutral = TempData.At(2);
|
refactor(ai): nine class-imbalance inputs down to two
The imbalance section offered nine controls for one job. Audited against the
code, five of them did not do what their names said at the shipped defaults:
AILogitPriorStrength DEAD - Inference.mqh's post-hoc prior early-returns
whenever the adjusted loss is on, which is default.
OversampleParity DEAD in training - Training.mqh gated the replay loop
on !useLogitAdjustedLoss (correctly, citing Buda et
al. 2018). Live only in the online-learning path.
EnableMinorityReplay DEAD as replay. It survived ONLY as a focal-gamma
damper - "replay minority bars through pass-2
oversampling" was a focal-loss switch.
ConstrainReplay DEAD as a cap; it only chose damper 0.125 vs 0.25.
UseStaticPrior An exact duplicate of FreezePriorCalibration - the two
were OR'd together in the single place either is read.
So they were not five mechanisms fighting; they were one mechanism plus eight
knobs that mostly described machinery that no longer ran. That is worse than
a real conflict, because the log agreed with the names: the label-cache line
printed "reps up to 28x (90% parity) (seeding era 0's class-balance
oversampling)" on every run, describing an oversampling pass that had been
switched off. It is fixed here too - it cost this session a wrong diagnosis.
The one genuine redundancy was focal loss, running at gamma*0.125 alongside
the adjusted loss: two corrections on the same axis, the exact stacking
failure this file already cited Buda et al. for in two other places, damped
by a replay flag whose replay path was itself dead. Removed rather than
re-tuned. The plateau ladder is unaffected - its escape is the learning-rate
warm restart; the gamma anneal beside it only ever stepped toward zero.
WHAT REMAINS is logit-adjusted loss (Menon et al. 2021) plus a prior freeze:
LogitAdjustTau 0 = off; replaces the separate EnableLogitAdjusted-
Loss boolean, since a strength dial where 0 already
means off does not need an on/off switch beside it.
FreezePriorCalibration unchanged.
It is the only one of the six corrections with a consistency guarantee, and
it is consistent for exactly the balanced-error metric checkpoint selection
already ranks on - so the loss and the deploy decision optimize one thing.
The online continual-learning path keeps its own alpha-balanced focal weight,
now as constants pinned to the removed inputs' shipped defaults, so its
behaviour is unchanged. It legitimately needs its own correction:
ApplyLogitAdjustment() only runs inside a training run, so a deployed model
that was reloaded carries no logit offsets and would otherwise stream 31:1
data into itself uncorrected.
The weights-filename fingerprint is BYTE-IDENTICAL. The focal slot was a
double fed to a %d conversion and had always emitted a literal 0; the |MR:
segment is written as the constant its shipped defaults produced. Dropping
either would have re-keyed every model and forced a from-scratch retrain of
the one topology currently converged and trading.
Also removed as orphans: FOCAL_GAMMA_PRESET, MAX_OVERSAMPLE_REPLICAS,
OVERSAMPLE_PARITY_FRACTION, PLATEAU_GAMMA_STEP, and the now-unreachable
"neutralized by prior correction" diagnostic.
Both builds compile 0 errors, 0 warnings. No retrain forced.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 11:46:57 -04:00
|
|
|
//--- Strict majority, ties to Neutral - the same rule as ApplyClassificationSoftmax(). The returned
|
|
|
|
|
//--- magnitude is a genuine probability, which the confidence floor and ConfidenceTier() read.
|
feat: fitted directional confidence threshold - selectivity gets a mechanism
The training loss and the selection metric wanted different things and only
the second one knew it. Logit-adjusted cross-entropy has no term for "how
often should I trade", so the head calls a direction on 87-91% of bars. The
selection metric is precision x coverage credit, saturating at the coverage
floor - above the floor extra calls earn NOTHING and only precision counts.
So selection wanted few good calls, the loss produced many mediocre ones, and
all selection could do was pick the least-bad era out of what it was handed.
Nothing pushed the model toward selectivity.
This gives the decision RULE the policy instead of distorting the loss (which
is estimating class probabilities correctly, and a probability estimate should
not be bent to encode a trading policy - Elkan 2001: estimate, then choose the
operating point separately). AdjustedSignalFromSoftmax now abstains unless the
winning direction's softmax margin over its best rival clears a fitted
threshold. Margin, not the winning probability: the latter moves with overall
calibration rather than with how close the decision actually was.
Fitted on IS, applied to OOS and live. Pass 2 already forward-passes every IS
sample, so the margin histogram is harvested there for free (primary
occurrences only, so the oversampled replay queue cannot skew the operating
point); the fit runs at the end of pass 2, BEFORE pass 3, so the deploy gate
grades the thresholded model on bars the threshold never saw. Fitting on
pass 3's own predictions would be choosing the operating point on the data
being graded - the best-of-N error corrected in five other places here.
Objective: maximise IS directional precision subject to still clearing the
SAME coverage floor the deploy gate uses (base rate x 0.25, re-derived
locally so the two cannot drift apart). Swept top-down in one pass; ties go
to the LOWER threshold, since equal precision for less coverage is strictly
worse. Under DIR_CONF_MIN_FIT_CALLS (200) it runs unthresholded rather than
on a guess.
The threshold is part of the MODEL, not the run: captured with
Net.CaptureWeights(), restored with the weights at both restore sites, and
appended to the .cfg under the same length-guard convention so a deployed
model reloads at the operating point its gate actually cleared. A pre-2026-08-09
.cfg reads 0.0, which is exactly the behaviour it was trained under.
Per-era line now prints "@margin>=X.XX" next to coverage, so a coverage drop
can be attributed to the operating point rather than guessed at.
Both build variants compile 0 errors / 0 warnings.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 15:04:37 -04:00
|
|
|
bool wantBuy = (pBuy > pSell && pBuy > pNeutral);
|
|
|
|
|
bool wantSell = (pSell > pBuy && pSell > pNeutral);
|
|
|
|
|
if(!wantBuy && !wantSell)
|
|
|
|
|
return 0.0;
|
2026-08-22 00:25:52 -04:00
|
|
|
//--- OPERATING POINT (2026-08-09). Argmax alone answers "which class is most likely"; it does
|
|
|
|
|
//--- not answer "is this worth trading", and those are different questions whenever the top two
|
|
|
|
|
//--- classes are nearly tied.
|
feat: fitted directional confidence threshold - selectivity gets a mechanism
The training loss and the selection metric wanted different things and only
the second one knew it. Logit-adjusted cross-entropy has no term for "how
often should I trade", so the head calls a direction on 87-91% of bars. The
selection metric is precision x coverage credit, saturating at the coverage
floor - above the floor extra calls earn NOTHING and only precision counts.
So selection wanted few good calls, the loss produced many mediocre ones, and
all selection could do was pick the least-bad era out of what it was handed.
Nothing pushed the model toward selectivity.
This gives the decision RULE the policy instead of distorting the loss (which
is estimating class probabilities correctly, and a probability estimate should
not be bent to encode a trading policy - Elkan 2001: estimate, then choose the
operating point separately). AdjustedSignalFromSoftmax now abstains unless the
winning direction's softmax margin over its best rival clears a fitted
threshold. Margin, not the winning probability: the latter moves with overall
calibration rather than with how close the decision actually was.
Fitted on IS, applied to OOS and live. Pass 2 already forward-passes every IS
sample, so the margin histogram is harvested there for free (primary
occurrences only, so the oversampled replay queue cannot skew the operating
point); the fit runs at the end of pass 2, BEFORE pass 3, so the deploy gate
grades the thresholded model on bars the threshold never saw. Fitting on
pass 3's own predictions would be choosing the operating point on the data
being graded - the best-of-N error corrected in five other places here.
Objective: maximise IS directional precision subject to still clearing the
SAME coverage floor the deploy gate uses (base rate x 0.25, re-derived
locally so the two cannot drift apart). Swept top-down in one pass; ties go
to the LOWER threshold, since equal precision for less coverage is strictly
worse. Under DIR_CONF_MIN_FIT_CALLS (200) it runs unthresholded rather than
on a guess.
The threshold is part of the MODEL, not the run: captured with
Net.CaptureWeights(), restored with the weights at both restore sites, and
appended to the .cfg under the same length-guard convention so a deployed
model reloads at the operating point its gate actually cleared. A pre-2026-08-09
.cfg reads 0.0, which is exactly the behaviour it was trained under.
Per-era line now prints "@margin>=X.XX" next to coverage, so a coverage drop
can be attributed to the operating point rather than guessed at.
Both build variants compile 0 errors / 0 warnings.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 15:04:37 -04:00
|
|
|
if(m_dirConfThreshold > 0.0)
|
|
|
|
|
{
|
|
|
|
|
double win = wantBuy ? pBuy : pSell;
|
|
|
|
|
double rival = wantBuy ? MathMax(pSell, pNeutral) : MathMax(pBuy, pNeutral);
|
|
|
|
|
if((win - rival) < m_dirConfThreshold)
|
|
|
|
|
return 0.0;
|
|
|
|
|
}
|
|
|
|
|
return wantBuy ? pBuy : -pSell;
|
|
|
|
|
}
|
|
|
|
|
//+------------------------------------------------------------------+
|
|
|
|
|
//| The statistic the operating point is expressed in - see the |
|
|
|
|
|
//| declaration. Reads the softmax ALREADY in TempData, so callers |
|
|
|
|
|
//| must have run ApplyClassificationSoftmax() first. |
|
|
|
|
|
//+------------------------------------------------------------------+
|
|
|
|
|
double CExpertSignalAIBase::DirectionalMargin(void)
|
|
|
|
|
{
|
|
|
|
|
if(TempData.Total() < 3)
|
|
|
|
|
return -1.0;
|
|
|
|
|
double pBuy = TempData.At(0), pSell = TempData.At(1), pNeutral = TempData.At(2);
|
refactor(ai): nine class-imbalance inputs down to two
The imbalance section offered nine controls for one job. Audited against the
code, five of them did not do what their names said at the shipped defaults:
AILogitPriorStrength DEAD - Inference.mqh's post-hoc prior early-returns
whenever the adjusted loss is on, which is default.
OversampleParity DEAD in training - Training.mqh gated the replay loop
on !useLogitAdjustedLoss (correctly, citing Buda et
al. 2018). Live only in the online-learning path.
EnableMinorityReplay DEAD as replay. It survived ONLY as a focal-gamma
damper - "replay minority bars through pass-2
oversampling" was a focal-loss switch.
ConstrainReplay DEAD as a cap; it only chose damper 0.125 vs 0.25.
UseStaticPrior An exact duplicate of FreezePriorCalibration - the two
were OR'd together in the single place either is read.
So they were not five mechanisms fighting; they were one mechanism plus eight
knobs that mostly described machinery that no longer ran. That is worse than
a real conflict, because the log agreed with the names: the label-cache line
printed "reps up to 28x (90% parity) (seeding era 0's class-balance
oversampling)" on every run, describing an oversampling pass that had been
switched off. It is fixed here too - it cost this session a wrong diagnosis.
The one genuine redundancy was focal loss, running at gamma*0.125 alongside
the adjusted loss: two corrections on the same axis, the exact stacking
failure this file already cited Buda et al. for in two other places, damped
by a replay flag whose replay path was itself dead. Removed rather than
re-tuned. The plateau ladder is unaffected - its escape is the learning-rate
warm restart; the gamma anneal beside it only ever stepped toward zero.
WHAT REMAINS is logit-adjusted loss (Menon et al. 2021) plus a prior freeze:
LogitAdjustTau 0 = off; replaces the separate EnableLogitAdjusted-
Loss boolean, since a strength dial where 0 already
means off does not need an on/off switch beside it.
FreezePriorCalibration unchanged.
It is the only one of the six corrections with a consistency guarantee, and
it is consistent for exactly the balanced-error metric checkpoint selection
already ranks on - so the loss and the deploy decision optimize one thing.
The online continual-learning path keeps its own alpha-balanced focal weight,
now as constants pinned to the removed inputs' shipped defaults, so its
behaviour is unchanged. It legitimately needs its own correction:
ApplyLogitAdjustment() only runs inside a training run, so a deployed model
that was reloaded carries no logit offsets and would otherwise stream 31:1
data into itself uncorrected.
The weights-filename fingerprint is BYTE-IDENTICAL. The focal slot was a
double fed to a %d conversion and had always emitted a literal 0; the |MR:
segment is written as the constant its shipped defaults produced. Dropping
either would have re-keyed every model and forced a from-scratch retrain of
the one topology currently converged and trading.
Also removed as orphans: FOCAL_GAMMA_PRESET, MAX_OVERSAMPLE_REPLICAS,
OVERSAMPLE_PARITY_FRACTION, PLATEAU_GAMMA_STEP, and the now-unreachable
"neutralized by prior correction" diagnostic.
Both builds compile 0 errors, 0 warnings. No retrain forced.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 11:46:57 -04:00
|
|
|
if(pBuy > pSell && pBuy > pNeutral)
|
feat: fitted directional confidence threshold - selectivity gets a mechanism
The training loss and the selection metric wanted different things and only
the second one knew it. Logit-adjusted cross-entropy has no term for "how
often should I trade", so the head calls a direction on 87-91% of bars. The
selection metric is precision x coverage credit, saturating at the coverage
floor - above the floor extra calls earn NOTHING and only precision counts.
So selection wanted few good calls, the loss produced many mediocre ones, and
all selection could do was pick the least-bad era out of what it was handed.
Nothing pushed the model toward selectivity.
This gives the decision RULE the policy instead of distorting the loss (which
is estimating class probabilities correctly, and a probability estimate should
not be bent to encode a trading policy - Elkan 2001: estimate, then choose the
operating point separately). AdjustedSignalFromSoftmax now abstains unless the
winning direction's softmax margin over its best rival clears a fitted
threshold. Margin, not the winning probability: the latter moves with overall
calibration rather than with how close the decision actually was.
Fitted on IS, applied to OOS and live. Pass 2 already forward-passes every IS
sample, so the margin histogram is harvested there for free (primary
occurrences only, so the oversampled replay queue cannot skew the operating
point); the fit runs at the end of pass 2, BEFORE pass 3, so the deploy gate
grades the thresholded model on bars the threshold never saw. Fitting on
pass 3's own predictions would be choosing the operating point on the data
being graded - the best-of-N error corrected in five other places here.
Objective: maximise IS directional precision subject to still clearing the
SAME coverage floor the deploy gate uses (base rate x 0.25, re-derived
locally so the two cannot drift apart). Swept top-down in one pass; ties go
to the LOWER threshold, since equal precision for less coverage is strictly
worse. Under DIR_CONF_MIN_FIT_CALLS (200) it runs unthresholded rather than
on a guess.
The threshold is part of the MODEL, not the run: captured with
Net.CaptureWeights(), restored with the weights at both restore sites, and
appended to the .cfg under the same length-guard convention so a deployed
model reloads at the operating point its gate actually cleared. A pre-2026-08-09
.cfg reads 0.0, which is exactly the behaviour it was trained under.
Per-era line now prints "@margin>=X.XX" next to coverage, so a coverage drop
can be attributed to the operating point rather than guessed at.
Both build variants compile 0 errors / 0 warnings.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 15:04:37 -04:00
|
|
|
return pBuy - MathMax(pSell, pNeutral);
|
refactor(ai): nine class-imbalance inputs down to two
The imbalance section offered nine controls for one job. Audited against the
code, five of them did not do what their names said at the shipped defaults:
AILogitPriorStrength DEAD - Inference.mqh's post-hoc prior early-returns
whenever the adjusted loss is on, which is default.
OversampleParity DEAD in training - Training.mqh gated the replay loop
on !useLogitAdjustedLoss (correctly, citing Buda et
al. 2018). Live only in the online-learning path.
EnableMinorityReplay DEAD as replay. It survived ONLY as a focal-gamma
damper - "replay minority bars through pass-2
oversampling" was a focal-loss switch.
ConstrainReplay DEAD as a cap; it only chose damper 0.125 vs 0.25.
UseStaticPrior An exact duplicate of FreezePriorCalibration - the two
were OR'd together in the single place either is read.
So they were not five mechanisms fighting; they were one mechanism plus eight
knobs that mostly described machinery that no longer ran. That is worse than
a real conflict, because the log agreed with the names: the label-cache line
printed "reps up to 28x (90% parity) (seeding era 0's class-balance
oversampling)" on every run, describing an oversampling pass that had been
switched off. It is fixed here too - it cost this session a wrong diagnosis.
The one genuine redundancy was focal loss, running at gamma*0.125 alongside
the adjusted loss: two corrections on the same axis, the exact stacking
failure this file already cited Buda et al. for in two other places, damped
by a replay flag whose replay path was itself dead. Removed rather than
re-tuned. The plateau ladder is unaffected - its escape is the learning-rate
warm restart; the gamma anneal beside it only ever stepped toward zero.
WHAT REMAINS is logit-adjusted loss (Menon et al. 2021) plus a prior freeze:
LogitAdjustTau 0 = off; replaces the separate EnableLogitAdjusted-
Loss boolean, since a strength dial where 0 already
means off does not need an on/off switch beside it.
FreezePriorCalibration unchanged.
It is the only one of the six corrections with a consistency guarantee, and
it is consistent for exactly the balanced-error metric checkpoint selection
already ranks on - so the loss and the deploy decision optimize one thing.
The online continual-learning path keeps its own alpha-balanced focal weight,
now as constants pinned to the removed inputs' shipped defaults, so its
behaviour is unchanged. It legitimately needs its own correction:
ApplyLogitAdjustment() only runs inside a training run, so a deployed model
that was reloaded carries no logit offsets and would otherwise stream 31:1
data into itself uncorrected.
The weights-filename fingerprint is BYTE-IDENTICAL. The focal slot was a
double fed to a %d conversion and had always emitted a literal 0; the |MR:
segment is written as the constant its shipped defaults produced. Dropping
either would have re-keyed every model and forced a from-scratch retrain of
the one topology currently converged and trading.
Also removed as orphans: FOCAL_GAMMA_PRESET, MAX_OVERSAMPLE_REPLICAS,
OVERSAMPLE_PARITY_FRACTION, PLATEAU_GAMMA_STEP, and the now-unreachable
"neutralized by prior correction" diagnostic.
Both builds compile 0 errors, 0 warnings. No retrain forced.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 11:46:57 -04:00
|
|
|
if(pSell > pBuy && pSell > pNeutral)
|
feat: fitted directional confidence threshold - selectivity gets a mechanism
The training loss and the selection metric wanted different things and only
the second one knew it. Logit-adjusted cross-entropy has no term for "how
often should I trade", so the head calls a direction on 87-91% of bars. The
selection metric is precision x coverage credit, saturating at the coverage
floor - above the floor extra calls earn NOTHING and only precision counts.
So selection wanted few good calls, the loss produced many mediocre ones, and
all selection could do was pick the least-bad era out of what it was handed.
Nothing pushed the model toward selectivity.
This gives the decision RULE the policy instead of distorting the loss (which
is estimating class probabilities correctly, and a probability estimate should
not be bent to encode a trading policy - Elkan 2001: estimate, then choose the
operating point separately). AdjustedSignalFromSoftmax now abstains unless the
winning direction's softmax margin over its best rival clears a fitted
threshold. Margin, not the winning probability: the latter moves with overall
calibration rather than with how close the decision actually was.
Fitted on IS, applied to OOS and live. Pass 2 already forward-passes every IS
sample, so the margin histogram is harvested there for free (primary
occurrences only, so the oversampled replay queue cannot skew the operating
point); the fit runs at the end of pass 2, BEFORE pass 3, so the deploy gate
grades the thresholded model on bars the threshold never saw. Fitting on
pass 3's own predictions would be choosing the operating point on the data
being graded - the best-of-N error corrected in five other places here.
Objective: maximise IS directional precision subject to still clearing the
SAME coverage floor the deploy gate uses (base rate x 0.25, re-derived
locally so the two cannot drift apart). Swept top-down in one pass; ties go
to the LOWER threshold, since equal precision for less coverage is strictly
worse. Under DIR_CONF_MIN_FIT_CALLS (200) it runs unthresholded rather than
on a guess.
The threshold is part of the MODEL, not the run: captured with
Net.CaptureWeights(), restored with the weights at both restore sites, and
appended to the .cfg under the same length-guard convention so a deployed
model reloads at the operating point its gate actually cleared. A pre-2026-08-09
.cfg reads 0.0, which is exactly the behaviour it was trained under.
Per-era line now prints "@margin>=X.XX" next to coverage, so a coverage drop
can be attributed to the operating point rather than guessed at.
Both build variants compile 0 errors / 0 warnings.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 15:04:37 -04:00
|
|
|
return pSell - MathMax(pBuy, pNeutral);
|
|
|
|
|
return -1.0; // Neutral won: no directional call, so no operating point applies
|
|
|
|
|
}
|
|
|
|
|
//+------------------------------------------------------------------+
|
fix: the operating point was fitted on bars the net had memorized
FitDirConfThreshold harvested its margin histogram from pass 2's own
backprop samples. Pairing every fit against the same era's OOS result
shows what that measured:
PAI era 1 IS 25% cov @ 66.1% (-0.8pp) -> OOS 64% (-3pp) gap +2.1pp
PAI era 76 IS 90% cov @ 79.6% (+12.7pp) -> OOS 65% (-2pp) gap +14.6pp
LSTM era 9 IS 77% cov @ 81.6% (+14.6pp) -> OOS 63% (-4pp) gap +18.6pp
The gap grows monotonically while OOS stays flat, so within a handful of
eras the curve stops describing behaviour on unseen bars. That is fatal
here specifically, because the objective branches on the SIGN of
(p - break-even): the memorized curve reads +12pp at 95% coverage, so
coverage x (p - p0) correctly maximises coverage and returns ~0.02 - fire
on every bar. The "p < p0 -> get more selective" branch, which is the
actual regime and the entire point of 983a6a3, could never fire because IS
never showed p < p0.
Carve a calibration slice out of the IS span - DIR_CONF_CALIB_PCT_OF_IS,
purged from backprop by one label horizon on BOTH sides (the far-side
purge is not optional: without it the newest training bars carry labels
partly decided by price action inside the slice, putting the memorization
straight back into the curve). Score it in a new chunked pass 2.5, after
pass 2 has trained and before pass 3 grades - the only position where the
histogram is simultaneously not-trained-on, not-graded, and current with
the weights it will be applied to.
Costs 15% of the training data. Worth it beyond honesty: the deploy gate
needs dirPrecPct > chance + EDGE_MIN_SIGMAS*SE, and a threshold pinned
near zero dilutes any edge concentrated in the confident bars across every
bar the model calls, driving dirPrecPct toward chance by construction. A
threshold that can be selective is the only mechanism by which a small,
concentrated edge could ever clear that gate.
Also: a sparse histogram now KEEPS the previous threshold instead of
resetting to 0.0. A failed measurement must not decay to the most exposed
setting in the range.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-10 15:58:18 -04:00
|
|
|
//| Clear the margin histogram at the start of the calibration walk. |
|
feat: fitted directional confidence threshold - selectivity gets a mechanism
The training loss and the selection metric wanted different things and only
the second one knew it. Logit-adjusted cross-entropy has no term for "how
often should I trade", so the head calls a direction on 87-91% of bars. The
selection metric is precision x coverage credit, saturating at the coverage
floor - above the floor extra calls earn NOTHING and only precision counts.
So selection wanted few good calls, the loss produced many mediocre ones, and
all selection could do was pick the least-bad era out of what it was handed.
Nothing pushed the model toward selectivity.
This gives the decision RULE the policy instead of distorting the loss (which
is estimating class probabilities correctly, and a probability estimate should
not be bent to encode a trading policy - Elkan 2001: estimate, then choose the
operating point separately). AdjustedSignalFromSoftmax now abstains unless the
winning direction's softmax margin over its best rival clears a fitted
threshold. Margin, not the winning probability: the latter moves with overall
calibration rather than with how close the decision actually was.
Fitted on IS, applied to OOS and live. Pass 2 already forward-passes every IS
sample, so the margin histogram is harvested there for free (primary
occurrences only, so the oversampled replay queue cannot skew the operating
point); the fit runs at the end of pass 2, BEFORE pass 3, so the deploy gate
grades the thresholded model on bars the threshold never saw. Fitting on
pass 3's own predictions would be choosing the operating point on the data
being graded - the best-of-N error corrected in five other places here.
Objective: maximise IS directional precision subject to still clearing the
SAME coverage floor the deploy gate uses (base rate x 0.25, re-derived
locally so the two cannot drift apart). Swept top-down in one pass; ties go
to the LOWER threshold, since equal precision for less coverage is strictly
worse. Under DIR_CONF_MIN_FIT_CALLS (200) it runs unthresholded rather than
on a guess.
The threshold is part of the MODEL, not the run: captured with
Net.CaptureWeights(), restored with the weights at both restore sites, and
appended to the .cfg under the same length-guard convention so a deployed
model reloads at the operating point its gate actually cleared. A pre-2026-08-09
.cfg reads 0.0, which is exactly the behaviour it was trained under.
Per-era line now prints "@margin>=X.XX" next to coverage, so a coverage drop
can be attributed to the operating point rather than guessed at.
Both build variants compile 0 errors / 0 warnings.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 15:04:37 -04:00
|
|
|
//+------------------------------------------------------------------+
|
|
|
|
|
void CExpertSignalAIBase::ResetDirConfHistogram(void)
|
|
|
|
|
{
|
|
|
|
|
ArrayInitialize(m_dirConfBinCalls, 0);
|
|
|
|
|
ArrayInitialize(m_dirConfBinHits, 0);
|
|
|
|
|
m_dirConfPrimaryBars = 0;
|
|
|
|
|
}
|
|
|
|
|
//+------------------------------------------------------------------+
|
fix: the operating point was fitted on bars the net had memorized
FitDirConfThreshold harvested its margin histogram from pass 2's own
backprop samples. Pairing every fit against the same era's OOS result
shows what that measured:
PAI era 1 IS 25% cov @ 66.1% (-0.8pp) -> OOS 64% (-3pp) gap +2.1pp
PAI era 76 IS 90% cov @ 79.6% (+12.7pp) -> OOS 65% (-2pp) gap +14.6pp
LSTM era 9 IS 77% cov @ 81.6% (+14.6pp) -> OOS 63% (-4pp) gap +18.6pp
The gap grows monotonically while OOS stays flat, so within a handful of
eras the curve stops describing behaviour on unseen bars. That is fatal
here specifically, because the objective branches on the SIGN of
(p - break-even): the memorized curve reads +12pp at 95% coverage, so
coverage x (p - p0) correctly maximises coverage and returns ~0.02 - fire
on every bar. The "p < p0 -> get more selective" branch, which is the
actual regime and the entire point of 983a6a3, could never fire because IS
never showed p < p0.
Carve a calibration slice out of the IS span - DIR_CONF_CALIB_PCT_OF_IS,
purged from backprop by one label horizon on BOTH sides (the far-side
purge is not optional: without it the newest training bars carry labels
partly decided by price action inside the slice, putting the memorization
straight back into the curve). Score it in a new chunked pass 2.5, after
pass 2 has trained and before pass 3 grades - the only position where the
histogram is simultaneously not-trained-on, not-graded, and current with
the weights it will be applied to.
Costs 15% of the training data. Worth it beyond honesty: the deploy gate
needs dirPrecPct > chance + EDGE_MIN_SIGMAS*SE, and a threshold pinned
near zero dilutes any edge concentrated in the confident bars across every
bar the model calls, driving dirPrecPct toward chance by construction. A
threshold that can be selective is the only mechanism by which a small,
concentrated edge could ever clear that gate.
Also: a sparse histogram now KEEPS the previous threshold instead of
resetting to 0.0. A failed measurement must not decay to the most exposed
setting in the range.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-10 15:58:18 -04:00
|
|
|
//| One calibration sample. isPrimaryBar survives from when this was |
|
|
|
|
|
//| harvested inside pass 2's oversampled replay queue, where counting |
|
|
|
|
|
//| duplicated minority bars would have fitted the operating point to |
|
|
|
|
|
//| a class balance the live model never sees (the same correction |
|
|
|
|
|
//| m_cumIsTotal makes - see its note in Training.mqh). The calibration |
|
|
|
|
|
//| walk visits each bar exactly once and passes true; the parameter |
|
|
|
|
|
//| stays so any future caller must state which it is. |
|
feat: fitted directional confidence threshold - selectivity gets a mechanism
The training loss and the selection metric wanted different things and only
the second one knew it. Logit-adjusted cross-entropy has no term for "how
often should I trade", so the head calls a direction on 87-91% of bars. The
selection metric is precision x coverage credit, saturating at the coverage
floor - above the floor extra calls earn NOTHING and only precision counts.
So selection wanted few good calls, the loss produced many mediocre ones, and
all selection could do was pick the least-bad era out of what it was handed.
Nothing pushed the model toward selectivity.
This gives the decision RULE the policy instead of distorting the loss (which
is estimating class probabilities correctly, and a probability estimate should
not be bent to encode a trading policy - Elkan 2001: estimate, then choose the
operating point separately). AdjustedSignalFromSoftmax now abstains unless the
winning direction's softmax margin over its best rival clears a fitted
threshold. Margin, not the winning probability: the latter moves with overall
calibration rather than with how close the decision actually was.
Fitted on IS, applied to OOS and live. Pass 2 already forward-passes every IS
sample, so the margin histogram is harvested there for free (primary
occurrences only, so the oversampled replay queue cannot skew the operating
point); the fit runs at the end of pass 2, BEFORE pass 3, so the deploy gate
grades the thresholded model on bars the threshold never saw. Fitting on
pass 3's own predictions would be choosing the operating point on the data
being graded - the best-of-N error corrected in five other places here.
Objective: maximise IS directional precision subject to still clearing the
SAME coverage floor the deploy gate uses (base rate x 0.25, re-derived
locally so the two cannot drift apart). Swept top-down in one pass; ties go
to the LOWER threshold, since equal precision for less coverage is strictly
worse. Under DIR_CONF_MIN_FIT_CALLS (200) it runs unthresholded rather than
on a guess.
The threshold is part of the MODEL, not the run: captured with
Net.CaptureWeights(), restored with the weights at both restore sites, and
appended to the .cfg under the same length-guard convention so a deployed
model reloads at the operating point its gate actually cleared. A pre-2026-08-09
.cfg reads 0.0, which is exactly the behaviour it was trained under.
Per-era line now prints "@margin>=X.XX" next to coverage, so a coverage drop
can be attributed to the operating point rather than guessed at.
Both build variants compile 0 errors / 0 warnings.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 15:04:37 -04:00
|
|
|
//+------------------------------------------------------------------+
|
|
|
|
|
void CExpertSignalAIBase::AccumulateDirConfSample(double margin, bool wasCorrect, bool isPrimaryBar)
|
|
|
|
|
{
|
|
|
|
|
if(!isPrimaryBar)
|
|
|
|
|
return;
|
|
|
|
|
//--- Counted BEFORE the directional test: this is the coverage denominator, so it has to be every
|
|
|
|
|
//--- primary bar the model scored, including the ones it called Neutral. Using only directional
|
|
|
|
|
//--- bars would make coverage 100% by construction at every threshold.
|
|
|
|
|
m_dirConfPrimaryBars++;
|
|
|
|
|
if(margin < 0.0)
|
|
|
|
|
return; // Neutral won - not a directional call
|
|
|
|
|
int bin = (int)(margin * DIR_CONF_THRESHOLD_BINS);
|
|
|
|
|
if(bin < 0)
|
|
|
|
|
bin = 0;
|
|
|
|
|
if(bin >= DIR_CONF_THRESHOLD_BINS)
|
|
|
|
|
bin = DIR_CONF_THRESHOLD_BINS - 1; // margin can reach exactly 1.0
|
|
|
|
|
m_dirConfBinCalls[bin]++;
|
|
|
|
|
if(wasCorrect)
|
|
|
|
|
m_dirConfBinHits[bin]++;
|
|
|
|
|
}
|
|
|
|
|
//+------------------------------------------------------------------+
|
2026-08-22 00:30:14 -04:00
|
|
|
//| Choose the operating point: the margin at which the model calls |
|
|
|
|
|
//| a direction as often as a direction actually occurs. One pass, |
|
|
|
|
|
//| top down, so the running totals are "calls at or above this bin" |
|
|
|
|
|
//| - the set a threshold there admits. |
|
feat: fitted directional confidence threshold - selectivity gets a mechanism
The training loss and the selection metric wanted different things and only
the second one knew it. Logit-adjusted cross-entropy has no term for "how
often should I trade", so the head calls a direction on 87-91% of bars. The
selection metric is precision x coverage credit, saturating at the coverage
floor - above the floor extra calls earn NOTHING and only precision counts.
So selection wanted few good calls, the loss produced many mediocre ones, and
all selection could do was pick the least-bad era out of what it was handed.
Nothing pushed the model toward selectivity.
This gives the decision RULE the policy instead of distorting the loss (which
is estimating class probabilities correctly, and a probability estimate should
not be bent to encode a trading policy - Elkan 2001: estimate, then choose the
operating point separately). AdjustedSignalFromSoftmax now abstains unless the
winning direction's softmax margin over its best rival clears a fitted
threshold. Margin, not the winning probability: the latter moves with overall
calibration rather than with how close the decision actually was.
Fitted on IS, applied to OOS and live. Pass 2 already forward-passes every IS
sample, so the margin histogram is harvested there for free (primary
occurrences only, so the oversampled replay queue cannot skew the operating
point); the fit runs at the end of pass 2, BEFORE pass 3, so the deploy gate
grades the thresholded model on bars the threshold never saw. Fitting on
pass 3's own predictions would be choosing the operating point on the data
being graded - the best-of-N error corrected in five other places here.
Objective: maximise IS directional precision subject to still clearing the
SAME coverage floor the deploy gate uses (base rate x 0.25, re-derived
locally so the two cannot drift apart). Swept top-down in one pass; ties go
to the LOWER threshold, since equal precision for less coverage is strictly
worse. Under DIR_CONF_MIN_FIT_CALLS (200) it runs unthresholded rather than
on a guess.
The threshold is part of the MODEL, not the run: captured with
Net.CaptureWeights(), restored with the weights at both restore sites, and
appended to the .cfg under the same length-guard convention so a deployed
model reloads at the operating point its gate actually cleared. A pre-2026-08-09
.cfg reads 0.0, which is exactly the behaviour it was trained under.
Per-era line now prints "@margin>=X.XX" next to coverage, so a coverage drop
can be attributed to the operating point rather than guessed at.
Both build variants compile 0 errors / 0 warnings.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 15:04:37 -04:00
|
|
|
//+------------------------------------------------------------------+
|
|
|
|
|
void CExpertSignalAIBase::FitDirConfThreshold(void)
|
|
|
|
|
{
|
|
|
|
|
long totalCalls = 0;
|
|
|
|
|
for(int b = 0; b < DIR_CONF_THRESHOLD_BINS; b++)
|
|
|
|
|
totalCalls += m_dirConfBinCalls[b];
|
|
|
|
|
if(totalCalls < DIR_CONF_MIN_FIT_CALLS || m_dirConfPrimaryBars <= 0)
|
|
|
|
|
{
|
2026-08-22 00:25:52 -04:00
|
|
|
//--- Not enough evidence to place an operating point. KEEP THE PREVIOUS ONE - the old
|
|
|
|
|
//--- behaviour here was to reset to 0.0, which is "call a direction on every bar", the single
|
|
|
|
|
//--- most exposed setting in the range.
|
feat: fitted directional confidence threshold - selectivity gets a mechanism
The training loss and the selection metric wanted different things and only
the second one knew it. Logit-adjusted cross-entropy has no term for "how
often should I trade", so the head calls a direction on 87-91% of bars. The
selection metric is precision x coverage credit, saturating at the coverage
floor - above the floor extra calls earn NOTHING and only precision counts.
So selection wanted few good calls, the loss produced many mediocre ones, and
all selection could do was pick the least-bad era out of what it was handed.
Nothing pushed the model toward selectivity.
This gives the decision RULE the policy instead of distorting the loss (which
is estimating class probabilities correctly, and a probability estimate should
not be bent to encode a trading policy - Elkan 2001: estimate, then choose the
operating point separately). AdjustedSignalFromSoftmax now abstains unless the
winning direction's softmax margin over its best rival clears a fitted
threshold. Margin, not the winning probability: the latter moves with overall
calibration rather than with how close the decision actually was.
Fitted on IS, applied to OOS and live. Pass 2 already forward-passes every IS
sample, so the margin histogram is harvested there for free (primary
occurrences only, so the oversampled replay queue cannot skew the operating
point); the fit runs at the end of pass 2, BEFORE pass 3, so the deploy gate
grades the thresholded model on bars the threshold never saw. Fitting on
pass 3's own predictions would be choosing the operating point on the data
being graded - the best-of-N error corrected in five other places here.
Objective: maximise IS directional precision subject to still clearing the
SAME coverage floor the deploy gate uses (base rate x 0.25, re-derived
locally so the two cannot drift apart). Swept top-down in one pass; ties go
to the LOWER threshold, since equal precision for less coverage is strictly
worse. Under DIR_CONF_MIN_FIT_CALLS (200) it runs unthresholded rather than
on a guess.
The threshold is part of the MODEL, not the run: captured with
Net.CaptureWeights(), restored with the weights at both restore sites, and
appended to the .cfg under the same length-guard convention so a deployed
model reloads at the operating point its gate actually cleared. A pre-2026-08-09
.cfg reads 0.0, which is exactly the behaviour it was trained under.
Per-era line now prints "@margin>=X.XX" next to coverage, so a coverage drop
can be attributed to the operating point rather than guessed at.
Both build variants compile 0 errors / 0 warnings.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 15:04:37 -04:00
|
|
|
if(!m_dirConfSparseWarned)
|
|
|
|
|
{
|
|
|
|
|
m_dirConfSparseWarned = true;
|
fix: the operating point was fitted on bars the net had memorized
FitDirConfThreshold harvested its margin histogram from pass 2's own
backprop samples. Pairing every fit against the same era's OOS result
shows what that measured:
PAI era 1 IS 25% cov @ 66.1% (-0.8pp) -> OOS 64% (-3pp) gap +2.1pp
PAI era 76 IS 90% cov @ 79.6% (+12.7pp) -> OOS 65% (-2pp) gap +14.6pp
LSTM era 9 IS 77% cov @ 81.6% (+14.6pp) -> OOS 63% (-4pp) gap +18.6pp
The gap grows monotonically while OOS stays flat, so within a handful of
eras the curve stops describing behaviour on unseen bars. That is fatal
here specifically, because the objective branches on the SIGN of
(p - break-even): the memorized curve reads +12pp at 95% coverage, so
coverage x (p - p0) correctly maximises coverage and returns ~0.02 - fire
on every bar. The "p < p0 -> get more selective" branch, which is the
actual regime and the entire point of 983a6a3, could never fire because IS
never showed p < p0.
Carve a calibration slice out of the IS span - DIR_CONF_CALIB_PCT_OF_IS,
purged from backprop by one label horizon on BOTH sides (the far-side
purge is not optional: without it the newest training bars carry labels
partly decided by price action inside the slice, putting the memorization
straight back into the curve). Score it in a new chunked pass 2.5, after
pass 2 has trained and before pass 3 grades - the only position where the
histogram is simultaneously not-trained-on, not-graded, and current with
the weights it will be applied to.
Costs 15% of the training data. Worth it beyond honesty: the deploy gate
needs dirPrecPct > chance + EDGE_MIN_SIGMAS*SE, and a threshold pinned
near zero dilutes any edge concentrated in the confident bars across every
bar the model calls, driving dirPrecPct toward chance by construction. A
threshold that can be selective is the only mechanism by which a small,
concentrated edge could ever clear that gate.
Also: a sparse histogram now KEEPS the previous threshold instead of
resetting to 0.0. A failed measurement must not decay to the most exposed
setting in the range.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-10 15:58:18 -04:00
|
|
|
Print(ID + StringFormat(": directional confidence threshold NOT refitted - only %d directional "
|
|
|
|
|
"calls in the held-out calibration slice this era (need %d). Keeping "
|
|
|
|
|
"the previous operating point %.2f; this is normal for the first eras "
|
|
|
|
|
"and self-corrects as the model starts calling directions.",
|
|
|
|
|
(int)totalCalls, DIR_CONF_MIN_FIT_CALLS, m_dirConfThreshold));
|
feat: fitted directional confidence threshold - selectivity gets a mechanism
The training loss and the selection metric wanted different things and only
the second one knew it. Logit-adjusted cross-entropy has no term for "how
often should I trade", so the head calls a direction on 87-91% of bars. The
selection metric is precision x coverage credit, saturating at the coverage
floor - above the floor extra calls earn NOTHING and only precision counts.
So selection wanted few good calls, the loss produced many mediocre ones, and
all selection could do was pick the least-bad era out of what it was handed.
Nothing pushed the model toward selectivity.
This gives the decision RULE the policy instead of distorting the loss (which
is estimating class probabilities correctly, and a probability estimate should
not be bent to encode a trading policy - Elkan 2001: estimate, then choose the
operating point separately). AdjustedSignalFromSoftmax now abstains unless the
winning direction's softmax margin over its best rival clears a fitted
threshold. Margin, not the winning probability: the latter moves with overall
calibration rather than with how close the decision actually was.
Fitted on IS, applied to OOS and live. Pass 2 already forward-passes every IS
sample, so the margin histogram is harvested there for free (primary
occurrences only, so the oversampled replay queue cannot skew the operating
point); the fit runs at the end of pass 2, BEFORE pass 3, so the deploy gate
grades the thresholded model on bars the threshold never saw. Fitting on
pass 3's own predictions would be choosing the operating point on the data
being graded - the best-of-N error corrected in five other places here.
Objective: maximise IS directional precision subject to still clearing the
SAME coverage floor the deploy gate uses (base rate x 0.25, re-derived
locally so the two cannot drift apart). Swept top-down in one pass; ties go
to the LOWER threshold, since equal precision for less coverage is strictly
worse. Under DIR_CONF_MIN_FIT_CALLS (200) it runs unthresholded rather than
on a guess.
The threshold is part of the MODEL, not the run: captured with
Net.CaptureWeights(), restored with the weights at both restore sites, and
appended to the .cfg under the same length-guard convention so a deployed
model reloads at the operating point its gate actually cleared. A pre-2026-08-09
.cfg reads 0.0, which is exactly the behaviour it was trained under.
Per-era line now prints "@margin>=X.XX" next to coverage, so a coverage drop
can be attributed to the operating point rather than guessed at.
Both build variants compile 0 errors / 0 warnings.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 15:04:37 -04:00
|
|
|
}
|
|
|
|
|
return;
|
|
|
|
|
}
|
feat(calibration): fit the operating point on the label rate instead of on edge
The margin threshold now sits where the model calls a direction as often as a
direction actually occurs. Nothing else.
WHY THE OLD OBJECTIVE HAD TO GO. It maximised `coverage x (precision -
breakEven)`, and this function's own comments were already the case against
it: over 98 consecutive fits of the shipped SP500 H4 model, correlation
between the chosen threshold and the win rate at it was -0.056, while the
era-to-era spread of that win rate (1.32pp) matched its own binomial SE
(1.25pp) to within 0.07pp. The margin does not rank trades. So the argmax
returned whichever of ~37 bins drew the luckiest sample, and the threshold
teleported 0.42 -> 0.04 -> 0.74 in three eras.
The response at the time was to build a null-of-the-maximum gate, an
effective-sample SE and a parsimony fallback to hold the noise down. All of
that is gone now, because fitting on calibration removes the problem instead
of bounding it: coverage is a ratio against a fixed denominator so it is well
determined at every bin, the target is a measured label rate rather than an
outcome, and nothing is maximised over a noisy curve so there is no best-of-N
to correct for. Net 174 lines out, 62 in.
It deliberately does not chase edge. It cannot - at ~0 measured edge no
operating point has more of it, and pretending otherwise is what produced a
threshold of 0.96 that still passed 60% of bars while the model called a
direction ~10x too often. The edge at the chosen point is still REPORTED,
just no longer what chooses it.
THREE READINGS OF ONE QUANTITY, AND THEY DISAGREE. "How often does a direction
occur" is measured in three places and gives ~7% (the scan's own tally), ~41%
(the era loop's counters, via this function's old coverage floor) and ~50%
(the ensemble gate's OOS base rate). They cannot all be right. Rather than
pick one silently, ScanDirectionalRatePct() and EraDirectionalRatePct() are
now named accessors, the fitter targets the SCAN - that is the tally the
operator reads, and the one "predict the labels as measured during the scan
phase" names - and the threshold line PRINTS BOTH every time it moves, so the
disagreement is on the record instead of buried in a derived floor.
The ensemble gate's own floor is deliberately NOT changed in this commit. If
the scan is right, a calibrated member covering ~7% of bars cannot clear a
12.4% floor and every model would fail the gate by construction; if the gate
is right, the scan tally is wrong. The CALIBRATION field added in 667f2bc
reports the OOS true class rates directly and settles it in one era - that
measurement comes first, and the floor follows it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-21 10:20:31 -04:00
|
|
|
//--- THE TARGET IS THE LABEL RATE. Call a direction as often as a direction actually occurs -
|
2026-08-22 00:25:52 -04:00
|
|
|
//--- nothing else. A whole null-of-the-maximum apparatus was then built to hold that down.
|
feat(calibration): fit the operating point on the label rate instead of on edge
The margin threshold now sits where the model calls a direction as often as a
direction actually occurs. Nothing else.
WHY THE OLD OBJECTIVE HAD TO GO. It maximised `coverage x (precision -
breakEven)`, and this function's own comments were already the case against
it: over 98 consecutive fits of the shipped SP500 H4 model, correlation
between the chosen threshold and the win rate at it was -0.056, while the
era-to-era spread of that win rate (1.32pp) matched its own binomial SE
(1.25pp) to within 0.07pp. The margin does not rank trades. So the argmax
returned whichever of ~37 bins drew the luckiest sample, and the threshold
teleported 0.42 -> 0.04 -> 0.74 in three eras.
The response at the time was to build a null-of-the-maximum gate, an
effective-sample SE and a parsimony fallback to hold the noise down. All of
that is gone now, because fitting on calibration removes the problem instead
of bounding it: coverage is a ratio against a fixed denominator so it is well
determined at every bin, the target is a measured label rate rather than an
outcome, and nothing is maximised over a noisy curve so there is no best-of-N
to correct for. Net 174 lines out, 62 in.
It deliberately does not chase edge. It cannot - at ~0 measured edge no
operating point has more of it, and pretending otherwise is what produced a
threshold of 0.96 that still passed 60% of bars while the model called a
direction ~10x too often. The edge at the chosen point is still REPORTED,
just no longer what chooses it.
THREE READINGS OF ONE QUANTITY, AND THEY DISAGREE. "How often does a direction
occur" is measured in three places and gives ~7% (the scan's own tally), ~41%
(the era loop's counters, via this function's old coverage floor) and ~50%
(the ensemble gate's OOS base rate). They cannot all be right. Rather than
pick one silently, ScanDirectionalRatePct() and EraDirectionalRatePct() are
now named accessors, the fitter targets the SCAN - that is the tally the
operator reads, and the one "predict the labels as measured during the scan
phase" names - and the threshold line PRINTS BOTH every time it moves, so the
disagreement is on the record instead of buried in a derived floor.
The ensemble gate's own floor is deliberately NOT changed in this commit. If
the scan is right, a calibrated member covering ~7% of bars cannot clear a
12.4% floor and every model would fail the gate by construction; if the gate
is right, the scan tally is wrong. The CALIBRATION field added in 667f2bc
reports the OOS true class rates directly and settles it in one era - that
measurement comes first, and the floor follows it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-21 10:20:31 -04:00
|
|
|
double targetPct = ScanDirectionalRatePct();
|
|
|
|
|
double eraPct = EraDirectionalRatePct();
|
|
|
|
|
bool fromScan = (targetPct >= 0.0);
|
|
|
|
|
if(!fromScan)
|
|
|
|
|
targetPct = eraPct; // resumed model with no prebuild - the era tally is the only measurement
|
|
|
|
|
if(targetPct < 0.0)
|
|
|
|
|
return; // nothing measured either way: keep the operating point we have
|
2026-08-22 00:25:52 -04:00
|
|
|
//--- Kept for the report only. The edge at the chosen point is worth SEEING - it is just no
|
|
|
|
|
//--- longer what chooses the point. The HORIZON-AWARE one, because this line is what the
|
|
|
|
|
//--- operator reads to judge a model.
|
feat(breakeven): the break-even every layer scores against prices a trade that always resolves
CostAdjustedBreakEvenPct is risk/(risk+reward) and has no horizon term. It is the win rate a trade
needs when it is CERTAIN to end at one barrier or the other. SimulateTradeOutcome has an explicit
branch for the case where it does not - runs out of horizon, closes at the last bar seen for
whatever P&L that is - so on this label geometry the figure describes a different trade than the
one being replayed.
The gap is measurable and large. Across 21 exit replays today on SP500 H4 the geometric figure read
34.5% while the EA's own R simulation crossed zero between 27.4% (lowest positive) and 28.9%
(highest non-positive). Independent corroboration: the zero-skill reference, computed empirically
over every scored bar as max(winLong,winShort)/bars, reads 25.4% - add cost and it lands on the
same ~28%. The geometric number is the outlier, and every edge printed against it was ~6.5pp too
pessimistic: LSTM's 30.6%-win era reported -4.0pp while its replay returned +0.075 R on the same
trades.
With a timeout share t paying a mean m R apiece, expectancy is w(1+RR) + t(1+m) - 1, so
w* = (1 - t(1+m)) / (1 + RR) = CostAdjustedBreakEvenPct x (1 - t(1+m))
which needs no new geometry - the existing figure already carries 1/(1+RR).
This commit MEASURES ONLY. The replay now separates timeout exits from barrier exits and latches
t and m for the next era to read (the accumulators are zeroed at era start and filled at era end,
so a mid-era reader sees zero trades and would fall back forever). Both break-evens print side by
side on the replay line with t and m beside them, and the threshold line's REPORTED edge - which
selects nothing - switches to the horizon-aware figure so the operator stops reading a wrong sign.
DELIBERATELY NOT CHANGED: LiveMetaGate's veto and the rung selector's BarrierMinReachPct still read
the geometric value. Both are decisions - the second re-derives geometry and therefore relabels -
and t and m have so far only been inferred from a zero-crossing, never seen on a log. One era of
this instrumentation settles that.
The file already contained the argument, one branch away, in the vote-exit comment: a vote exit
produces a CONTINUOUS payoff, not a win or a loss, and that is why an exit-aware gate cannot go on
scoring win-rate against a fixed break-even. A horizon timeout is the same thing, and unlike vote
exits it is on by default.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-21 19:26:25 -04:00
|
|
|
double breakEvenPct = EmpiricalBreakEvenPct();
|
feat(calibration): fit the operating point on the label rate instead of on edge
The margin threshold now sits where the model calls a direction as often as a
direction actually occurs. Nothing else.
WHY THE OLD OBJECTIVE HAD TO GO. It maximised `coverage x (precision -
breakEven)`, and this function's own comments were already the case against
it: over 98 consecutive fits of the shipped SP500 H4 model, correlation
between the chosen threshold and the win rate at it was -0.056, while the
era-to-era spread of that win rate (1.32pp) matched its own binomial SE
(1.25pp) to within 0.07pp. The margin does not rank trades. So the argmax
returned whichever of ~37 bins drew the luckiest sample, and the threshold
teleported 0.42 -> 0.04 -> 0.74 in three eras.
The response at the time was to build a null-of-the-maximum gate, an
effective-sample SE and a parsimony fallback to hold the noise down. All of
that is gone now, because fitting on calibration removes the problem instead
of bounding it: coverage is a ratio against a fixed denominator so it is well
determined at every bin, the target is a measured label rate rather than an
outcome, and nothing is maximised over a noisy curve so there is no best-of-N
to correct for. Net 174 lines out, 62 in.
It deliberately does not chase edge. It cannot - at ~0 measured edge no
operating point has more of it, and pretending otherwise is what produced a
threshold of 0.96 that still passed 60% of bars while the model called a
direction ~10x too often. The edge at the chosen point is still REPORTED,
just no longer what chooses it.
THREE READINGS OF ONE QUANTITY, AND THEY DISAGREE. "How often does a direction
occur" is measured in three places and gives ~7% (the scan's own tally), ~41%
(the era loop's counters, via this function's old coverage floor) and ~50%
(the ensemble gate's OOS base rate). They cannot all be right. Rather than
pick one silently, ScanDirectionalRatePct() and EraDirectionalRatePct() are
now named accessors, the fitter targets the SCAN - that is the tally the
operator reads, and the one "predict the labels as measured during the scan
phase" names - and the threshold line PRINTS BOTH every time it moves, so the
disagreement is on the record instead of buried in a derived floor.
The ensemble gate's own floor is deliberately NOT changed in this commit. If
the scan is right, a calibrated member covering ~7% of bars cannot clear a
12.4% floor and every model would fail the gate by construction; if the gate
is right, the scan tally is wrong. The CALIBRATION field added in 667f2bc
reports the OOS true class rates directly and settles it in one era - that
measurement comes first, and the floor follows it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-21 10:20:31 -04:00
|
|
|
//--- Running "at or above this bin" totals, which is exactly the population a threshold there
|
|
|
|
|
//--- admits. Coverage therefore rises monotonically as the sweep descends, so |coverage - target|
|
|
|
|
|
//--- is V-shaped and the first minimum found is the answer.
|
fix(diagnostics+calibration): the frozen-layer reading was a broken ruler; gate the operating point on a null of the maximum
Two defects behind the "training is highly unstable" report, from 101 eras of
SP500 H4 PAI logs. Neither was the optimizer.
1) LayerLearningReport's dW/W for BN layers divided by the WHOLE packed block.
getWeightsBN concatenates the outgoing dense matrix, gamma/beta, the running
mean/variance, the Adam moments AND BN_OPT_NX - the forward-pass scratch copy
of the normalized input. At era 101 bn1's dense matrix normed 15.1 against a
block norm of 15430.3, of which NX alone was 15429.3: the weights were 0.098%
of their own denominator, a 1022x inflation. NX is also near-constant between
era-end reports (same last forward pass), which pins the numerator down too,
so the layer read "bn1:0.000%" for 101 consecutive eras and was diagnosed as a
frozen first layer. It was the ruler that was broken. The ratio now covers
trainable parameters only (dense matrix + gamma + beta); mean/var/NX/Adam are
excluded. NX is reported separately because it is a health signal in its own
right - bn5 read nx 6.8e6 over 16 neurons, ~1.7e6 per unit against a healthy
~1.0, which is what a near-zero running variance in the denominator looks like.
NO historical dW/W reading on a bn* layer is admissible evidence that a layer
did or did not train. That includes every such claim in this repo's notes.
2) FitDirConfThreshold took a bare argmax of coverage x (precision - breakEven)
over 50 bins. Measured across 98 consecutive fits:
correlation(chosen threshold, win rate at it) = -0.056 over 0.00..0.74
win rate stdev across fits = 1.32pp
binomial SE of that win rate at ~1430 calls = 1.25pp
The correlation is zero - the margin does not rank trades - and the era-to-era
spread IS its own sampling error to within 0.07pp. So the objective was
coverage x (3.4 +/- 1.3) and the argmax over ~37 eligible bins returned
whichever bin drew the luckiest sample. The threshold teleported
0.42 -> 0.04 -> 0.74 in three eras, swinging OOS coverage 0% -> 39%, leaving
the era win rate measured on 1-5 calls and swinging 0% <-> 100%. That is the
entire reported instability.
The argmax is now adopted only if it beats a DETERMINISTIC fallback - the most
selective bin still clearing the coverage floor, chosen from the margin
distribution alone and never from a win rate - by more than a best-of-N
maximum could manage on noise, sqrt(2 ln N) standard errors. Same null-of-the-
maximum correction the deploy gate already applies to model selection.
A plain one-standard-error band was tried first and is NOT sufficient: its
edge is bestScore - bestSE, and with a 2.3pp edge against a 1.25pp SE that
edge is itself +/-50%, so the admitted set would still wander by half its own
width every era. The fallback has to be independent of the noisy quantity.
Simulated on the observed numbers: falls back every era at the current 2.3pp
edge (stable), adopts the argmax once a real edge reaches ~5pp.
NOT COMPILED - user compiles.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-16 23:44:26 -04:00
|
|
|
double binCov[DIR_CONF_THRESHOLD_BINS];
|
|
|
|
|
double binPrec[DIR_CONF_THRESHOLD_BINS];
|
feat: fitted directional confidence threshold - selectivity gets a mechanism
The training loss and the selection metric wanted different things and only
the second one knew it. Logit-adjusted cross-entropy has no term for "how
often should I trade", so the head calls a direction on 87-91% of bars. The
selection metric is precision x coverage credit, saturating at the coverage
floor - above the floor extra calls earn NOTHING and only precision counts.
So selection wanted few good calls, the loss produced many mediocre ones, and
all selection could do was pick the least-bad era out of what it was handed.
Nothing pushed the model toward selectivity.
This gives the decision RULE the policy instead of distorting the loss (which
is estimating class probabilities correctly, and a probability estimate should
not be bent to encode a trading policy - Elkan 2001: estimate, then choose the
operating point separately). AdjustedSignalFromSoftmax now abstains unless the
winning direction's softmax margin over its best rival clears a fitted
threshold. Margin, not the winning probability: the latter moves with overall
calibration rather than with how close the decision actually was.
Fitted on IS, applied to OOS and live. Pass 2 already forward-passes every IS
sample, so the margin histogram is harvested there for free (primary
occurrences only, so the oversampled replay queue cannot skew the operating
point); the fit runs at the end of pass 2, BEFORE pass 3, so the deploy gate
grades the thresholded model on bars the threshold never saw. Fitting on
pass 3's own predictions would be choosing the operating point on the data
being graded - the best-of-N error corrected in five other places here.
Objective: maximise IS directional precision subject to still clearing the
SAME coverage floor the deploy gate uses (base rate x 0.25, re-derived
locally so the two cannot drift apart). Swept top-down in one pass; ties go
to the LOWER threshold, since equal precision for less coverage is strictly
worse. Under DIR_CONF_MIN_FIT_CALLS (200) it runs unthresholded rather than
on a guess.
The threshold is part of the MODEL, not the run: captured with
Net.CaptureWeights(), restored with the weights at both restore sites, and
appended to the .cfg under the same length-guard convention so a deployed
model reloads at the operating point its gate actually cleared. A pre-2026-08-09
.cfg reads 0.0, which is exactly the behaviour it was trained under.
Per-era line now prints "@margin>=X.XX" next to coverage, so a coverage drop
can be attributed to the operating point rather than guessed at.
Both build variants compile 0 errors / 0 warnings.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 15:04:37 -04:00
|
|
|
long runCalls = 0, runHits = 0;
|
feat(calibration): fit the operating point on the label rate instead of on edge
The margin threshold now sits where the model calls a direction as often as a
direction actually occurs. Nothing else.
WHY THE OLD OBJECTIVE HAD TO GO. It maximised `coverage x (precision -
breakEven)`, and this function's own comments were already the case against
it: over 98 consecutive fits of the shipped SP500 H4 model, correlation
between the chosen threshold and the win rate at it was -0.056, while the
era-to-era spread of that win rate (1.32pp) matched its own binomial SE
(1.25pp) to within 0.07pp. The margin does not rank trades. So the argmax
returned whichever of ~37 bins drew the luckiest sample, and the threshold
teleported 0.42 -> 0.04 -> 0.74 in three eras.
The response at the time was to build a null-of-the-maximum gate, an
effective-sample SE and a parsimony fallback to hold the noise down. All of
that is gone now, because fitting on calibration removes the problem instead
of bounding it: coverage is a ratio against a fixed denominator so it is well
determined at every bin, the target is a measured label rate rather than an
outcome, and nothing is maximised over a noisy curve so there is no best-of-N
to correct for. Net 174 lines out, 62 in.
It deliberately does not chase edge. It cannot - at ~0 measured edge no
operating point has more of it, and pretending otherwise is what produced a
threshold of 0.96 that still passed 60% of bars while the model called a
direction ~10x too often. The edge at the chosen point is still REPORTED,
just no longer what chooses it.
THREE READINGS OF ONE QUANTITY, AND THEY DISAGREE. "How often does a direction
occur" is measured in three places and gives ~7% (the scan's own tally), ~41%
(the era loop's counters, via this function's old coverage floor) and ~50%
(the ensemble gate's OOS base rate). They cannot all be right. Rather than
pick one silently, ScanDirectionalRatePct() and EraDirectionalRatePct() are
now named accessors, the fitter targets the SCAN - that is the tally the
operator reads, and the one "predict the labels as measured during the scan
phase" names - and the threshold line PRINTS BOTH every time it moves, so the
disagreement is on the record instead of buried in a derived floor.
The ensemble gate's own floor is deliberately NOT changed in this commit. If
the scan is right, a calibrated member covering ~7% of bars cannot clear a
12.4% floor and every model would fail the gate by construction; if the gate
is right, the scan tally is wrong. The CALIBRATION field added in 667f2bc
reports the OOS true class rates directly and settles it in one era - that
measurement comes first, and the floor follows it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-21 10:20:31 -04:00
|
|
|
double bestGap = DBL_MAX;
|
|
|
|
|
int bestBin = -1;
|
feat: fitted directional confidence threshold - selectivity gets a mechanism
The training loss and the selection metric wanted different things and only
the second one knew it. Logit-adjusted cross-entropy has no term for "how
often should I trade", so the head calls a direction on 87-91% of bars. The
selection metric is precision x coverage credit, saturating at the coverage
floor - above the floor extra calls earn NOTHING and only precision counts.
So selection wanted few good calls, the loss produced many mediocre ones, and
all selection could do was pick the least-bad era out of what it was handed.
Nothing pushed the model toward selectivity.
This gives the decision RULE the policy instead of distorting the loss (which
is estimating class probabilities correctly, and a probability estimate should
not be bent to encode a trading policy - Elkan 2001: estimate, then choose the
operating point separately). AdjustedSignalFromSoftmax now abstains unless the
winning direction's softmax margin over its best rival clears a fitted
threshold. Margin, not the winning probability: the latter moves with overall
calibration rather than with how close the decision actually was.
Fitted on IS, applied to OOS and live. Pass 2 already forward-passes every IS
sample, so the margin histogram is harvested there for free (primary
occurrences only, so the oversampled replay queue cannot skew the operating
point); the fit runs at the end of pass 2, BEFORE pass 3, so the deploy gate
grades the thresholded model on bars the threshold never saw. Fitting on
pass 3's own predictions would be choosing the operating point on the data
being graded - the best-of-N error corrected in five other places here.
Objective: maximise IS directional precision subject to still clearing the
SAME coverage floor the deploy gate uses (base rate x 0.25, re-derived
locally so the two cannot drift apart). Swept top-down in one pass; ties go
to the LOWER threshold, since equal precision for less coverage is strictly
worse. Under DIR_CONF_MIN_FIT_CALLS (200) it runs unthresholded rather than
on a guess.
The threshold is part of the MODEL, not the run: captured with
Net.CaptureWeights(), restored with the weights at both restore sites, and
appended to the .cfg under the same length-guard convention so a deployed
model reloads at the operating point its gate actually cleared. A pre-2026-08-09
.cfg reads 0.0, which is exactly the behaviour it was trained under.
Per-era line now prints "@margin>=X.XX" next to coverage, so a coverage drop
can be attributed to the operating point rather than guessed at.
Both build variants compile 0 errors / 0 warnings.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 15:04:37 -04:00
|
|
|
for(int b = DIR_CONF_THRESHOLD_BINS - 1; b >= 0; b--)
|
|
|
|
|
{
|
fix(diagnostics+calibration): the frozen-layer reading was a broken ruler; gate the operating point on a null of the maximum
Two defects behind the "training is highly unstable" report, from 101 eras of
SP500 H4 PAI logs. Neither was the optimizer.
1) LayerLearningReport's dW/W for BN layers divided by the WHOLE packed block.
getWeightsBN concatenates the outgoing dense matrix, gamma/beta, the running
mean/variance, the Adam moments AND BN_OPT_NX - the forward-pass scratch copy
of the normalized input. At era 101 bn1's dense matrix normed 15.1 against a
block norm of 15430.3, of which NX alone was 15429.3: the weights were 0.098%
of their own denominator, a 1022x inflation. NX is also near-constant between
era-end reports (same last forward pass), which pins the numerator down too,
so the layer read "bn1:0.000%" for 101 consecutive eras and was diagnosed as a
frozen first layer. It was the ruler that was broken. The ratio now covers
trainable parameters only (dense matrix + gamma + beta); mean/var/NX/Adam are
excluded. NX is reported separately because it is a health signal in its own
right - bn5 read nx 6.8e6 over 16 neurons, ~1.7e6 per unit against a healthy
~1.0, which is what a near-zero running variance in the denominator looks like.
NO historical dW/W reading on a bn* layer is admissible evidence that a layer
did or did not train. That includes every such claim in this repo's notes.
2) FitDirConfThreshold took a bare argmax of coverage x (precision - breakEven)
over 50 bins. Measured across 98 consecutive fits:
correlation(chosen threshold, win rate at it) = -0.056 over 0.00..0.74
win rate stdev across fits = 1.32pp
binomial SE of that win rate at ~1430 calls = 1.25pp
The correlation is zero - the margin does not rank trades - and the era-to-era
spread IS its own sampling error to within 0.07pp. So the objective was
coverage x (3.4 +/- 1.3) and the argmax over ~37 eligible bins returned
whichever bin drew the luckiest sample. The threshold teleported
0.42 -> 0.04 -> 0.74 in three eras, swinging OOS coverage 0% -> 39%, leaving
the era win rate measured on 1-5 calls and swinging 0% <-> 100%. That is the
entire reported instability.
The argmax is now adopted only if it beats a DETERMINISTIC fallback - the most
selective bin still clearing the coverage floor, chosen from the margin
distribution alone and never from a win rate - by more than a best-of-N
maximum could manage on noise, sqrt(2 ln N) standard errors. Same null-of-the-
maximum correction the deploy gate already applies to model selection.
A plain one-standard-error band was tried first and is NOT sufficient: its
edge is bestScore - bestSE, and with a 2.3pp edge against a 1.25pp SE that
edge is itself +/-50%, so the admitted set would still wander by half its own
width every era. The fallback has to be independent of the noisy quantity.
Simulated on the observed numbers: falls back every era at the current 2.3pp
edge (stable), adopts the argmax once a real edge reaches ~5pp.
NOT COMPILED - user compiles.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-16 23:44:26 -04:00
|
|
|
binCov[b] = 0.0;
|
|
|
|
|
binPrec[b] = 0.0;
|
feat: fitted directional confidence threshold - selectivity gets a mechanism
The training loss and the selection metric wanted different things and only
the second one knew it. Logit-adjusted cross-entropy has no term for "how
often should I trade", so the head calls a direction on 87-91% of bars. The
selection metric is precision x coverage credit, saturating at the coverage
floor - above the floor extra calls earn NOTHING and only precision counts.
So selection wanted few good calls, the loss produced many mediocre ones, and
all selection could do was pick the least-bad era out of what it was handed.
Nothing pushed the model toward selectivity.
This gives the decision RULE the policy instead of distorting the loss (which
is estimating class probabilities correctly, and a probability estimate should
not be bent to encode a trading policy - Elkan 2001: estimate, then choose the
operating point separately). AdjustedSignalFromSoftmax now abstains unless the
winning direction's softmax margin over its best rival clears a fitted
threshold. Margin, not the winning probability: the latter moves with overall
calibration rather than with how close the decision actually was.
Fitted on IS, applied to OOS and live. Pass 2 already forward-passes every IS
sample, so the margin histogram is harvested there for free (primary
occurrences only, so the oversampled replay queue cannot skew the operating
point); the fit runs at the end of pass 2, BEFORE pass 3, so the deploy gate
grades the thresholded model on bars the threshold never saw. Fitting on
pass 3's own predictions would be choosing the operating point on the data
being graded - the best-of-N error corrected in five other places here.
Objective: maximise IS directional precision subject to still clearing the
SAME coverage floor the deploy gate uses (base rate x 0.25, re-derived
locally so the two cannot drift apart). Swept top-down in one pass; ties go
to the LOWER threshold, since equal precision for less coverage is strictly
worse. Under DIR_CONF_MIN_FIT_CALLS (200) it runs unthresholded rather than
on a guess.
The threshold is part of the MODEL, not the run: captured with
Net.CaptureWeights(), restored with the weights at both restore sites, and
appended to the .cfg under the same length-guard convention so a deployed
model reloads at the operating point its gate actually cleared. A pre-2026-08-09
.cfg reads 0.0, which is exactly the behaviour it was trained under.
Per-era line now prints "@margin>=X.XX" next to coverage, so a coverage drop
can be attributed to the operating point rather than guessed at.
Both build variants compile 0 errors / 0 warnings.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 15:04:37 -04:00
|
|
|
runCalls += m_dirConfBinCalls[b];
|
|
|
|
|
runHits += m_dirConfBinHits[b];
|
|
|
|
|
if(runCalls <= 0)
|
|
|
|
|
continue;
|
feat(calibration): fit the operating point on the label rate instead of on edge
The margin threshold now sits where the model calls a direction as often as a
direction actually occurs. Nothing else.
WHY THE OLD OBJECTIVE HAD TO GO. It maximised `coverage x (precision -
breakEven)`, and this function's own comments were already the case against
it: over 98 consecutive fits of the shipped SP500 H4 model, correlation
between the chosen threshold and the win rate at it was -0.056, while the
era-to-era spread of that win rate (1.32pp) matched its own binomial SE
(1.25pp) to within 0.07pp. The margin does not rank trades. So the argmax
returned whichever of ~37 bins drew the luckiest sample, and the threshold
teleported 0.42 -> 0.04 -> 0.74 in three eras.
The response at the time was to build a null-of-the-maximum gate, an
effective-sample SE and a parsimony fallback to hold the noise down. All of
that is gone now, because fitting on calibration removes the problem instead
of bounding it: coverage is a ratio against a fixed denominator so it is well
determined at every bin, the target is a measured label rate rather than an
outcome, and nothing is maximised over a noisy curve so there is no best-of-N
to correct for. Net 174 lines out, 62 in.
It deliberately does not chase edge. It cannot - at ~0 measured edge no
operating point has more of it, and pretending otherwise is what produced a
threshold of 0.96 that still passed 60% of bars while the model called a
direction ~10x too often. The edge at the chosen point is still REPORTED,
just no longer what chooses it.
THREE READINGS OF ONE QUANTITY, AND THEY DISAGREE. "How often does a direction
occur" is measured in three places and gives ~7% (the scan's own tally), ~41%
(the era loop's counters, via this function's old coverage floor) and ~50%
(the ensemble gate's OOS base rate). They cannot all be right. Rather than
pick one silently, ScanDirectionalRatePct() and EraDirectionalRatePct() are
now named accessors, the fitter targets the SCAN - that is the tally the
operator reads, and the one "predict the labels as measured during the scan
phase" names - and the threshold line PRINTS BOTH every time it moves, so the
disagreement is on the record instead of buried in a derived floor.
The ensemble gate's own floor is deliberately NOT changed in this commit. If
the scan is right, a calibrated member covering ~7% of bars cannot clear a
12.4% floor and every model would fail the gate by construction; if the gate
is right, the scan tally is wrong. The CALIBRATION field added in 667f2bc
reports the OOS true class rates directly and settles it in one era - that
measurement comes first, and the floor follows it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-21 10:20:31 -04:00
|
|
|
binCov[b] = 100.0 * (double)runCalls / m_dirConfPrimaryBars;
|
|
|
|
|
binPrec[b] = 100.0 * (double)runHits / runCalls;
|
|
|
|
|
double gap = MathAbs(binCov[b] - targetPct);
|
|
|
|
|
//--- Strict <, so a tie keeps the MORE selective bin - the sweep reaches it first. Same
|
|
|
|
|
//--- tie-break direction the expectancy version used, and for the same reason.
|
|
|
|
|
if(gap < bestGap)
|
feat: fitted directional confidence threshold - selectivity gets a mechanism
The training loss and the selection metric wanted different things and only
the second one knew it. Logit-adjusted cross-entropy has no term for "how
often should I trade", so the head calls a direction on 87-91% of bars. The
selection metric is precision x coverage credit, saturating at the coverage
floor - above the floor extra calls earn NOTHING and only precision counts.
So selection wanted few good calls, the loss produced many mediocre ones, and
all selection could do was pick the least-bad era out of what it was handed.
Nothing pushed the model toward selectivity.
This gives the decision RULE the policy instead of distorting the loss (which
is estimating class probabilities correctly, and a probability estimate should
not be bent to encode a trading policy - Elkan 2001: estimate, then choose the
operating point separately). AdjustedSignalFromSoftmax now abstains unless the
winning direction's softmax margin over its best rival clears a fitted
threshold. Margin, not the winning probability: the latter moves with overall
calibration rather than with how close the decision actually was.
Fitted on IS, applied to OOS and live. Pass 2 already forward-passes every IS
sample, so the margin histogram is harvested there for free (primary
occurrences only, so the oversampled replay queue cannot skew the operating
point); the fit runs at the end of pass 2, BEFORE pass 3, so the deploy gate
grades the thresholded model on bars the threshold never saw. Fitting on
pass 3's own predictions would be choosing the operating point on the data
being graded - the best-of-N error corrected in five other places here.
Objective: maximise IS directional precision subject to still clearing the
SAME coverage floor the deploy gate uses (base rate x 0.25, re-derived
locally so the two cannot drift apart). Swept top-down in one pass; ties go
to the LOWER threshold, since equal precision for less coverage is strictly
worse. Under DIR_CONF_MIN_FIT_CALLS (200) it runs unthresholded rather than
on a guess.
The threshold is part of the MODEL, not the run: captured with
Net.CaptureWeights(), restored with the weights at both restore sites, and
appended to the .cfg under the same length-guard convention so a deployed
model reloads at the operating point its gate actually cleared. A pre-2026-08-09
.cfg reads 0.0, which is exactly the behaviour it was trained under.
Per-era line now prints "@margin>=X.XX" next to coverage, so a coverage drop
can be attributed to the operating point rather than guessed at.
Both build variants compile 0 errors / 0 warnings.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 15:04:37 -04:00
|
|
|
{
|
feat(calibration): fit the operating point on the label rate instead of on edge
The margin threshold now sits where the model calls a direction as often as a
direction actually occurs. Nothing else.
WHY THE OLD OBJECTIVE HAD TO GO. It maximised `coverage x (precision -
breakEven)`, and this function's own comments were already the case against
it: over 98 consecutive fits of the shipped SP500 H4 model, correlation
between the chosen threshold and the win rate at it was -0.056, while the
era-to-era spread of that win rate (1.32pp) matched its own binomial SE
(1.25pp) to within 0.07pp. The margin does not rank trades. So the argmax
returned whichever of ~37 bins drew the luckiest sample, and the threshold
teleported 0.42 -> 0.04 -> 0.74 in three eras.
The response at the time was to build a null-of-the-maximum gate, an
effective-sample SE and a parsimony fallback to hold the noise down. All of
that is gone now, because fitting on calibration removes the problem instead
of bounding it: coverage is a ratio against a fixed denominator so it is well
determined at every bin, the target is a measured label rate rather than an
outcome, and nothing is maximised over a noisy curve so there is no best-of-N
to correct for. Net 174 lines out, 62 in.
It deliberately does not chase edge. It cannot - at ~0 measured edge no
operating point has more of it, and pretending otherwise is what produced a
threshold of 0.96 that still passed 60% of bars while the model called a
direction ~10x too often. The edge at the chosen point is still REPORTED,
just no longer what chooses it.
THREE READINGS OF ONE QUANTITY, AND THEY DISAGREE. "How often does a direction
occur" is measured in three places and gives ~7% (the scan's own tally), ~41%
(the era loop's counters, via this function's old coverage floor) and ~50%
(the ensemble gate's OOS base rate). They cannot all be right. Rather than
pick one silently, ScanDirectionalRatePct() and EraDirectionalRatePct() are
now named accessors, the fitter targets the SCAN - that is the tally the
operator reads, and the one "predict the labels as measured during the scan
phase" names - and the threshold line PRINTS BOTH every time it moves, so the
disagreement is on the record instead of buried in a derived floor.
The ensemble gate's own floor is deliberately NOT changed in this commit. If
the scan is right, a calibrated member covering ~7% of bars cannot clear a
12.4% floor and every model would fail the gate by construction; if the gate
is right, the scan tally is wrong. The CALIBRATION field added in 667f2bc
reports the OOS true class rates directly and settles it in one era - that
measurement comes first, and the floor follows it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-21 10:20:31 -04:00
|
|
|
bestGap = gap;
|
fix(diagnostics+calibration): the frozen-layer reading was a broken ruler; gate the operating point on a null of the maximum
Two defects behind the "training is highly unstable" report, from 101 eras of
SP500 H4 PAI logs. Neither was the optimizer.
1) LayerLearningReport's dW/W for BN layers divided by the WHOLE packed block.
getWeightsBN concatenates the outgoing dense matrix, gamma/beta, the running
mean/variance, the Adam moments AND BN_OPT_NX - the forward-pass scratch copy
of the normalized input. At era 101 bn1's dense matrix normed 15.1 against a
block norm of 15430.3, of which NX alone was 15429.3: the weights were 0.098%
of their own denominator, a 1022x inflation. NX is also near-constant between
era-end reports (same last forward pass), which pins the numerator down too,
so the layer read "bn1:0.000%" for 101 consecutive eras and was diagnosed as a
frozen first layer. It was the ruler that was broken. The ratio now covers
trainable parameters only (dense matrix + gamma + beta); mean/var/NX/Adam are
excluded. NX is reported separately because it is a health signal in its own
right - bn5 read nx 6.8e6 over 16 neurons, ~1.7e6 per unit against a healthy
~1.0, which is what a near-zero running variance in the denominator looks like.
NO historical dW/W reading on a bn* layer is admissible evidence that a layer
did or did not train. That includes every such claim in this repo's notes.
2) FitDirConfThreshold took a bare argmax of coverage x (precision - breakEven)
over 50 bins. Measured across 98 consecutive fits:
correlation(chosen threshold, win rate at it) = -0.056 over 0.00..0.74
win rate stdev across fits = 1.32pp
binomial SE of that win rate at ~1430 calls = 1.25pp
The correlation is zero - the margin does not rank trades - and the era-to-era
spread IS its own sampling error to within 0.07pp. So the objective was
coverage x (3.4 +/- 1.3) and the argmax over ~37 eligible bins returned
whichever bin drew the luckiest sample. The threshold teleported
0.42 -> 0.04 -> 0.74 in three eras, swinging OOS coverage 0% -> 39%, leaving
the era win rate measured on 1-5 calls and swinging 0% <-> 100%. That is the
entire reported instability.
The argmax is now adopted only if it beats a DETERMINISTIC fallback - the most
selective bin still clearing the coverage floor, chosen from the margin
distribution alone and never from a win rate - by more than a best-of-N
maximum could manage on noise, sqrt(2 ln N) standard errors. Same null-of-the-
maximum correction the deploy gate already applies to model selection.
A plain one-standard-error band was tried first and is NOT sufficient: its
edge is bestScore - bestSE, and with a 2.3pp edge against a 1.25pp SE that
edge is itself +/-50%, so the admitted set would still wander by half its own
width every era. The fallback has to be independent of the noisy quantity.
Simulated on the observed numbers: falls back every era at the current 2.3pp
edge (stable), adopts the argmax once a real edge reaches ~5pp.
NOT COMPILED - user compiles.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-16 23:44:26 -04:00
|
|
|
bestBin = b;
|
feat: fitted directional confidence threshold - selectivity gets a mechanism
The training loss and the selection metric wanted different things and only
the second one knew it. Logit-adjusted cross-entropy has no term for "how
often should I trade", so the head calls a direction on 87-91% of bars. The
selection metric is precision x coverage credit, saturating at the coverage
floor - above the floor extra calls earn NOTHING and only precision counts.
So selection wanted few good calls, the loss produced many mediocre ones, and
all selection could do was pick the least-bad era out of what it was handed.
Nothing pushed the model toward selectivity.
This gives the decision RULE the policy instead of distorting the loss (which
is estimating class probabilities correctly, and a probability estimate should
not be bent to encode a trading policy - Elkan 2001: estimate, then choose the
operating point separately). AdjustedSignalFromSoftmax now abstains unless the
winning direction's softmax margin over its best rival clears a fitted
threshold. Margin, not the winning probability: the latter moves with overall
calibration rather than with how close the decision actually was.
Fitted on IS, applied to OOS and live. Pass 2 already forward-passes every IS
sample, so the margin histogram is harvested there for free (primary
occurrences only, so the oversampled replay queue cannot skew the operating
point); the fit runs at the end of pass 2, BEFORE pass 3, so the deploy gate
grades the thresholded model on bars the threshold never saw. Fitting on
pass 3's own predictions would be choosing the operating point on the data
being graded - the best-of-N error corrected in five other places here.
Objective: maximise IS directional precision subject to still clearing the
SAME coverage floor the deploy gate uses (base rate x 0.25, re-derived
locally so the two cannot drift apart). Swept top-down in one pass; ties go
to the LOWER threshold, since equal precision for less coverage is strictly
worse. Under DIR_CONF_MIN_FIT_CALLS (200) it runs unthresholded rather than
on a guess.
The threshold is part of the MODEL, not the run: captured with
Net.CaptureWeights(), restored with the weights at both restore sites, and
appended to the .cfg under the same length-guard convention so a deployed
model reloads at the operating point its gate actually cleared. A pre-2026-08-09
.cfg reads 0.0, which is exactly the behaviour it was trained under.
Per-era line now prints "@margin>=X.XX" next to coverage, so a coverage drop
can be attributed to the operating point rather than guessed at.
Both build variants compile 0 errors / 0 warnings.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 15:04:37 -04:00
|
|
|
}
|
|
|
|
|
}
|
fix(diagnostics+calibration): the frozen-layer reading was a broken ruler; gate the operating point on a null of the maximum
Two defects behind the "training is highly unstable" report, from 101 eras of
SP500 H4 PAI logs. Neither was the optimizer.
1) LayerLearningReport's dW/W for BN layers divided by the WHOLE packed block.
getWeightsBN concatenates the outgoing dense matrix, gamma/beta, the running
mean/variance, the Adam moments AND BN_OPT_NX - the forward-pass scratch copy
of the normalized input. At era 101 bn1's dense matrix normed 15.1 against a
block norm of 15430.3, of which NX alone was 15429.3: the weights were 0.098%
of their own denominator, a 1022x inflation. NX is also near-constant between
era-end reports (same last forward pass), which pins the numerator down too,
so the layer read "bn1:0.000%" for 101 consecutive eras and was diagnosed as a
frozen first layer. It was the ruler that was broken. The ratio now covers
trainable parameters only (dense matrix + gamma + beta); mean/var/NX/Adam are
excluded. NX is reported separately because it is a health signal in its own
right - bn5 read nx 6.8e6 over 16 neurons, ~1.7e6 per unit against a healthy
~1.0, which is what a near-zero running variance in the denominator looks like.
NO historical dW/W reading on a bn* layer is admissible evidence that a layer
did or did not train. That includes every such claim in this repo's notes.
2) FitDirConfThreshold took a bare argmax of coverage x (precision - breakEven)
over 50 bins. Measured across 98 consecutive fits:
correlation(chosen threshold, win rate at it) = -0.056 over 0.00..0.74
win rate stdev across fits = 1.32pp
binomial SE of that win rate at ~1430 calls = 1.25pp
The correlation is zero - the margin does not rank trades - and the era-to-era
spread IS its own sampling error to within 0.07pp. So the objective was
coverage x (3.4 +/- 1.3) and the argmax over ~37 eligible bins returned
whichever bin drew the luckiest sample. The threshold teleported
0.42 -> 0.04 -> 0.74 in three eras, swinging OOS coverage 0% -> 39%, leaving
the era win rate measured on 1-5 calls and swinging 0% <-> 100%. That is the
entire reported instability.
The argmax is now adopted only if it beats a DETERMINISTIC fallback - the most
selective bin still clearing the coverage floor, chosen from the margin
distribution alone and never from a win rate - by more than a best-of-N
maximum could manage on noise, sqrt(2 ln N) standard errors. Same null-of-the-
maximum correction the deploy gate already applies to model selection.
A plain one-standard-error band was tried first and is NOT sufficient: its
edge is bestScore - bestSE, and with a 2.3pp edge against a 1.25pp SE that
edge is itself +/-50%, so the admitted set would still wander by half its own
width every era. The fallback has to be independent of the noisy quantity.
Simulated on the observed numbers: falls back every era at the current 2.3pp
edge (stable), adopts the argmax once a real edge reaches ~5pp.
NOT COMPILED - user compiles.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-16 23:44:26 -04:00
|
|
|
if(bestBin < 0)
|
feat: fitted directional confidence threshold - selectivity gets a mechanism
The training loss and the selection metric wanted different things and only
the second one knew it. Logit-adjusted cross-entropy has no term for "how
often should I trade", so the head calls a direction on 87-91% of bars. The
selection metric is precision x coverage credit, saturating at the coverage
floor - above the floor extra calls earn NOTHING and only precision counts.
So selection wanted few good calls, the loss produced many mediocre ones, and
all selection could do was pick the least-bad era out of what it was handed.
Nothing pushed the model toward selectivity.
This gives the decision RULE the policy instead of distorting the loss (which
is estimating class probabilities correctly, and a probability estimate should
not be bent to encode a trading policy - Elkan 2001: estimate, then choose the
operating point separately). AdjustedSignalFromSoftmax now abstains unless the
winning direction's softmax margin over its best rival clears a fitted
threshold. Margin, not the winning probability: the latter moves with overall
calibration rather than with how close the decision actually was.
Fitted on IS, applied to OOS and live. Pass 2 already forward-passes every IS
sample, so the margin histogram is harvested there for free (primary
occurrences only, so the oversampled replay queue cannot skew the operating
point); the fit runs at the end of pass 2, BEFORE pass 3, so the deploy gate
grades the thresholded model on bars the threshold never saw. Fitting on
pass 3's own predictions would be choosing the operating point on the data
being graded - the best-of-N error corrected in five other places here.
Objective: maximise IS directional precision subject to still clearing the
SAME coverage floor the deploy gate uses (base rate x 0.25, re-derived
locally so the two cannot drift apart). Swept top-down in one pass; ties go
to the LOWER threshold, since equal precision for less coverage is strictly
worse. Under DIR_CONF_MIN_FIT_CALLS (200) it runs unthresholded rather than
on a guess.
The threshold is part of the MODEL, not the run: captured with
Net.CaptureWeights(), restored with the weights at both restore sites, and
appended to the .cfg under the same length-guard convention so a deployed
model reloads at the operating point its gate actually cleared. A pre-2026-08-09
.cfg reads 0.0, which is exactly the behaviour it was trained under.
Per-era line now prints "@margin>=X.XX" next to coverage, so a coverage drop
can be attributed to the operating point rather than guessed at.
Both build variants compile 0 errors / 0 warnings.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 15:04:37 -04:00
|
|
|
return;
|
|
|
|
|
double prevThresh = m_dirConfThreshold;
|
feat(calibration): fit the operating point on the label rate instead of on edge
The margin threshold now sits where the model calls a direction as often as a
direction actually occurs. Nothing else.
WHY THE OLD OBJECTIVE HAD TO GO. It maximised `coverage x (precision -
breakEven)`, and this function's own comments were already the case against
it: over 98 consecutive fits of the shipped SP500 H4 model, correlation
between the chosen threshold and the win rate at it was -0.056, while the
era-to-era spread of that win rate (1.32pp) matched its own binomial SE
(1.25pp) to within 0.07pp. The margin does not rank trades. So the argmax
returned whichever of ~37 bins drew the luckiest sample, and the threshold
teleported 0.42 -> 0.04 -> 0.74 in three eras.
The response at the time was to build a null-of-the-maximum gate, an
effective-sample SE and a parsimony fallback to hold the noise down. All of
that is gone now, because fitting on calibration removes the problem instead
of bounding it: coverage is a ratio against a fixed denominator so it is well
determined at every bin, the target is a measured label rate rather than an
outcome, and nothing is maximised over a noisy curve so there is no best-of-N
to correct for. Net 174 lines out, 62 in.
It deliberately does not chase edge. It cannot - at ~0 measured edge no
operating point has more of it, and pretending otherwise is what produced a
threshold of 0.96 that still passed 60% of bars while the model called a
direction ~10x too often. The edge at the chosen point is still REPORTED,
just no longer what chooses it.
THREE READINGS OF ONE QUANTITY, AND THEY DISAGREE. "How often does a direction
occur" is measured in three places and gives ~7% (the scan's own tally), ~41%
(the era loop's counters, via this function's old coverage floor) and ~50%
(the ensemble gate's OOS base rate). They cannot all be right. Rather than
pick one silently, ScanDirectionalRatePct() and EraDirectionalRatePct() are
now named accessors, the fitter targets the SCAN - that is the tally the
operator reads, and the one "predict the labels as measured during the scan
phase" names - and the threshold line PRINTS BOTH every time it moves, so the
disagreement is on the record instead of buried in a derived floor.
The ensemble gate's own floor is deliberately NOT changed in this commit. If
the scan is right, a calibrated member covering ~7% of bars cannot clear a
12.4% floor and every model would fail the gate by construction; if the gate
is right, the scan tally is wrong. The CALIBRATION field added in 667f2bc
reports the OOS true class rates directly and settles it in one era - that
measurement comes first, and the floor follows it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-21 10:20:31 -04:00
|
|
|
m_dirConfThreshold = (double)bestBin / DIR_CONF_THRESHOLD_BINS;
|
|
|
|
|
//--- Logged when it moves a bin AND the era cadence is due. The threshold updates every era
|
|
|
|
|
//--- regardless; only the announcement is throttled.
|
feat(logs): throttle the settled per-era diagnostics - measured 22MB/9.5h of confirmed-working systems
Measured from the journal (2026-08-19): the era deep-dive line (~2KB) plus
the excursion verdict, tier re-rank, calibration move, barrier hold and
selection-regressed note each printed EVERY era for EVERY member - ~940
eras/member/day - long after the systems they watch were confirmed
working. Yesterday's file was 1.3GB (70% of it the news-filter calendar
spam the sweep fix already removed).
VerboseMode returns as an INPUT (demoted 2026-08-01 for the marketplace;
that track is dead since the 2026-08-16 pivot) and gains a second job:
false throttles each settled per-era print to eras 0-3 plus every
TRAIN_LOG_EVERY_ERAS-th (25 ~= one deep-dive per ~15min per member);
true restores the per-era firehose, flippable live.
Never throttled: anything that marks a CHANGE - new bests, restores +
eta decays, plateau stage transitions, deploy approvals, warnings,
errors, the label-cache/adoption one-shots, and the combined-vote gate
line (the active system's primary telemetry, still every era).
Semantic fixes over blanket gating:
- barrier hold now ARMS silently and prints only when the hold outlasts
the 2-min report interval - a brief hold every era is the design, the
long hold is the watchdog case the line exists for;
- the ensemble deploy REFUSAL prints immediately when its reason
changes (that is a finding), on cadence when unchanged;
- the filtered-view census prints when its RESULT moves (drawn count,
or strongest vote by >=2pp) and at least every 10th sweep.
NOT COMPILED - user compiles in MetaEditor.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-19 09:36:15 -04:00
|
|
|
if(TrainLogDue() && MathAbs(m_dirConfThreshold - prevThresh) >= 1.0 / DIR_CONF_THRESHOLD_BINS)
|
feat(calibration): fit the operating point on the label rate instead of on edge
The margin threshold now sits where the model calls a direction as often as a
direction actually occurs. Nothing else.
WHY THE OLD OBJECTIVE HAD TO GO. It maximised `coverage x (precision -
breakEven)`, and this function's own comments were already the case against
it: over 98 consecutive fits of the shipped SP500 H4 model, correlation
between the chosen threshold and the win rate at it was -0.056, while the
era-to-era spread of that win rate (1.32pp) matched its own binomial SE
(1.25pp) to within 0.07pp. The margin does not rank trades. So the argmax
returned whichever of ~37 bins drew the luckiest sample, and the threshold
teleported 0.42 -> 0.04 -> 0.74 in three eras.
The response at the time was to build a null-of-the-maximum gate, an
effective-sample SE and a parsimony fallback to hold the noise down. All of
that is gone now, because fitting on calibration removes the problem instead
of bounding it: coverage is a ratio against a fixed denominator so it is well
determined at every bin, the target is a measured label rate rather than an
outcome, and nothing is maximised over a noisy curve so there is no best-of-N
to correct for. Net 174 lines out, 62 in.
It deliberately does not chase edge. It cannot - at ~0 measured edge no
operating point has more of it, and pretending otherwise is what produced a
threshold of 0.96 that still passed 60% of bars while the model called a
direction ~10x too often. The edge at the chosen point is still REPORTED,
just no longer what chooses it.
THREE READINGS OF ONE QUANTITY, AND THEY DISAGREE. "How often does a direction
occur" is measured in three places and gives ~7% (the scan's own tally), ~41%
(the era loop's counters, via this function's old coverage floor) and ~50%
(the ensemble gate's OOS base rate). They cannot all be right. Rather than
pick one silently, ScanDirectionalRatePct() and EraDirectionalRatePct() are
now named accessors, the fitter targets the SCAN - that is the tally the
operator reads, and the one "predict the labels as measured during the scan
phase" names - and the threshold line PRINTS BOTH every time it moves, so the
disagreement is on the record instead of buried in a derived floor.
The ensemble gate's own floor is deliberately NOT changed in this commit. If
the scan is right, a calibrated member covering ~7% of bars cannot clear a
12.4% floor and every model would fail the gate by construction; if the gate
is right, the scan tally is wrong. The CALIBRATION field added in 667f2bc
reports the OOS true class rates directly and settles it in one era - that
measurement comes first, and the floor follows it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-21 10:20:31 -04:00
|
|
|
Print(ID + StringFormat(": directional confidence threshold %.2f -> %.2f, fitted on CALIBRATION"
|
|
|
|
|
" over %d held-out bars: calls a direction on %.1f%% of them against a"
|
|
|
|
|
" %s-measured label rate of %.1f%% (miss %.1fpp - the closest of %d bins)."
|
|
|
|
|
" Below this winner-vs-rival margin the model abstains."
|
feat(breakeven): the break-even every layer scores against prices a trade that always resolves
CostAdjustedBreakEvenPct is risk/(risk+reward) and has no horizon term. It is the win rate a trade
needs when it is CERTAIN to end at one barrier or the other. SimulateTradeOutcome has an explicit
branch for the case where it does not - runs out of horizon, closes at the last bar seen for
whatever P&L that is - so on this label geometry the figure describes a different trade than the
one being replayed.
The gap is measurable and large. Across 21 exit replays today on SP500 H4 the geometric figure read
34.5% while the EA's own R simulation crossed zero between 27.4% (lowest positive) and 28.9%
(highest non-positive). Independent corroboration: the zero-skill reference, computed empirically
over every scored bar as max(winLong,winShort)/bars, reads 25.4% - add cost and it lands on the
same ~28%. The geometric number is the outlier, and every edge printed against it was ~6.5pp too
pessimistic: LSTM's 30.6%-win era reported -4.0pp while its replay returned +0.075 R on the same
trades.
With a timeout share t paying a mean m R apiece, expectancy is w(1+RR) + t(1+m) - 1, so
w* = (1 - t(1+m)) / (1 + RR) = CostAdjustedBreakEvenPct x (1 - t(1+m))
which needs no new geometry - the existing figure already carries 1/(1+RR).
This commit MEASURES ONLY. The replay now separates timeout exits from barrier exits and latches
t and m for the next era to read (the accumulators are zeroed at era start and filled at era end,
so a mid-era reader sees zero trades and would fall back forever). Both break-evens print side by
side on the replay line with t and m beside them, and the threshold line's REPORTED edge - which
selects nothing - switches to the horizon-aware figure so the operator stops reading a wrong sign.
DELIBERATELY NOT CHANGED: LiveMetaGate's veto and the rung selector's BarrierMinReachPct still read
the geometric value. Both are decisions - the second re-derives geometry and therefore relabels -
and t and m have so far only been inferred from a zero-crossing, never seen on a log. One era of
this instrumentation settles that.
The file already contained the argument, one branch away, in the vote-exit comment: a vote exit
produces a CONTINUOUS payoff, not a win or a loss, and that is why an exit-aware gate cannot go on
scoring win-rate against a fixed break-even. A horizon timeout is the same thing, and unlike vote
exits it is on by default.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-21 19:26:25 -04:00
|
|
|
" | Win rate at this point %.1f%% vs %.1f%% horizon-aware break-even"
|
|
|
|
|
" (edge %.1fpp; geometric %.1f%%) -"
|
feat(calibration): fit the operating point on the label rate instead of on edge
The margin threshold now sits where the model calls a direction as often as a
direction actually occurs. Nothing else.
WHY THE OLD OBJECTIVE HAD TO GO. It maximised `coverage x (precision -
breakEven)`, and this function's own comments were already the case against
it: over 98 consecutive fits of the shipped SP500 H4 model, correlation
between the chosen threshold and the win rate at it was -0.056, while the
era-to-era spread of that win rate (1.32pp) matched its own binomial SE
(1.25pp) to within 0.07pp. The margin does not rank trades. So the argmax
returned whichever of ~37 bins drew the luckiest sample, and the threshold
teleported 0.42 -> 0.04 -> 0.74 in three eras.
The response at the time was to build a null-of-the-maximum gate, an
effective-sample SE and a parsimony fallback to hold the noise down. All of
that is gone now, because fitting on calibration removes the problem instead
of bounding it: coverage is a ratio against a fixed denominator so it is well
determined at every bin, the target is a measured label rate rather than an
outcome, and nothing is maximised over a noisy curve so there is no best-of-N
to correct for. Net 174 lines out, 62 in.
It deliberately does not chase edge. It cannot - at ~0 measured edge no
operating point has more of it, and pretending otherwise is what produced a
threshold of 0.96 that still passed 60% of bars while the model called a
direction ~10x too often. The edge at the chosen point is still REPORTED,
just no longer what chooses it.
THREE READINGS OF ONE QUANTITY, AND THEY DISAGREE. "How often does a direction
occur" is measured in three places and gives ~7% (the scan's own tally), ~41%
(the era loop's counters, via this function's old coverage floor) and ~50%
(the ensemble gate's OOS base rate). They cannot all be right. Rather than
pick one silently, ScanDirectionalRatePct() and EraDirectionalRatePct() are
now named accessors, the fitter targets the SCAN - that is the tally the
operator reads, and the one "predict the labels as measured during the scan
phase" names - and the threshold line PRINTS BOTH every time it moves, so the
disagreement is on the record instead of buried in a derived floor.
The ensemble gate's own floor is deliberately NOT changed in this commit. If
the scan is right, a calibrated member covering ~7% of bars cannot clear a
12.4% floor and every model would fail the gate by construction; if the gate
is right, the scan tally is wrong. The CALIBRATION field added in 667f2bc
reports the OOS true class rates directly and settles it in one era - that
measurement comes first, and the floor follows it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-21 10:20:31 -04:00
|
|
|
" REPORTED, not optimised: the margin does not rank these trades, and"
|
|
|
|
|
" selecting on that curve is what made this threshold thrash."
|
|
|
|
|
" | Scan says %.1f%%, era loop says %.1f%% - if these disagree the"
|
|
|
|
|
" populations differ and the scan is the one the operator reads.",
|
|
|
|
|
prevThresh, m_dirConfThreshold, (int)m_dirConfPrimaryBars,
|
|
|
|
|
binCov[bestBin], fromScan ? "scan" : "era-loop", targetPct, bestGap,
|
|
|
|
|
DIR_CONF_THRESHOLD_BINS, binPrec[bestBin], breakEvenPct,
|
feat(breakeven): the break-even every layer scores against prices a trade that always resolves
CostAdjustedBreakEvenPct is risk/(risk+reward) and has no horizon term. It is the win rate a trade
needs when it is CERTAIN to end at one barrier or the other. SimulateTradeOutcome has an explicit
branch for the case where it does not - runs out of horizon, closes at the last bar seen for
whatever P&L that is - so on this label geometry the figure describes a different trade than the
one being replayed.
The gap is measurable and large. Across 21 exit replays today on SP500 H4 the geometric figure read
34.5% while the EA's own R simulation crossed zero between 27.4% (lowest positive) and 28.9%
(highest non-positive). Independent corroboration: the zero-skill reference, computed empirically
over every scored bar as max(winLong,winShort)/bars, reads 25.4% - add cost and it lands on the
same ~28%. The geometric number is the outlier, and every edge printed against it was ~6.5pp too
pessimistic: LSTM's 30.6%-win era reported -4.0pp while its replay returned +0.075 R on the same
trades.
With a timeout share t paying a mean m R apiece, expectancy is w(1+RR) + t(1+m) - 1, so
w* = (1 - t(1+m)) / (1 + RR) = CostAdjustedBreakEvenPct x (1 - t(1+m))
which needs no new geometry - the existing figure already carries 1/(1+RR).
This commit MEASURES ONLY. The replay now separates timeout exits from barrier exits and latches
t and m for the next era to read (the accumulators are zeroed at era start and filled at era end,
so a mid-era reader sees zero trades and would fall back forever). Both break-evens print side by
side on the replay line with t and m beside them, and the threshold line's REPORTED edge - which
selects nothing - switches to the horizon-aware figure so the operator stops reading a wrong sign.
DELIBERATELY NOT CHANGED: LiveMetaGate's veto and the rung selector's BarrierMinReachPct still read
the geometric value. Both are decisions - the second re-derives geometry and therefore relabels -
and t and m have so far only been inferred from a zero-crossing, never seen on a log. One era of
this instrumentation settles that.
The file already contained the argument, one branch away, in the vote-exit comment: a vote exit
produces a CONTINUOUS payoff, not a win or a loss, and that is why an exit-aware gate cannot go on
scoring win-rate against a fixed break-even. A horizon timeout is the same thing, and unlike vote
exits it is on by default.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-21 19:26:25 -04:00
|
|
|
binPrec[bestBin] - breakEvenPct, CostAdjustedBreakEvenPct(),
|
feat(calibration): fit the operating point on the label rate instead of on edge
The margin threshold now sits where the model calls a direction as often as a
direction actually occurs. Nothing else.
WHY THE OLD OBJECTIVE HAD TO GO. It maximised `coverage x (precision -
breakEven)`, and this function's own comments were already the case against
it: over 98 consecutive fits of the shipped SP500 H4 model, correlation
between the chosen threshold and the win rate at it was -0.056, while the
era-to-era spread of that win rate (1.32pp) matched its own binomial SE
(1.25pp) to within 0.07pp. The margin does not rank trades. So the argmax
returned whichever of ~37 bins drew the luckiest sample, and the threshold
teleported 0.42 -> 0.04 -> 0.74 in three eras.
The response at the time was to build a null-of-the-maximum gate, an
effective-sample SE and a parsimony fallback to hold the noise down. All of
that is gone now, because fitting on calibration removes the problem instead
of bounding it: coverage is a ratio against a fixed denominator so it is well
determined at every bin, the target is a measured label rate rather than an
outcome, and nothing is maximised over a noisy curve so there is no best-of-N
to correct for. Net 174 lines out, 62 in.
It deliberately does not chase edge. It cannot - at ~0 measured edge no
operating point has more of it, and pretending otherwise is what produced a
threshold of 0.96 that still passed 60% of bars while the model called a
direction ~10x too often. The edge at the chosen point is still REPORTED,
just no longer what chooses it.
THREE READINGS OF ONE QUANTITY, AND THEY DISAGREE. "How often does a direction
occur" is measured in three places and gives ~7% (the scan's own tally), ~41%
(the era loop's counters, via this function's old coverage floor) and ~50%
(the ensemble gate's OOS base rate). They cannot all be right. Rather than
pick one silently, ScanDirectionalRatePct() and EraDirectionalRatePct() are
now named accessors, the fitter targets the SCAN - that is the tally the
operator reads, and the one "predict the labels as measured during the scan
phase" names - and the threshold line PRINTS BOTH every time it moves, so the
disagreement is on the record instead of buried in a derived floor.
The ensemble gate's own floor is deliberately NOT changed in this commit. If
the scan is right, a calibrated member covering ~7% of bars cannot clear a
12.4% floor and every model would fail the gate by construction; if the gate
is right, the scan tally is wrong. The CALIBRATION field added in 667f2bc
reports the OOS true class rates directly and settles it in one era - that
measurement comes first, and the floor follows it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-21 10:20:31 -04:00
|
|
|
ScanDirectionalRatePct(), eraPct));
|
refactor: split CExpertSignalAIBase implementation by responsibility
ExpertSignalAIBase.mqh was 8216 lines: the class declaration followed by 87
method bodies covering training, labelling, feature extraction, persistence,
chart drawing, online learning, the GA auto-tuner and inference, all in one
file. Train() alone is 1492 lines; a change to arrow drawing meant scrolling
past the era loop.
Moved the bodies into Expert\AIBase\, included at the bottom of the original
after the class declaration:
Training.mqh 1607 era loop, plateau ladder, checkpoint select, deploy
Features.mqh 1093 indicator creation + per-bar input feature vector
ChartUI.mqh 634 arrows, arrow persistence, status panel, cleanup
Persistence.mqh 492 .stats/.cfg sidecars, CPU-inference validation, copy
OnlineLearning.mqh 461 live continual learning, EMA shadow, OOS simulator
Labels.mqh 309 ZigZag pivot labels, async label-cache prebuild
AutoTune.mqh 275 genetic tuner (population, crossover, halving)
Inference.mqh 235 softmax, prior calibration, class priors
ExpertSignalAIBase.mqh 8216 -> 3131 (declaration + topology build only)
This is a pure relocation - verified mechanically, not by eye: HEAD's file
reconstructed from the eight partials plus the surviving remainder is
byte-identical to HEAD, span for span (scratchpad verify_split.py). No
declaration moved, no signature changed, no code rewritten, so behaviour is
unchanged by construction.
Compiles 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 00:42:45 -04:00
|
|
|
}
|
|
|
|
|
//+------------------------------------------------------------------+
|
|
|
|
|
//| EMA-updates the persisted true class base rates from a finished |
|
|
|
|
|
//| era's true class counts. First real measurement seeds directly; |
|
|
|
|
|
//| thereafter blended with the same smoothing as the accuracy/ |
|
|
|
|
|
//| confidence EMAs so one noisy era can't swing the live decision. |
|
|
|
|
|
//+------------------------------------------------------------------+
|
|
|
|
|
void CExpertSignalAIBase::UpdateClassPriors(long buyCnt, long sellCnt, long neutralCnt)
|
|
|
|
|
{
|
|
|
|
|
long tot = buyCnt + sellCnt + neutralCnt;
|
|
|
|
|
if(tot <= 0)
|
|
|
|
|
return;
|
|
|
|
|
double pb = (double)buyCnt / tot, ps = (double)sellCnt / tot, pn = (double)neutralCnt / tot;
|
|
|
|
|
if(m_priorNeutral <= 0.0) // first real measurement
|
|
|
|
|
{
|
|
|
|
|
m_priorBuy = pb;
|
|
|
|
|
m_priorSell = ps;
|
|
|
|
|
m_priorNeutral = pn;
|
|
|
|
|
return;
|
|
|
|
|
}
|
refactor(ai): nine class-imbalance inputs down to two
The imbalance section offered nine controls for one job. Audited against the
code, five of them did not do what their names said at the shipped defaults:
AILogitPriorStrength DEAD - Inference.mqh's post-hoc prior early-returns
whenever the adjusted loss is on, which is default.
OversampleParity DEAD in training - Training.mqh gated the replay loop
on !useLogitAdjustedLoss (correctly, citing Buda et
al. 2018). Live only in the online-learning path.
EnableMinorityReplay DEAD as replay. It survived ONLY as a focal-gamma
damper - "replay minority bars through pass-2
oversampling" was a focal-loss switch.
ConstrainReplay DEAD as a cap; it only chose damper 0.125 vs 0.25.
UseStaticPrior An exact duplicate of FreezePriorCalibration - the two
were OR'd together in the single place either is read.
So they were not five mechanisms fighting; they were one mechanism plus eight
knobs that mostly described machinery that no longer ran. That is worse than
a real conflict, because the log agreed with the names: the label-cache line
printed "reps up to 28x (90% parity) (seeding era 0's class-balance
oversampling)" on every run, describing an oversampling pass that had been
switched off. It is fixed here too - it cost this session a wrong diagnosis.
The one genuine redundancy was focal loss, running at gamma*0.125 alongside
the adjusted loss: two corrections on the same axis, the exact stacking
failure this file already cited Buda et al. for in two other places, damped
by a replay flag whose replay path was itself dead. Removed rather than
re-tuned. The plateau ladder is unaffected - its escape is the learning-rate
warm restart; the gamma anneal beside it only ever stepped toward zero.
WHAT REMAINS is logit-adjusted loss (Menon et al. 2021) plus a prior freeze:
LogitAdjustTau 0 = off; replaces the separate EnableLogitAdjusted-
Loss boolean, since a strength dial where 0 already
means off does not need an on/off switch beside it.
FreezePriorCalibration unchanged.
It is the only one of the six corrections with a consistency guarantee, and
it is consistent for exactly the balanced-error metric checkpoint selection
already ranks on - so the loss and the deploy decision optimize one thing.
The online continual-learning path keeps its own alpha-balanced focal weight,
now as constants pinned to the removed inputs' shipped defaults, so its
behaviour is unchanged. It legitimately needs its own correction:
ApplyLogitAdjustment() only runs inside a training run, so a deployed model
that was reloaded carries no logit offsets and would otherwise stream 31:1
data into itself uncorrected.
The weights-filename fingerprint is BYTE-IDENTICAL. The focal slot was a
double fed to a %d conversion and had always emitted a literal 0; the |MR:
segment is written as the constant its shipped defaults produced. Dropping
either would have re-keyed every model and forced a from-scratch retrain of
the one topology currently converged and trading.
Also removed as orphans: FOCAL_GAMMA_PRESET, MAX_OVERSAMPLE_REPLICAS,
OVERSAMPLE_PARITY_FRACTION, PLATEAU_GAMMA_STEP, and the now-unreachable
"neutralized by prior correction" diagnostic.
Both builds compile 0 errors, 0 warnings. No retrain forced.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 11:46:57 -04:00
|
|
|
//--- (Was `m_useStaticPrior || m_freezePriorCalibration`. Those were two separate user-facing inputs
|
|
|
|
|
//--- whose only effect anywhere in the codebase was this one OR - two controls for one decision.
|
|
|
|
|
//--- UseStaticPrior was removed 2026-07-31; see the class-imbalance audit in Variables\Inputs.mqh.)
|
|
|
|
|
if(m_freezePriorCalibration)
|
refactor: split CExpertSignalAIBase implementation by responsibility
ExpertSignalAIBase.mqh was 8216 lines: the class declaration followed by 87
method bodies covering training, labelling, feature extraction, persistence,
chart drawing, online learning, the GA auto-tuner and inference, all in one
file. Train() alone is 1492 lines; a change to arrow drawing meant scrolling
past the era loop.
Moved the bodies into Expert\AIBase\, included at the bottom of the original
after the class declaration:
Training.mqh 1607 era loop, plateau ladder, checkpoint select, deploy
Features.mqh 1093 indicator creation + per-bar input feature vector
ChartUI.mqh 634 arrows, arrow persistence, status panel, cleanup
Persistence.mqh 492 .stats/.cfg sidecars, CPU-inference validation, copy
OnlineLearning.mqh 461 live continual learning, EMA shadow, OOS simulator
Labels.mqh 309 ZigZag pivot labels, async label-cache prebuild
AutoTune.mqh 275 genetic tuner (population, crossover, halving)
Inference.mqh 235 softmax, prior calibration, class priors
ExpertSignalAIBase.mqh 8216 -> 3131 (declaration + topology build only)
This is a pure relocation - verified mechanically, not by eye: HEAD's file
reconstructed from the eight partials plus the surviving remainder is
byte-identical to HEAD, span for span (scratchpad verify_split.py). No
declaration moved, no signature changed, no code rewritten, so behaviour is
unchanged by construction.
Compiles 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 00:42:45 -04:00
|
|
|
return;
|
|
|
|
|
double k = Net.recentAverageSmoothingFactor;
|
|
|
|
|
if(k < 1.0)
|
|
|
|
|
k = 1.0;
|
|
|
|
|
m_priorBuy += (pb - m_priorBuy) / k;
|
|
|
|
|
m_priorSell += (ps - m_priorSell) / k;
|
|
|
|
|
m_priorNeutral += (pn - m_priorNeutral) / k;
|
|
|
|
|
}
|
|
|
|
|
//+------------------------------------------------------------------+
|
feat(ai): logit-adjusted loss, replacing oversampling and the post-hoc prior
Menon et al. 2021 (ICLR), "Long-tail learning via logit adjustment": add
tau*log(prior_c) to each class logit inside the training gradient. Softmax
CE on adjusted logits is consistent for BALANCED error - the metric
checkpoint selection already ranks on - so the loss and the deploy decision
finally optimize the same thing.
The engine already computed a true softmax + categorical-CE gradient and
wrote it over the per-neuron sigmoid delta, so this is an offset added to
three logits in the two places that gradient is built (backProp scalar path
and backPropOCL). No backend, kernel or DLL change; the forward pass and
every inference path are untouched, which is the point - the network learns
to absorb the offset, so its raw argmax becomes the balanced-optimal
decision with nothing applied at inference.
Replaces rather than stacks. Minority replay is disabled while this is on,
and the post-hoc inference prior is forced off. Stacking is not a
theoretical worry: simulated on the measured 1118/1119/34298 distribution
in the weak-signal regime, plain CE collapses to Neutral (33.4% balanced,
Buy 0%), replay reaches 48.1%, logit adjustment 50.9% with better balance -
and BOTH together score 45.4% with Neutral recall at 0%, worse than either
alone. Buda et al. 2018 predicts exactly that.
Motivation from the six-chart run: every topology took one direction to
~50% recall and abandoned the other, the direction chosen arbitrarily (the
batch-norm control went Buy 1% / Sell 42%, the inverse of the other five).
One era in 1,301 cleared the per-class recall floor.
Fingerprinted conditionally, so the converged 60.7% models on disk keep
their filenames and stay loadable as the fallback.
Both builds compile 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 19:05:14 -04:00
|
|
|
//| Installs the training-time logit offsets - see the declaration. |
|
|
|
|
|
//+------------------------------------------------------------------+
|
|
|
|
|
void CExpertSignalAIBase::ApplyLogitAdjustment(void)
|
|
|
|
|
{
|
|
|
|
|
if(CheckPointer(Net) == POINTER_INVALID)
|
|
|
|
|
return;
|
refactor(ai): nine class-imbalance inputs down to two
The imbalance section offered nine controls for one job. Audited against the
code, five of them did not do what their names said at the shipped defaults:
AILogitPriorStrength DEAD - Inference.mqh's post-hoc prior early-returns
whenever the adjusted loss is on, which is default.
OversampleParity DEAD in training - Training.mqh gated the replay loop
on !useLogitAdjustedLoss (correctly, citing Buda et
al. 2018). Live only in the online-learning path.
EnableMinorityReplay DEAD as replay. It survived ONLY as a focal-gamma
damper - "replay minority bars through pass-2
oversampling" was a focal-loss switch.
ConstrainReplay DEAD as a cap; it only chose damper 0.125 vs 0.25.
UseStaticPrior An exact duplicate of FreezePriorCalibration - the two
were OR'd together in the single place either is read.
So they were not five mechanisms fighting; they were one mechanism plus eight
knobs that mostly described machinery that no longer ran. That is worse than
a real conflict, because the log agreed with the names: the label-cache line
printed "reps up to 28x (90% parity) (seeding era 0's class-balance
oversampling)" on every run, describing an oversampling pass that had been
switched off. It is fixed here too - it cost this session a wrong diagnosis.
The one genuine redundancy was focal loss, running at gamma*0.125 alongside
the adjusted loss: two corrections on the same axis, the exact stacking
failure this file already cited Buda et al. for in two other places, damped
by a replay flag whose replay path was itself dead. Removed rather than
re-tuned. The plateau ladder is unaffected - its escape is the learning-rate
warm restart; the gamma anneal beside it only ever stepped toward zero.
WHAT REMAINS is logit-adjusted loss (Menon et al. 2021) plus a prior freeze:
LogitAdjustTau 0 = off; replaces the separate EnableLogitAdjusted-
Loss boolean, since a strength dial where 0 already
means off does not need an on/off switch beside it.
FreezePriorCalibration unchanged.
It is the only one of the six corrections with a consistency guarantee, and
it is consistent for exactly the balanced-error metric checkpoint selection
already ranks on - so the loss and the deploy decision optimize one thing.
The online continual-learning path keeps its own alpha-balanced focal weight,
now as constants pinned to the removed inputs' shipped defaults, so its
behaviour is unchanged. It legitimately needs its own correction:
ApplyLogitAdjustment() only runs inside a training run, so a deployed model
that was reloaded carries no logit offsets and would otherwise stream 31:1
data into itself uncorrected.
The weights-filename fingerprint is BYTE-IDENTICAL. The focal slot was a
double fed to a %d conversion and had always emitted a literal 0; the |MR:
segment is written as the constant its shipped defaults produced. Dropping
either would have re-keyed every model and forced a from-scratch retrain of
the one topology currently converged and trading.
Also removed as orphans: FOCAL_GAMMA_PRESET, MAX_OVERSAMPLE_REPLICAS,
OVERSAMPLE_PARITY_FRACTION, PLATEAU_GAMMA_STEP, and the now-unreachable
"neutralized by prior correction" diagnostic.
Both builds compile 0 errors, 0 warnings. No retrain forced.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 11:46:57 -04:00
|
|
|
if(m_logitAdjustTau <= 0.0)
|
feat(ai): logit-adjusted loss, replacing oversampling and the post-hoc prior
Menon et al. 2021 (ICLR), "Long-tail learning via logit adjustment": add
tau*log(prior_c) to each class logit inside the training gradient. Softmax
CE on adjusted logits is consistent for BALANCED error - the metric
checkpoint selection already ranks on - so the loss and the deploy decision
finally optimize the same thing.
The engine already computed a true softmax + categorical-CE gradient and
wrote it over the per-neuron sigmoid delta, so this is an offset added to
three logits in the two places that gradient is built (backProp scalar path
and backPropOCL). No backend, kernel or DLL change; the forward pass and
every inference path are untouched, which is the point - the network learns
to absorb the offset, so its raw argmax becomes the balanced-optimal
decision with nothing applied at inference.
Replaces rather than stacks. Minority replay is disabled while this is on,
and the post-hoc inference prior is forced off. Stacking is not a
theoretical worry: simulated on the measured 1118/1119/34298 distribution
in the weak-signal regime, plain CE collapses to Neutral (33.4% balanced,
Buy 0%), replay reaches 48.1%, logit adjustment 50.9% with better balance -
and BOTH together score 45.4% with Neutral recall at 0%, worse than either
alone. Buda et al. 2018 predicts exactly that.
Motivation from the six-chart run: every topology took one direction to
~50% recall and abandoned the other, the direction chosen arbitrarily (the
batch-norm control went Buy 1% / Sell 42%, the inverse of the other five).
One era in 1,301 cleared the per-class recall floor.
Fingerprinted conditionally, so the converged 60.7% models on disk keep
their filenames and stay loadable as the fallback.
Both builds compile 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 19:05:14 -04:00
|
|
|
{
|
|
|
|
|
//--- Clear rather than merely skip: the input can be turned off on a chart that already installed
|
|
|
|
|
//--- offsets this session, and a stale adjustment would keep biasing the gradient silently.
|
|
|
|
|
Net.ClearLogitAdjustment();
|
|
|
|
|
return;
|
|
|
|
|
}
|
2026-08-22 00:25:52 -04:00
|
|
|
//--- Priors not measured yet (era 0 before the first tally, or a model with no .stats): leave
|
|
|
|
|
//--- the gradient unadjusted rather than guessing a distribution. The next era installs them.
|
feat(ai): logit-adjusted loss, replacing oversampling and the post-hoc prior
Menon et al. 2021 (ICLR), "Long-tail learning via logit adjustment": add
tau*log(prior_c) to each class logit inside the training gradient. Softmax
CE on adjusted logits is consistent for BALANCED error - the metric
checkpoint selection already ranks on - so the loss and the deploy decision
finally optimize the same thing.
The engine already computed a true softmax + categorical-CE gradient and
wrote it over the per-neuron sigmoid delta, so this is an offset added to
three logits in the two places that gradient is built (backProp scalar path
and backPropOCL). No backend, kernel or DLL change; the forward pass and
every inference path are untouched, which is the point - the network learns
to absorb the offset, so its raw argmax becomes the balanced-optimal
decision with nothing applied at inference.
Replaces rather than stacks. Minority replay is disabled while this is on,
and the post-hoc inference prior is forced off. Stacking is not a
theoretical worry: simulated on the measured 1118/1119/34298 distribution
in the weak-signal regime, plain CE collapses to Neutral (33.4% balanced,
Buy 0%), replay reaches 48.1%, logit adjustment 50.9% with better balance -
and BOTH together score 45.4% with Neutral recall at 0%, worse than either
alone. Buda et al. 2018 predicts exactly that.
Motivation from the six-chart run: every topology took one direction to
~50% recall and abandoned the other, the direction chosen arbitrarily (the
batch-norm control went Buy 1% / Sell 42%, the inverse of the other five).
One era in 1,301 cleared the per-class recall floor.
Fingerprinted conditionally, so the converged 60.7% models on disk keep
their filenames and stay loadable as the fallback.
Both builds compile 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 19:05:14 -04:00
|
|
|
if(m_priorBuy <= 0.0 || m_priorSell <= 0.0 || m_priorNeutral <= 0.0)
|
|
|
|
|
{
|
fix: the imbalance correction never ran during the auto-tune search
Neutral collapse on all four topologies by era 5 with a 2:6 barrier
(recall Buy 0% / Sell 0% / Neutral 100%), and the panel stuck on
"measuring...". One root cause, and it was not the barrier.
The labels were fine: Buy 25.4% / Sell 22.0% / Neutral 52.5%, which is
exactly gambler's ruin for m=2,k=6 (2/8 = 25% per side), with only 0.1%
of Neutral coming from the vertical barrier - so the new m*k horizon
scaling is right, arguably generous.
What was broken: Train()'s era-start block wrapped UpdateClassPriors() in
`if(!m_evalMode)`. The auto-tune GA scores every candidate in eval mode,
and AutoTuneIndicators ships ON, so on a default configuration EVERY era
of the search ran with unmeasured priors. ApplyLogitAdjustment() requires
measured priors; without them it calls ClearLogitAdjustment() and returns.
So the entire search trained under PLAIN cross-entropy. With a 52.5%
majority class the optimum of plain CE is "always predict Neutral", and
that is precisely what all four models found. The panel followed: its
counters only advance on bars the model CALLED Buy or Sell, so a
collapsed model leaves them at zero and the line reads "measuring..."
forever.
This was latent, not new. It has been true for every auto-tuned run, but
it was invisible while the labels were near-balanced - last night's
accidental 1:1 barrier gave 43/40/17, where plain CE has no majority to
collapse into. Widening the stop to 2*ATR (correctly - 1*ATR is too tight
to survive noise) moved Neutral to the majority and exposed it.
The guard's stated fear cannot happen. These priors are measured from the
LABEL distribution, and the tuner only perturbs indicator periods
(MA/RSI/MACD/Ichimoku/AD). The barrier label depends on ATR, SL_Mode and
TP_Mode - none of which the search touches - so every candidate sees
byte-identical labels and identical priors. There is nothing to
contaminate. What the guard actually protected was the .stats write, and
that is gated separately: eval candidates never checkpoint and never
persist.
Also, because this is the THIRD quiet no-op to cost a run in this
codebase (after the fictional oversampling log line and the shadow-blend
skip):
- ApplyLogitAdjustment() now WARNS when it declines to install, instead
of silently clearing. A mechanism that cannot announce it is not
running is indistinguishable from one that is.
- The panel distinguishes "measuring..." (before era 1, nothing scored
yet - an honest warm-up) from "no directional calls yet" (eras trained,
zero calls - a finding, not a wait).
Both builds compile 0 errors / 0 warnings. No retrain forced by this
commit itself, but the collapsed models must be discarded.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 00:46:24 -04:00
|
|
|
if(m_eraCount > 0 && !m_logitAdjustSkipWarned)
|
|
|
|
|
{
|
|
|
|
|
m_logitAdjustSkipWarned = true;
|
|
|
|
|
Print(ID + ": WARNING - class-imbalance correction is NOT running at era " +
|
|
|
|
|
IntegerToString(m_eraCount) + ": the class priors have never been measured (Buy " +
|
|
|
|
|
DoubleToString(m_priorBuy, 4) + " Sell " + DoubleToString(m_priorSell, 4) + " Neutral " +
|
|
|
|
|
DoubleToString(m_priorNeutral, 4) + "). Training is falling back to plain cross-entropy, "
|
|
|
|
|
"which on a skewed label set collapses to the majority class.");
|
|
|
|
|
}
|
feat(ai): logit-adjusted loss, replacing oversampling and the post-hoc prior
Menon et al. 2021 (ICLR), "Long-tail learning via logit adjustment": add
tau*log(prior_c) to each class logit inside the training gradient. Softmax
CE on adjusted logits is consistent for BALANCED error - the metric
checkpoint selection already ranks on - so the loss and the deploy decision
finally optimize the same thing.
The engine already computed a true softmax + categorical-CE gradient and
wrote it over the per-neuron sigmoid delta, so this is an offset added to
three logits in the two places that gradient is built (backProp scalar path
and backPropOCL). No backend, kernel or DLL change; the forward pass and
every inference path are untouched, which is the point - the network learns
to absorb the offset, so its raw argmax becomes the balanced-optimal
decision with nothing applied at inference.
Replaces rather than stacks. Minority replay is disabled while this is on,
and the post-hoc inference prior is forced off. Stacking is not a
theoretical worry: simulated on the measured 1118/1119/34298 distribution
in the weak-signal regime, plain CE collapses to Neutral (33.4% balanced,
Buy 0%), replay reaches 48.1%, logit adjustment 50.9% with better balance -
and BOTH together score 45.4% with Neutral recall at 0%, worse than either
alone. Buda et al. 2018 predicts exactly that.
Motivation from the six-chart run: every topology took one direction to
~50% recall and abandoned the other, the direction chosen arbitrarily (the
batch-norm control went Buy 1% / Sell 42%, the inverse of the other five).
One era in 1,301 cleared the per-class recall floor.
Fingerprinted conditionally, so the converged 60.7% models on disk keep
their filenames and stay loadable as the fallback.
Both builds compile 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 19:05:14 -04:00
|
|
|
Net.ClearLogitAdjustment();
|
|
|
|
|
return;
|
|
|
|
|
}
|
fix(ai): cap logit-adjustment strength to the head's usable logit range
tau=1.0 inverted the collapse instead of curing it. The head is SIGMOID, so
each output is bounded to [0,1] and the widest logit gap the net can express
between two classes is CLASS_LOGIT_SCALE * (1-0) = 6. The offsets are
tau*log(prior_c), whose spread on this 30:1 imbalance is 3.42 - so tau=1.0
spent 57% of the ENTIRE expressible range on the prior correction.
The network did the only thing available to it: saturate Buy/Sell outputs to
1.0 to overcome a -3.42 training handicap. The offsets are absent at
inference, so that surplus made every bar directional. Measured across all
five still-training charts: Neutral recall 0%, directional calls on ~100% of
bars, win rate 5-7% against a ~6% base rate - no information whatsoever -
while balanced accuracy read a flattering 58-64% because two of its three
terms sat near 95%. OOS accuracy 6%.
Menon et al. assume an unbounded logit head where a 3.42 shift is negligible
against the reachable range. It is not negligible here, so the strength is
now expressed RELATIVE to the range actually available:
tau_eff = min(tau_cfg, LOGIT_ADJUST_MAX_RANGE_FRACTION * SCALE / spread)
At 20% that gives tau 0.35 on this data. Deliberately a fraction rather than
a tau ceiling: it stays correct if CLASS_LOGIT_SCALE changes, if the head
becomes unbounded, or on any symbol whose imbalance differs. The input
remains effective below the cap, so dialling it down needs no rebuild.
Simulated at a signal strength where the task is genuinely learnable, the
precision/recall frontier is monotone: tau 1.0 -> 49.6% call rate at 6.4%
precision (base rate 6.1%, i.e. worthless); tau 0.35 -> 2.0% at 15.5%;
tau 0.15 -> 0.2% at 33.3%. The capped value lands in the same regime the
pre-logit-adjustment run occupied (1-6% of bars at 20-35% win rate).
Also logs the measured priors, the spread, and whether the cap bound.
Both builds compile 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 23:20:07 -04:00
|
|
|
//--- Effective tau, capped so the offsets cannot swamp the head's usable logit range - see
|
2026-08-22 00:25:52 -04:00
|
|
|
//--- LOGIT_ADJUST_MAX_RANGE_FRACTION.
|
fix(ai): cap logit-adjustment strength to the head's usable logit range
tau=1.0 inverted the collapse instead of curing it. The head is SIGMOID, so
each output is bounded to [0,1] and the widest logit gap the net can express
between two classes is CLASS_LOGIT_SCALE * (1-0) = 6. The offsets are
tau*log(prior_c), whose spread on this 30:1 imbalance is 3.42 - so tau=1.0
spent 57% of the ENTIRE expressible range on the prior correction.
The network did the only thing available to it: saturate Buy/Sell outputs to
1.0 to overcome a -3.42 training handicap. The offsets are absent at
inference, so that surplus made every bar directional. Measured across all
five still-training charts: Neutral recall 0%, directional calls on ~100% of
bars, win rate 5-7% against a ~6% base rate - no information whatsoever -
while balanced accuracy read a flattering 58-64% because two of its three
terms sat near 95%. OOS accuracy 6%.
Menon et al. assume an unbounded logit head where a 3.42 shift is negligible
against the reachable range. It is not negligible here, so the strength is
now expressed RELATIVE to the range actually available:
tau_eff = min(tau_cfg, LOGIT_ADJUST_MAX_RANGE_FRACTION * SCALE / spread)
At 20% that gives tau 0.35 on this data. Deliberately a fraction rather than
a tau ceiling: it stays correct if CLASS_LOGIT_SCALE changes, if the head
becomes unbounded, or on any symbol whose imbalance differs. The input
remains effective below the cap, so dialling it down needs no rebuild.
Simulated at a signal strength where the task is genuinely learnable, the
precision/recall frontier is monotone: tau 1.0 -> 49.6% call rate at 6.4%
precision (base rate 6.1%, i.e. worthless); tau 0.35 -> 2.0% at 15.5%;
tau 0.15 -> 0.2% at 33.3%. The capped value lands in the same regime the
pre-logit-adjustment run occupied (1-6% of bars at 20-35% win rate).
Also logs the measured priors, the spread, and whether the cap bound.
Both builds compile 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 23:20:07 -04:00
|
|
|
double lb = MathLog(m_priorBuy), ls = MathLog(m_priorSell), lnn = MathLog(m_priorNeutral);
|
2026-08-21 12:09:07 -04:00
|
|
|
//--- ALL THREE CLASSES, WITH ONE INVARIANT: THE ABSTAIN CLASS IS NEVER SUBSIDISED.
|
|
|
|
|
double mid = (lb + ls + lnn) / 3.0;
|
|
|
|
|
double spread = MathMax(lb, MathMax(ls, lnn)) - MathMin(lb, MathMin(ls, lnn));
|
fix(ai): cap logit-adjustment strength to the head's usable logit range
tau=1.0 inverted the collapse instead of curing it. The head is SIGMOID, so
each output is bounded to [0,1] and the widest logit gap the net can express
between two classes is CLASS_LOGIT_SCALE * (1-0) = 6. The offsets are
tau*log(prior_c), whose spread on this 30:1 imbalance is 3.42 - so tau=1.0
spent 57% of the ENTIRE expressible range on the prior correction.
The network did the only thing available to it: saturate Buy/Sell outputs to
1.0 to overcome a -3.42 training handicap. The offsets are absent at
inference, so that surplus made every bar directional. Measured across all
five still-training charts: Neutral recall 0%, directional calls on ~100% of
bars, win rate 5-7% against a ~6% base rate - no information whatsoever -
while balanced accuracy read a flattering 58-64% because two of its three
terms sat near 95%. OOS accuracy 6%.
Menon et al. assume an unbounded logit head where a 3.42 shift is negligible
against the reachable range. It is not negligible here, so the strength is
now expressed RELATIVE to the range actually available:
tau_eff = min(tau_cfg, LOGIT_ADJUST_MAX_RANGE_FRACTION * SCALE / spread)
At 20% that gives tau 0.35 on this data. Deliberately a fraction rather than
a tau ceiling: it stays correct if CLASS_LOGIT_SCALE changes, if the head
becomes unbounded, or on any symbol whose imbalance differs. The input
remains effective below the cap, so dialling it down needs no rebuild.
Simulated at a signal strength where the task is genuinely learnable, the
precision/recall frontier is monotone: tau 1.0 -> 49.6% call rate at 6.4%
precision (base rate 6.1%, i.e. worthless); tau 0.35 -> 2.0% at 15.5%;
tau 0.15 -> 0.2% at 33.3%. The capped value lands in the same regime the
pre-logit-adjustment run occupied (1-6% of bars at 20-35% win rate).
Also logs the measured priors, the spread, and whether the cap bound.
Both builds compile 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 23:20:07 -04:00
|
|
|
double tauEff = m_logitAdjustTau;
|
|
|
|
|
if(spread > 0.0)
|
|
|
|
|
{
|
|
|
|
|
double cap = LOGIT_ADJUST_MAX_RANGE_FRACTION * CLASS_LOGIT_SCALE / spread;
|
|
|
|
|
if(tauEff > cap)
|
|
|
|
|
tauEff = cap;
|
|
|
|
|
}
|
|
|
|
|
if(!m_logitAdjustLogged)
|
|
|
|
|
{
|
|
|
|
|
m_logitAdjustLogged = true;
|
2026-08-21 12:09:07 -04:00
|
|
|
//--- Says which way the abstain class is being pushed, because that is the whole question this
|
|
|
|
|
//--- correction has got wrong in both directions before.
|
fix(ai): cap logit-adjustment strength to the head's usable logit range
tau=1.0 inverted the collapse instead of curing it. The head is SIGMOID, so
each output is bounded to [0,1] and the widest logit gap the net can express
between two classes is CLASS_LOGIT_SCALE * (1-0) = 6. The offsets are
tau*log(prior_c), whose spread on this 30:1 imbalance is 3.42 - so tau=1.0
spent 57% of the ENTIRE expressible range on the prior correction.
The network did the only thing available to it: saturate Buy/Sell outputs to
1.0 to overcome a -3.42 training handicap. The offsets are absent at
inference, so that surplus made every bar directional. Measured across all
five still-training charts: Neutral recall 0%, directional calls on ~100% of
bars, win rate 5-7% against a ~6% base rate - no information whatsoever -
while balanced accuracy read a flattering 58-64% because two of its three
terms sat near 95%. OOS accuracy 6%.
Menon et al. assume an unbounded logit head where a 3.42 shift is negligible
against the reachable range. It is not negligible here, so the strength is
now expressed RELATIVE to the range actually available:
tau_eff = min(tau_cfg, LOGIT_ADJUST_MAX_RANGE_FRACTION * SCALE / spread)
At 20% that gives tau 0.35 on this data. Deliberately a fraction rather than
a tau ceiling: it stays correct if CLASS_LOGIT_SCALE changes, if the head
becomes unbounded, or on any symbol whose imbalance differs. The input
remains effective below the cap, so dialling it down needs no rebuild.
Simulated at a signal strength where the task is genuinely learnable, the
precision/recall frontier is monotone: tau 1.0 -> 49.6% call rate at 6.4%
precision (base rate 6.1%, i.e. worthless); tau 0.35 -> 2.0% at 15.5%;
tau 0.15 -> 0.2% at 33.3%. The capped value lands in the same regime the
pre-logit-adjustment run occupied (1-6% of bars at 20-35% win rate).
Also logs the measured priors, the spread, and whether the cap bound.
Both builds compile 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 23:20:07 -04:00
|
|
|
Print(ID + ": logit adjustment - measured priors Buy " + DoubleToString(m_priorBuy * 100.0, 2) +
|
|
|
|
|
"% Sell " + DoubleToString(m_priorSell * 100.0, 2) + "% Neutral " +
|
2026-08-21 12:09:07 -04:00
|
|
|
DoubleToString(m_priorNeutral * 100.0, 2) + "% | APPLIED across all three, log-prior"
|
|
|
|
|
" spread " + DoubleToString(spread, 2) + " | Neutral offset " +
|
|
|
|
|
DoubleToString(MathMax(tauEff * (lnn - mid), 0.0), 2) +
|
fix(imbalance): the class-imbalance correction was subsidising the abstain class
NOT COMPILED - user compiles.
Root cause of the Neutral collapse. Logit adjustment (Menon et al. 2020) makes a
classifier Bayes-optimal for BALANCED error by subsidising rare classes. It was
wired here when Neutral was the DOMINANT class - the "big move up / big move down
/ nothing much" era, where the correction pulled the model off the majority.
The triple-barrier relabel (b4a704d) inverted the distribution. The barriers are
now the EA's own SL/TP, so ~89% of bars RESOLVE and only timeouts are Neutral.
Measured on SP500 H4, from the EA's own log:
measured priors Buy 48.26% Sell 41.13% Neutral 10.61%
log-prior spread 1.52 | tau 1.00 CAPPED to 0.79
Neutral became the RAREST class, so the correction started subsidising it - by
tau*(log pB - log pN) = 1.20 logits. With no directional edge to overcome that
(direction is closed at best-of-999, p=1.0000), the model took the free lunch:
OOS recall Buy:1% Sell:0% Neutral:100%
OOS raw out spread avg 0.9993 (softmax saturated, near one-hot)
dW/W bn1 0.000% bn3 2.0% bn5 6.5% (input weight block frozen; head twitching)
The anti-collapse mechanism was the collapse. The recall gate needs >=40% on all
three classes, so nothing could ever deploy and the plateau ladder burned eras.
Present in both runs today (b6b5 froze bn1 by era ~719, 17ae by ~169), so it
predates this week's work.
FIX: the correction now spans the DECIDABLE classes only, Buy against Sell,
centred on their midpoint, with Neutral pinned at offset 0. Neutral is the
ABSTAIN outcome and abstention already has a better owner - m_dirConfThreshold,
refitted every era on the held-out calibration band against a coverage floor and
the measured break-even. Subsidising the abstain class does that job twice and
spends the whole correction suppressing the only decisions that can pay.
What still gets corrected is real: a trending symbol resolves more long barriers
than short, and uncorrected the model inherits that as a standing directional
bias. Here it is log(0.4826)-log(0.4113) = 0.16, so the offsets are tiny - the
correct answer, not a broken one. The two traded classes were already balanced;
the old spread of 1.52 only ever described how rare a timeout is.
Everything is derived from the measured distribution, as requested - offsets from
the priors, cap from the resulting spread. tau itself is deliberately NOT fitted:
tuning it against the same data that selects the checkpoint would add another
search dimension to a project that has been burned by exactly that. tau=1 is the
theory value and the cap (now ~9.5x looser at spread 0.16) will rarely bind.
Log line now reports both spreads and, when the abstain class is the rarest, says
how much the old form would have boosted it. Fingerprint |LA:<tau> -> |LA:<tau>:BS
so models trained under the all-three form re-key instead of resuming.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-16 22:43:30 -04:00
|
|
|
(m_priorNeutral < m_priorBuy && m_priorNeutral < m_priorSell
|
2026-08-21 12:09:07 -04:00
|
|
|
? " - CLAMPED to zero: Neutral is the RAREST class here, and paying the model to abstain is"
|
|
|
|
|
" the 2026-08-16 collapse. Uncapped it would have been " +
|
|
|
|
|
DoubleToString(tauEff * (lnn - mid), 2)
|
|
|
|
|
: " - Neutral is over-represented, so this PENALISES abstention") +
|
fix(imbalance): the class-imbalance correction was subsidising the abstain class
NOT COMPILED - user compiles.
Root cause of the Neutral collapse. Logit adjustment (Menon et al. 2020) makes a
classifier Bayes-optimal for BALANCED error by subsidising rare classes. It was
wired here when Neutral was the DOMINANT class - the "big move up / big move down
/ nothing much" era, where the correction pulled the model off the majority.
The triple-barrier relabel (b4a704d) inverted the distribution. The barriers are
now the EA's own SL/TP, so ~89% of bars RESOLVE and only timeouts are Neutral.
Measured on SP500 H4, from the EA's own log:
measured priors Buy 48.26% Sell 41.13% Neutral 10.61%
log-prior spread 1.52 | tau 1.00 CAPPED to 0.79
Neutral became the RAREST class, so the correction started subsidising it - by
tau*(log pB - log pN) = 1.20 logits. With no directional edge to overcome that
(direction is closed at best-of-999, p=1.0000), the model took the free lunch:
OOS recall Buy:1% Sell:0% Neutral:100%
OOS raw out spread avg 0.9993 (softmax saturated, near one-hot)
dW/W bn1 0.000% bn3 2.0% bn5 6.5% (input weight block frozen; head twitching)
The anti-collapse mechanism was the collapse. The recall gate needs >=40% on all
three classes, so nothing could ever deploy and the plateau ladder burned eras.
Present in both runs today (b6b5 froze bn1 by era ~719, 17ae by ~169), so it
predates this week's work.
FIX: the correction now spans the DECIDABLE classes only, Buy against Sell,
centred on their midpoint, with Neutral pinned at offset 0. Neutral is the
ABSTAIN outcome and abstention already has a better owner - m_dirConfThreshold,
refitted every era on the held-out calibration band against a coverage floor and
the measured break-even. Subsidising the abstain class does that job twice and
spends the whole correction suppressing the only decisions that can pay.
What still gets corrected is real: a trending symbol resolves more long barriers
than short, and uncorrected the model inherits that as a standing directional
bias. Here it is log(0.4826)-log(0.4113) = 0.16, so the offsets are tiny - the
correct answer, not a broken one. The two traded classes were already balanced;
the old spread of 1.52 only ever described how rare a timeout is.
Everything is derived from the measured distribution, as requested - offsets from
the priors, cap from the resulting spread. tau itself is deliberately NOT fitted:
tuning it against the same data that selects the checkpoint would add another
search dimension to a project that has been burned by exactly that. tau=1 is the
theory value and the cap (now ~9.5x looser at spread 0.16) will rarely bind.
Log line now reports both spreads and, when the abstain class is the rarest, says
how much the old form would have boosted it. Fingerprint |LA:<tau> -> |LA:<tau>:BS
so models trained under the all-three form re-key instead of resuming.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-16 22:43:30 -04:00
|
|
|
" against a logit range of " +
|
fix(ai): cap logit-adjustment strength to the head's usable logit range
tau=1.0 inverted the collapse instead of curing it. The head is SIGMOID, so
each output is bounded to [0,1] and the widest logit gap the net can express
between two classes is CLASS_LOGIT_SCALE * (1-0) = 6. The offsets are
tau*log(prior_c), whose spread on this 30:1 imbalance is 3.42 - so tau=1.0
spent 57% of the ENTIRE expressible range on the prior correction.
The network did the only thing available to it: saturate Buy/Sell outputs to
1.0 to overcome a -3.42 training handicap. The offsets are absent at
inference, so that surplus made every bar directional. Measured across all
five still-training charts: Neutral recall 0%, directional calls on ~100% of
bars, win rate 5-7% against a ~6% base rate - no information whatsoever -
while balanced accuracy read a flattering 58-64% because two of its three
terms sat near 95%. OOS accuracy 6%.
Menon et al. assume an unbounded logit head where a 3.42 shift is negligible
against the reachable range. It is not negligible here, so the strength is
now expressed RELATIVE to the range actually available:
tau_eff = min(tau_cfg, LOGIT_ADJUST_MAX_RANGE_FRACTION * SCALE / spread)
At 20% that gives tau 0.35 on this data. Deliberately a fraction rather than
a tau ceiling: it stays correct if CLASS_LOGIT_SCALE changes, if the head
becomes unbounded, or on any symbol whose imbalance differs. The input
remains effective below the cap, so dialling it down needs no rebuild.
Simulated at a signal strength where the task is genuinely learnable, the
precision/recall frontier is monotone: tau 1.0 -> 49.6% call rate at 6.4%
precision (base rate 6.1%, i.e. worthless); tau 0.35 -> 2.0% at 15.5%;
tau 0.15 -> 0.2% at 33.3%. The capped value lands in the same regime the
pre-logit-adjustment run occupied (1-6% of bars at 20-35% win rate).
Also logs the measured priors, the spread, and whether the cap bound.
Both builds compile 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 23:20:07 -04:00
|
|
|
DoubleToString(CLASS_LOGIT_SCALE, 1) + " | tau " + DoubleToString(m_logitAdjustTau, 2) +
|
|
|
|
|
(tauEff < m_logitAdjustTau
|
|
|
|
|
? " CAPPED to " + DoubleToString(tauEff, 2) + " (uncapped it would consume " +
|
|
|
|
|
DoubleToString(100.0 * spread * m_logitAdjustTau / CLASS_LOGIT_SCALE, 0) +
|
|
|
|
|
"% of the range and saturate the head)"
|
|
|
|
|
: " (uncapped - within budget)"));
|
|
|
|
|
}
|
feat(ai): logit-adjusted loss, replacing oversampling and the post-hoc prior
Menon et al. 2021 (ICLR), "Long-tail learning via logit adjustment": add
tau*log(prior_c) to each class logit inside the training gradient. Softmax
CE on adjusted logits is consistent for BALANCED error - the metric
checkpoint selection already ranks on - so the loss and the deploy decision
finally optimize the same thing.
The engine already computed a true softmax + categorical-CE gradient and
wrote it over the per-neuron sigmoid delta, so this is an offset added to
three logits in the two places that gradient is built (backProp scalar path
and backPropOCL). No backend, kernel or DLL change; the forward pass and
every inference path are untouched, which is the point - the network learns
to absorb the offset, so its raw argmax becomes the balanced-optimal
decision with nothing applied at inference.
Replaces rather than stacks. Minority replay is disabled while this is on,
and the post-hoc inference prior is forced off. Stacking is not a
theoretical worry: simulated on the measured 1118/1119/34298 distribution
in the weak-signal regime, plain CE collapses to Neutral (33.4% balanced,
Buy 0%), replay reaches 48.1%, logit adjustment 50.9% with better balance -
and BOTH together score 45.4% with Neutral recall at 0%, worse than either
alone. Buda et al. 2018 predicts exactly that.
Motivation from the six-chart run: every topology took one direction to
~50% recall and abandoned the other, the direction chosen arbitrarily (the
batch-norm control went Buy 1% / Sell 42%, the inverse of the other five).
One era in 1,301 cleared the per-class recall floor.
Fingerprinted conditionally, so the converged 60.7% models on disk keep
their filenames and stay loadable as the fallback.
Both builds compile 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 19:05:14 -04:00
|
|
|
//--- ORDERED to match the output layer: [0]=Buy, [1]=Sell, [2]=Neutral - the order
|
|
|
|
|
//--- BuildFreshTopology emits and the order the softmax gradient reads (AI\Network.mqh).
|
|
|
|
|
double offsets[3];
|
fix(imbalance): the class-imbalance correction was subsidising the abstain class
NOT COMPILED - user compiles.
Root cause of the Neutral collapse. Logit adjustment (Menon et al. 2020) makes a
classifier Bayes-optimal for BALANCED error by subsidising rare classes. It was
wired here when Neutral was the DOMINANT class - the "big move up / big move down
/ nothing much" era, where the correction pulled the model off the majority.
The triple-barrier relabel (b4a704d) inverted the distribution. The barriers are
now the EA's own SL/TP, so ~89% of bars RESOLVE and only timeouts are Neutral.
Measured on SP500 H4, from the EA's own log:
measured priors Buy 48.26% Sell 41.13% Neutral 10.61%
log-prior spread 1.52 | tau 1.00 CAPPED to 0.79
Neutral became the RAREST class, so the correction started subsidising it - by
tau*(log pB - log pN) = 1.20 logits. With no directional edge to overcome that
(direction is closed at best-of-999, p=1.0000), the model took the free lunch:
OOS recall Buy:1% Sell:0% Neutral:100%
OOS raw out spread avg 0.9993 (softmax saturated, near one-hot)
dW/W bn1 0.000% bn3 2.0% bn5 6.5% (input weight block frozen; head twitching)
The anti-collapse mechanism was the collapse. The recall gate needs >=40% on all
three classes, so nothing could ever deploy and the plateau ladder burned eras.
Present in both runs today (b6b5 froze bn1 by era ~719, 17ae by ~169), so it
predates this week's work.
FIX: the correction now spans the DECIDABLE classes only, Buy against Sell,
centred on their midpoint, with Neutral pinned at offset 0. Neutral is the
ABSTAIN outcome and abstention already has a better owner - m_dirConfThreshold,
refitted every era on the held-out calibration band against a coverage floor and
the measured break-even. Subsidising the abstain class does that job twice and
spends the whole correction suppressing the only decisions that can pay.
What still gets corrected is real: a trending symbol resolves more long barriers
than short, and uncorrected the model inherits that as a standing directional
bias. Here it is log(0.4826)-log(0.4113) = 0.16, so the offsets are tiny - the
correct answer, not a broken one. The two traded classes were already balanced;
the old spread of 1.52 only ever described how rare a timeout is.
Everything is derived from the measured distribution, as requested - offsets from
the priors, cap from the resulting spread. tau itself is deliberately NOT fitted:
tuning it against the same data that selects the checkpoint would add another
search dimension to a project that has been burned by exactly that. tau=1 is the
theory value and the cap (now ~9.5x looser at spread 0.16) will rarely bind.
Log line now reports both spreads and, when the abstain class is the rarest, says
how much the old form would have boosted it. Fingerprint |LA:<tau> -> |LA:<tau>:BS
so models trained under the all-three form re-key instead of resuming.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-16 22:43:30 -04:00
|
|
|
offsets[0] = tauEff * (lb - mid);
|
|
|
|
|
offsets[1] = tauEff * (ls - mid);
|
2026-08-21 12:09:07 -04:00
|
|
|
//--- A NEGATIVE offset here is a subsidy: it makes the head produce a larger raw Neutral logit to
|
|
|
|
|
//--- classify Neutral correctly, which is exactly what wins Neutral more bars at inference. Clamped
|
|
|
|
|
//--- away. A positive one penalises over-abstention and is allowed through.
|
|
|
|
|
offsets[2] = MathMax(tauEff * (lnn - mid), 0.0);
|
feat(ai): logit-adjusted loss, replacing oversampling and the post-hoc prior
Menon et al. 2021 (ICLR), "Long-tail learning via logit adjustment": add
tau*log(prior_c) to each class logit inside the training gradient. Softmax
CE on adjusted logits is consistent for BALANCED error - the metric
checkpoint selection already ranks on - so the loss and the deploy decision
finally optimize the same thing.
The engine already computed a true softmax + categorical-CE gradient and
wrote it over the per-neuron sigmoid delta, so this is an offset added to
three logits in the two places that gradient is built (backProp scalar path
and backPropOCL). No backend, kernel or DLL change; the forward pass and
every inference path are untouched, which is the point - the network learns
to absorb the offset, so its raw argmax becomes the balanced-optimal
decision with nothing applied at inference.
Replaces rather than stacks. Minority replay is disabled while this is on,
and the post-hoc inference prior is forced off. Stacking is not a
theoretical worry: simulated on the measured 1118/1119/34298 distribution
in the weak-signal regime, plain CE collapses to Neutral (33.4% balanced,
Buy 0%), replay reaches 48.1%, logit adjustment 50.9% with better balance -
and BOTH together score 45.4% with Neutral recall at 0%, worse than either
alone. Buda et al. 2018 predicts exactly that.
Motivation from the six-chart run: every topology took one direction to
~50% recall and abandoned the other, the direction chosen arbitrarily (the
batch-norm control went Buy 1% / Sell 42%, the inverse of the other five).
One era in 1,301 cleared the per-class recall floor.
Fingerprinted conditionally, so the converged 60.7% models on disk keep
their filenames and stay loadable as the fallback.
Both builds compile 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 19:05:14 -04:00
|
|
|
Net.SetLogitAdjustment(offsets);
|
|
|
|
|
}
|
|
|
|
|
//+------------------------------------------------------------------+
|
2026-08-22 00:30:14 -04:00
|
|
|
//| Converts a double to ENUM_SIGNAL. |
|
refactor: split CExpertSignalAIBase implementation by responsibility
ExpertSignalAIBase.mqh was 8216 lines: the class declaration followed by 87
method bodies covering training, labelling, feature extraction, persistence,
chart drawing, online learning, the GA auto-tuner and inference, all in one
file. Train() alone is 1492 lines; a change to arrow drawing meant scrolling
past the era loop.
Moved the bodies into Expert\AIBase\, included at the bottom of the original
after the class declaration:
Training.mqh 1607 era loop, plateau ladder, checkpoint select, deploy
Features.mqh 1093 indicator creation + per-bar input feature vector
ChartUI.mqh 634 arrows, arrow persistence, status panel, cleanup
Persistence.mqh 492 .stats/.cfg sidecars, CPU-inference validation, copy
OnlineLearning.mqh 461 live continual learning, EMA shadow, OOS simulator
Labels.mqh 309 ZigZag pivot labels, async label-cache prebuild
AutoTune.mqh 275 genetic tuner (population, crossover, halving)
Inference.mqh 235 softmax, prior calibration, class priors
ExpertSignalAIBase.mqh 8216 -> 3131 (declaration + topology build only)
This is a pure relocation - verified mechanically, not by eye: HEAD's file
reconstructed from the eight partials plus the surviving remainder is
byte-identical to HEAD, span for span (scratchpad verify_split.py). No
declaration moved, no signature changed, no code rewritten, so behaviour is
unchanged by construction.
Compiles 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 00:42:45 -04:00
|
|
|
//+------------------------------------------------------------------+
|
|
|
|
|
ENUM_SIGNAL CExpertSignalAIBase::DoubleToSignal(double value)
|
|
|
|
|
{
|
|
|
|
|
value = NormalizeDouble(value, 2); // Round 'value' to two decimal places
|
|
|
|
|
if(value < -1.0 || value > 1.0)
|
|
|
|
|
return Undefine; // out of range, e.g. the -2 "not yet studied" sentinel
|
|
|
|
|
if(m_outputNeuronsCount == 3)
|
|
|
|
|
{
|
|
|
|
|
if(value > 0.0)
|
|
|
|
|
return Buy;
|
|
|
|
|
if(value < 0.0)
|
|
|
|
|
return Sell;
|
|
|
|
|
return Neutral;
|
|
|
|
|
}
|
|
|
|
|
if(value > 0.50)
|
|
|
|
|
return Buy;
|
|
|
|
|
if(value < -0.50)
|
|
|
|
|
return Sell;
|
|
|
|
|
return Neutral;
|
|
|
|
|
}
|
feat(hud): per-member neuron lines + a vote label that moves as the nets learn
Both 2026-08-19 reports were the same staleness: every source behind the
label was an ERA artifact (live cache refills at pass-3 completion, the
snapshot copies once per era, dPrevSignal is the frozen purge-band edge
bar) - so the readout stepped at era cadence at best, stayed glued to
one direction, and lagged the era counter.
DisplayInference(): throttled (4s, 1s across an era boundary),
SIDE-EFFECT-FREE forward of the current decision bar (window ending on
bar 1, same question the live path asks) through the LEARNER net.
Batch-norm running stats are bracketed frozen/RESTORED via the new
CNet::GetBatchNormFrozen() + CNeuronBatchNormOCL::StatsFrozen() - restore,
not unfreeze, because a display tick can land between pass-3 chunks whose
whole scan holds them frozen. Writes nothing a trading or training path
reads (dPrevSignal, NMS state, tallies, watermarks all untouched;
RefreshLatestSignal is not reusable here precisely because it writes all
of them). LSTM safe by construction: h/c zeroed per forward.
ProspectiveVote() reads the fresh forward as its FIRST source; the
era-artifact chain becomes the fallback (meta head, warm-up, window
holes).
DisplayHudLine(): the reference library's training label, per ensemble
member - name, output activations (softmax probs or raw scalar), the
decision, its weighted vote (the exact consensus numerator term), era,
recent average error, "(trn)" while not vote-capable. Rendered under the
vote line in RefreshVoteReadout BEFORE the live-vote defer (member lines
are telemetry, not tradable readings), coloured by the member's own
direction in muted tones - the vote line's strict
green-only-when-it-would-trade rule is untouched.
NOT COMPILED - user compiles in MetaEditor.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-19 08:48:34 -04:00
|
|
|
//+------------------------------------------------------------------+
|
2026-08-22 00:30:14 -04:00
|
|
|
//| Throttled, SIDE-EFFECT-FREE forward of the current decision bar, |
|
|
|
|
|
//| for display only (the HUD member lines and the prospective |
|
|
|
|
|
//| vote). |
|
feat(hud): per-member neuron lines + a vote label that moves as the nets learn
Both 2026-08-19 reports were the same staleness: every source behind the
label was an ERA artifact (live cache refills at pass-3 completion, the
snapshot copies once per era, dPrevSignal is the frozen purge-band edge
bar) - so the readout stepped at era cadence at best, stayed glued to
one direction, and lagged the era counter.
DisplayInference(): throttled (4s, 1s across an era boundary),
SIDE-EFFECT-FREE forward of the current decision bar (window ending on
bar 1, same question the live path asks) through the LEARNER net.
Batch-norm running stats are bracketed frozen/RESTORED via the new
CNet::GetBatchNormFrozen() + CNeuronBatchNormOCL::StatsFrozen() - restore,
not unfreeze, because a display tick can land between pass-3 chunks whose
whole scan holds them frozen. Writes nothing a trading or training path
reads (dPrevSignal, NMS state, tallies, watermarks all untouched;
RefreshLatestSignal is not reusable here precisely because it writes all
of them). LSTM safe by construction: h/c zeroed per forward.
ProspectiveVote() reads the fresh forward as its FIRST source; the
era-artifact chain becomes the fallback (meta head, warm-up, window
holes).
DisplayHudLine(): the reference library's training label, per ensemble
member - name, output activations (softmax probs or raw scalar), the
decision, its weighted vote (the exact consensus numerator term), era,
recent average error, "(trn)" while not vote-capable. Rendered under the
vote line in RefreshVoteReadout BEFORE the live-vote defer (member lines
are telemetry, not tradable readings), coloured by the member's own
direction in muted tones - the vote line's strict
green-only-when-it-would-trade rule is untouched.
NOT COMPILED - user compiles in MetaEditor.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-19 08:48:34 -04:00
|
|
|
//+------------------------------------------------------------------+
|
2026-08-19 10:07:12 -04:00
|
|
|
//--- (uint) so the throttle comparison is unsigned-vs-unsigned: age is a uint tick delta, and
|
|
|
|
|
//--- the ternary picking between these is a runtime expression the compiler cannot constant-
|
|
|
|
|
//--- fold, so bare int literals here drew a sign-mismatch warning (reported 2026-08-19).
|
|
|
|
|
#define DISPLAY_FWD_MIN_MS ((uint)4000)
|
|
|
|
|
#define DISPLAY_FWD_ERA_MS ((uint)1000)
|
feat(hud): per-member neuron lines + a vote label that moves as the nets learn
Both 2026-08-19 reports were the same staleness: every source behind the
label was an ERA artifact (live cache refills at pass-3 completion, the
snapshot copies once per era, dPrevSignal is the frozen purge-band edge
bar) - so the readout stepped at era cadence at best, stayed glued to
one direction, and lagged the era counter.
DisplayInference(): throttled (4s, 1s across an era boundary),
SIDE-EFFECT-FREE forward of the current decision bar (window ending on
bar 1, same question the live path asks) through the LEARNER net.
Batch-norm running stats are bracketed frozen/RESTORED via the new
CNet::GetBatchNormFrozen() + CNeuronBatchNormOCL::StatsFrozen() - restore,
not unfreeze, because a display tick can land between pass-3 chunks whose
whole scan holds them frozen. Writes nothing a trading or training path
reads (dPrevSignal, NMS state, tallies, watermarks all untouched;
RefreshLatestSignal is not reusable here precisely because it writes all
of them). LSTM safe by construction: h/c zeroed per forward.
ProspectiveVote() reads the fresh forward as its FIRST source; the
era-artifact chain becomes the fallback (meta head, warm-up, window
holes).
DisplayHudLine(): the reference library's training label, per ensemble
member - name, output activations (softmax probs or raw scalar), the
decision, its weighted vote (the exact consensus numerator term), era,
recent average error, "(trn)" while not vote-capable. Rendered under the
vote line in RefreshVoteReadout BEFORE the live-vote defer (member lines
are telemetry, not tradable readings), coloured by the member's own
direction in muted tones - the vote line's strict
green-only-when-it-would-trade rule is untouched.
NOT COMPILED - user compiles in MetaEditor.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-19 08:48:34 -04:00
|
|
|
bool CExpertSignalAIBase::DisplayInference(void)
|
|
|
|
|
{
|
|
|
|
|
//--- Meta head consumes fired candidates, not a bare bar window - a candidate-less forward is
|
|
|
|
|
//--- width-mismatched against its input layer. Same guard as RefreshConvergedSignal.
|
|
|
|
|
if(IsMetaTarget())
|
|
|
|
|
return false;
|
|
|
|
|
if(CheckPointer(Net) == POINTER_INVALID)
|
|
|
|
|
return false;
|
|
|
|
|
if(m_outputNeuronsCount != 1 && m_outputNeuronsCount != 3)
|
|
|
|
|
return false;
|
|
|
|
|
uint now = GetTickCount();
|
|
|
|
|
uint age = now - m_dispStamp; // unsigned subtraction survives the 49-day wrap
|
|
|
|
|
bool eraMoved = ((long)m_eraCount != m_dispEra);
|
|
|
|
|
if(m_dispStamp != 0 && age < (eraMoved ? DISPLAY_FWD_ERA_MS : DISPLAY_FWD_MIN_MS))
|
|
|
|
|
return m_dispValid; // serve the cache (or keep failing quietly) until the throttle opens
|
|
|
|
|
m_dispStamp = now;
|
|
|
|
|
if(!BuildFeatureWindow(1))
|
|
|
|
|
return m_dispValid; // window not buildable (warm-up, indicator hole): keep the last read
|
|
|
|
|
//--- Save/restore, NOT set/clear - see the header. Frozen, this forward is a pure function.
|
|
|
|
|
bool bnWasFrozen = Net.GetBatchNormFrozen();
|
|
|
|
|
if(!bnWasFrozen)
|
|
|
|
|
Net.SetBatchNormFrozen(true);
|
|
|
|
|
bool fwdOk = Net.feedForward(TempData);
|
|
|
|
|
if(fwdOk)
|
|
|
|
|
Net.getResults(TempData);
|
|
|
|
|
if(!bnWasFrozen)
|
|
|
|
|
Net.SetBatchNormFrozen(false);
|
|
|
|
|
if(!fwdOk)
|
|
|
|
|
return m_dispValid;
|
|
|
|
|
if(m_outputNeuronsCount == 1)
|
|
|
|
|
{
|
|
|
|
|
double v = TempData.At(0);
|
|
|
|
|
if(!MathIsValidNumber(v))
|
|
|
|
|
return m_dispValid; // NaN net: keep the last finite read, the NaN latch reports elsewhere
|
|
|
|
|
m_dispProbs[0] = v;
|
|
|
|
|
m_dispProbs[1] = 0.0;
|
|
|
|
|
m_dispProbs[2] = 0.0;
|
|
|
|
|
m_dispSignal = v;
|
|
|
|
|
}
|
|
|
|
|
else
|
|
|
|
|
{
|
|
|
|
|
//--- Same two calls, same order, as the live decision in RefreshLatestSignal: softmax INTO
|
2026-08-22 00:25:52 -04:00
|
|
|
//--- TempData, then the strict-majority read.
|
feat(hud): per-member neuron lines + a vote label that moves as the nets learn
Both 2026-08-19 reports were the same staleness: every source behind the
label was an ERA artifact (live cache refills at pass-3 completion, the
snapshot copies once per era, dPrevSignal is the frozen purge-band edge
bar) - so the readout stepped at era cadence at best, stayed glued to
one direction, and lagged the era counter.
DisplayInference(): throttled (4s, 1s across an era boundary),
SIDE-EFFECT-FREE forward of the current decision bar (window ending on
bar 1, same question the live path asks) through the LEARNER net.
Batch-norm running stats are bracketed frozen/RESTORED via the new
CNet::GetBatchNormFrozen() + CNeuronBatchNormOCL::StatsFrozen() - restore,
not unfreeze, because a display tick can land between pass-3 chunks whose
whole scan holds them frozen. Writes nothing a trading or training path
reads (dPrevSignal, NMS state, tallies, watermarks all untouched;
RefreshLatestSignal is not reusable here precisely because it writes all
of them). LSTM safe by construction: h/c zeroed per forward.
ProspectiveVote() reads the fresh forward as its FIRST source; the
era-artifact chain becomes the fallback (meta head, warm-up, window
holes).
DisplayHudLine(): the reference library's training label, per ensemble
member - name, output activations (softmax probs or raw scalar), the
decision, its weighted vote (the exact consensus numerator term), era,
recent average error, "(trn)" while not vote-capable. Rendered under the
vote line in RefreshVoteReadout BEFORE the live-vote defer (member lines
are telemetry, not tradable readings), coloured by the member's own
direction in muted tones - the vote line's strict
green-only-when-it-would-trade rule is untouched.
NOT COMPILED - user compiles in MetaEditor.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-19 08:48:34 -04:00
|
|
|
ApplyClassificationSoftmax();
|
|
|
|
|
double p0 = TempData.At(0), p1 = TempData.At(1), p2 = TempData.At(2);
|
|
|
|
|
if(!MathIsValidNumber(p0) || !MathIsValidNumber(p1) || !MathIsValidNumber(p2))
|
|
|
|
|
return m_dispValid;
|
|
|
|
|
m_dispProbs[0] = p0; // Buy
|
|
|
|
|
m_dispProbs[1] = p1; // Sell
|
|
|
|
|
m_dispProbs[2] = p2; // Neutral
|
|
|
|
|
m_dispSignal = AdjustedSignalFromSoftmax();
|
|
|
|
|
}
|
|
|
|
|
m_dispEra = (long)m_eraCount;
|
|
|
|
|
m_dispValid = true;
|
|
|
|
|
return true;
|
|
|
|
|
}
|
refactor: split CExpertSignalAIBase implementation by responsibility
ExpertSignalAIBase.mqh was 8216 lines: the class declaration followed by 87
method bodies covering training, labelling, feature extraction, persistence,
chart drawing, online learning, the GA auto-tuner and inference, all in one
file. Train() alone is 1492 lines; a change to arrow drawing meant scrolling
past the era loop.
Moved the bodies into Expert\AIBase\, included at the bottom of the original
after the class declaration:
Training.mqh 1607 era loop, plateau ladder, checkpoint select, deploy
Features.mqh 1093 indicator creation + per-bar input feature vector
ChartUI.mqh 634 arrows, arrow persistence, status panel, cleanup
Persistence.mqh 492 .stats/.cfg sidecars, CPU-inference validation, copy
OnlineLearning.mqh 461 live continual learning, EMA shadow, OOS simulator
Labels.mqh 309 ZigZag pivot labels, async label-cache prebuild
AutoTune.mqh 275 genetic tuner (population, crossover, halving)
Inference.mqh 235 softmax, prior calibration, class priors
ExpertSignalAIBase.mqh 8216 -> 3131 (declaration + topology build only)
This is a pure relocation - verified mechanically, not by eye: HEAD's file
reconstructed from the eight partials plus the surviving remainder is
byte-identical to HEAD, span for span (scratchpad verify_split.py). No
declaration moved, no signature changed, no code rewritten, so behaviour is
unchanged by construction.
Compiles 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 00:42:45 -04:00
|
|
|
#endif // WARRIOR_AIBASE_INFERENCE_MQH
|