Warrior_EA/Expert/AIBase/Training.mqh

1663 lines
113 KiB
MQL5
Raw Permalink Normal View History

feat(ai): logit-adjusted loss, replacing oversampling and the post-hoc prior Menon et al. 2021 (ICLR), "Long-tail learning via logit adjustment": add tau*log(prior_c) to each class logit inside the training gradient. Softmax CE on adjusted logits is consistent for BALANCED error - the metric checkpoint selection already ranks on - so the loss and the deploy decision finally optimize the same thing. The engine already computed a true softmax + categorical-CE gradient and wrote it over the per-neuron sigmoid delta, so this is an offset added to three logits in the two places that gradient is built (backProp scalar path and backPropOCL). No backend, kernel or DLL change; the forward pass and every inference path are untouched, which is the point - the network learns to absorb the offset, so its raw argmax becomes the balanced-optimal decision with nothing applied at inference. Replaces rather than stacks. Minority replay is disabled while this is on, and the post-hoc inference prior is forced off. Stacking is not a theoretical worry: simulated on the measured 1118/1119/34298 distribution in the weak-signal regime, plain CE collapses to Neutral (33.4% balanced, Buy 0%), replay reaches 48.1%, logit adjustment 50.9% with better balance - and BOTH together score 45.4% with Neutral recall at 0%, worse than either alone. Buda et al. 2018 predicts exactly that. Motivation from the six-chart run: every topology took one direction to ~50% recall and abandoned the other, the direction chosen arbitrarily (the batch-norm control went Buy 1% / Sell 42%, the inverse of the other five). One era in 1,301 cleared the per-class recall floor. Fingerprinted conditionally, so the converged 60.7% models on disk keep their filenames and stay loadable as the fallback. Both builds compile 0 errors, 0 warnings. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 19:05:14 -04:00
//+------------------------------------------------------------------+
refactor: split CExpertSignalAIBase implementation by responsibility ExpertSignalAIBase.mqh was 8216 lines: the class declaration followed by 87 method bodies covering training, labelling, feature extraction, persistence, chart drawing, online learning, the GA auto-tuner and inference, all in one file. Train() alone is 1492 lines; a change to arrow drawing meant scrolling past the era loop. Moved the bodies into Expert\AIBase\, included at the bottom of the original after the class declaration: Training.mqh 1607 era loop, plateau ladder, checkpoint select, deploy Features.mqh 1093 indicator creation + per-bar input feature vector ChartUI.mqh 634 arrows, arrow persistence, status panel, cleanup Persistence.mqh 492 .stats/.cfg sidecars, CPU-inference validation, copy OnlineLearning.mqh 461 live continual learning, EMA shadow, OOS simulator Labels.mqh 309 ZigZag pivot labels, async label-cache prebuild AutoTune.mqh 275 genetic tuner (population, crossover, halving) Inference.mqh 235 softmax, prior calibration, class priors ExpertSignalAIBase.mqh 8216 -> 3131 (declaration + topology build only) This is a pure relocation - verified mechanically, not by eye: HEAD's file reconstructed from the eight partials plus the surviving remainder is byte-identical to HEAD, span for span (scratchpad verify_split.py). No declaration moved, no signature changed, no code rewritten, so behaviour is unchanged by construction. Compiles 0 errors, 0 warnings. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 00:42:45 -04:00
//| Warrior_EA |
//| AnimateDread |
//| |
//| Era loop, plateau ladder, checkpoint selection, deploy/finalise.|
//| |
//| PARTIAL IMPLEMENTATION FILE - not standalone. |
//| This holds CExpertSignalAIBase method BODIES only. The class |
//| declaration lives in Expert\ExpertSignalAIBase.mqh, which |
//| #includes this file at the bottom, after the declaration. Do not |
//| include it anywhere else and do not compile it on its own. |
//| |
//| Split out purely to make the 8216-line original navigable; the |
//| code inside was moved verbatim, not rewritten. |
//+------------------------------------------------------------------+
#ifndef WARRIOR_AIBASE_TRAINING_MQH
#define WARRIOR_AIBASE_TRAINING_MQH
//+------------------------------------------------------------------+
//| Training and Signal Methods |
//+------------------------------------------------------------------+
void CExpertSignalAIBase::Train(datetime StartTrainBar = 0)
{
const int STABILITY_WINDOW = 3; // consecutive eras the OOS accuracy must hold steady for
const double STABILITY_TOLERANCE = 2.0; // max spread (percentage points) across that window
// Max wall-clock work per call before yielding - see m_trainRunActive's declaration comment for
// why chunking exists at all. TuneIndicatorsAndTrain()/Train() only run ONCE per dispatched
// "New Bar" custom chart event (see OnChartEventHandler - id 1001 calls it exactly once, then
// clears bEventStudy so ScheduleTrainingIfNeeded() can arm the next one), so the real throughput
// ceiling in practice is however fast MT5 itself pumps/dispatches that custom event - NOT this
// constant. Raising the OnTimer interval (5s->250ms) had ~zero effect for exactly that reason:
// ticks/chart events were already redispatching far more often than the timer alone would. Since
// per-event dispatch overhead is roughly fixed, doing more compute per event (fewer, larger
// chunks) cuts wall-clock training time roughly in proportion, but MT5 has only this one thread -
// the panel/chart can only respond to input in the gap between chunks, so 500ms made it feel
// unresponsive unless clicks landed in that narrow window. Lowered back to 120ms to keep the UI
// reactive. Raised to 200ms 2026-07-26 (throughput became the bigger complaint, as flagged above) -
// a deliberate middle ground between the reactive-but-slow 120ms and the previously-rejected 500ms,
// not a return to that. Watch panel drag/click feel after this change; back off toward 120ms if it
// regresses, or raise further only in small steps if it doesn't.
const uint TRAIN_TIME_BUDGET_MS = 200;
//---
//--- Never block the calling thread while paused/stopped - just decline this call (or finalize a
//--- run that just got stopped) and let the next scheduled call check again, so Pause/Resume/Stop
//--- and everything else on the control panel stays responsive instead of Sleep()-ing the one
//--- MQL5 thread this chart has.
if(m_trainingPaused && !IsStopped() && !m_trainingStopRequested)
return;
bool stop = IsStopped() || m_trainingStopRequested;
if(stop)
{
if(m_trainRunActive)
FinalizeTrainRun();
if(m_simOosRunActive)
{
delete m_simOosNet;
m_simOosNet = NULL;
m_simOosRunActive = false;
}
return;
}
//--- Evaluation-only continual-learning OOS simulation walk in progress (see StartOosContinualSimulation):
//--- give it exclusive occupancy of this call, same chunked budget as the real era loop below, so a
//--- large OOS window can't freeze the UI in one shot. While it's active no real-training
//--- ResizeBuffers()/RefreshData() runs, so the price/ATR/time buffers it reads stay frozen for its
//--- whole walk - it never has to worry about the label cache's shifting-index invalidation below.
if(m_simOosRunActive)
{
AdvanceOosSimulationChunk();
return;
}
//--- Eager label-cache pre-build in progress (see StartLabelCachePrebuild/AdvanceLabelCachePrebuild) -
//--- same exclusive-occupancy/chunking treatment as the OOS simulation walk above, so it can't freeze
//--- the UI on a large study window either. m_trainRunActive stays false for its whole duration, so
//--- once it completes, Train() falls through to the normal !m_trainRunActive setup below and era 0
//--- starts with the oversampling ratio it just seeded.
if(m_labelPrebuildActive)
{
AdvanceLabelCachePrebuild();
return;
}
if(!m_trainRunActive)
{
//--- Wait (briefly, bounded, non-blocking across calls) for the terminal to finish syncing this
//--- symbol/period's history from the broker before computing the training window. Bars(symbol,
//--- period) - the hard cap on how many bars the era loop below will ever process - reflects
//--- whatever's synced SO FAR, not necessarily the true total; starting before sync completes
//--- would let that cap (and therefore the "Bar X of Y" progress display) silently grow between
//--- eras as more history trickles in.
if(!SeriesInfoInteger(m_symbol.Name(), PERIOD_CURRENT, SERIES_SYNCHRONIZED))
{
uint syncNowTick = GetTickCount();
if(m_syncWaitStartTick == 0)
m_syncWaitStartTick = syncNowTick;
if(syncNowTick - m_syncWaitStartTick < 5000)
return; // retry on the next scheduled call instead of blocking here
Print(ID + ": WARNING - history for " + m_symbol.Name() + " " + EnumToString(PERIOD_CURRENT) + " did not finish syncing after 5s; training window may still grow as more history arrives");
}
m_syncWaitStartTick = 0;
//--- 3 no-op passes before the era loop ever runs for a fresh start (see m_warmupPassesRemaining's
//--- declaration comment) - each is its own separately-scheduled Train() call (this whole method
//--- just returns, deferring to the next "New Bar"/timer-driven call), giving MT5's history sync
//--- several real, wall-clock-separated chances to settle on top of the 5s soft wait just above,
//--- before training commits to a bar count and starts populating the label cache below.
if(m_warmupPassesRemaining > 0)
{
m_warmupPassesRemaining--;
PrintVerbose(ID + ": warm-up pass " + IntegerToString(3 - m_warmupPassesRemaining) + " of 3 (letting history sync settle before training starts)");
return;
}
MqlDateTime start_time;
TimeCurrent(start_time);
start_time.year -= m_studyPeriod;
if(start_time.year <= 0 || start_time.year < m_minTrainYear)
start_time.year = m_minTrainYear;
datetime st_time = StructToTime(start_time);
dtStudied = MathMax(StartTrainBar, st_time);
//--- Clamp the requested training window to what actually exists: a StudyPeriod (or MinTrainYear)
//--- reaching further back than the symbol's real history (e.g. broker feed only starts 2019, but
//--- StudyPeriod=20 asks for 20 years back from today) must not change how much data training
//--- thinks it has to cover - only the genuinely available bars ever get processed regardless, so
//--- clamping here just keeps the displayed window/progress honest instead of implying a longer
//--- span than actually exists.
datetime firstAvailableBar = (datetime)SeriesInfoInteger(m_symbol.Name(), PERIOD_CURRENT, SERIES_FIRSTDATE);
if(firstAvailableBar > 0 && dtStudied < firstAvailableBar)
dtStudied = firstAvailableBar;
//--- OOS-based objective + stability tracking: training only "converges" once the objective is
//--- met AND OOS accuracy has held inside a tight band for the last few eras, so a single lucky
//--- era can't get locked in as the final model. The best-scoring era's weights are checkpointed
//--- to an agent-local scratch file (not FILE_COMMON) and restored at the end - this works inside
//--- the tester too, unlike Net.Save()/Load() which are disabled there.
m_oosWindow.Clear();
m_bestOosForecast = -1;
m_bestBalancedOos = -1;
m_bestPassedRecall = false;
m_haveOosCheckpoint = false;
m_oosStable = false;
m_objectiveMet = false;
m_erasSinceCooldown = 0;
m_eraResumePending = false;
//--- Plateau ladder starts fresh with this run, and the annealed gamma resets to the CONFIGURED
//--- value (see m_focalGammaRuntime) so each run re-walks the schedule from its own starting point.
m_erasSinceBestBalanced = 0;
m_plateauStage = 0;
m_focalGammaRuntime = m_focalGamma;
//--- One-time eager pre-scan for a fresh start (see m_labelCachePrebuilt's declaration comment) -
//--- kick it off and defer era 0 until it's done, so era 0 can start with a real class-balance
//--- oversampling ratio instead of the reps=1 fallback. Routed via the m_labelPrebuildActive gate
//--- above on every subsequent call until it completes.
if(!m_labelCachePrebuilt)
{
StartLabelCachePrebuild();
return;
}
m_trainRunActive = true;
}
int bars, totalIter, oosCutoff, i;
bool add_loop;
if(!m_eraResumePending)
{
int barsNow = (int)MathMin(Bars(m_symbol.Name(), PERIOD_CURRENT, dtStudied, TimeCurrent()) + m_historyBars, Bars(m_symbol.Name(), PERIOD_CURRENT));
if(!ResizeBuffers(barsNow) || !RefreshData())
{
FinalizeTrainRun();
return;
}
bars = barsNow;
add_loop = false;
//--- Label/feature cache invalidation: MQL5 timeseries indices are always relative to "now"
//--- (index 0 = current bar), so every new closed candle shifts every older bar's index - a
//--- cache keyed by index would silently misalign the moment that happens. See
//--- EnsureBarCachesCapacity() for why `bars` + m_Time.GetData(0) are the correct/sufficient
//--- invalidation keys.
//--- When a wipe happens MID-RUN (a new candle closed while training was still going - e.g. the
//--- market reopening after the weekend), the label cache comes back empty and the lazy per-bar
//--- fallback (ComputeLabelForBar) labels everything Neutral by design (recent pivots are
//--- unconfirmable) - so continuing on a wiped cache silently turns the REST OF THE RUN into
//--- training AND scoring against an all-Neutral world. Observed 2026-07-19: eras 18-20 started
//--- right after the Sunday session open - IS error collapsed 0.44->0.22, OOS "accuracy" soared
//--- to 84.9% with Buy/Sell recall n/a and era time halved, the convergence machinery happily
//--- rewarding all-Neutral predictions on 100%-Neutral relabeled truth. Re-arm the same chunked
//--- eager prebuild that seeded era 0 and defer this era until it completes; its completion
//--- re-seeds the oversampling tallies (m_prebuildSeedPending) so the next era's reps reflect
//--- the freshly relabeled window.
if(EnsureBarCachesCapacity(bars) && m_labelCachePrebuilt)
{
StartLabelCachePrebuild();
return;
}
//--- freeze the just-finished era's true class totals for this new era's oversampling ratio (see
//--- m_prevEraTrueBuyCount's declaration comment) before resetting the live counters below - EXCEPT
//--- right after StartLabelCachePrebuild()/AdvanceLabelCachePrebuild() seeded them for era 0: the
//--- live m_trueBuyCount/Sell/Neutral tally is still all-zero at that point (nothing trained yet),
//--- so copying it here would silently stomp the real upfront tally right back to reps=1.
if(m_prebuildSeedPending)
m_prebuildSeedPending = false;
else
{
m_prevEraTrueBuyCount = m_trueBuyCount;
m_prevEraTrueSellCount = m_trueSellCount;
m_prevEraTrueNeutralCount = m_trueNeutralCount;
}
//--- Natural class base rates for the live logit-adjusted decision (see AdjustedSignalFromSoftmax):
//--- derived from the same just-finished-era true class totals the oversampling ratio uses, so live
//--- calibrates to exactly the distribution the model was measured against. Both branches above
//--- leave m_prevEraTrue* holding the freshest real tally (prebuild-seeded on era 0, copied here
//--- otherwise), so updating from them here covers both paths. Skipped while scoring a throwaway GA
//--- candidate (m_evalMode) so the search can't contaminate the deployed calibration - the single
//--- final retrain (m_evalMode=false) re-measures it cleanly for the model that actually ships.
if(!m_evalMode)
UpdateClassPriors(m_prevEraTrueBuyCount, m_prevEraTrueSellCount, m_prevEraTrueNeutralCount);
feat(ai): logit-adjusted loss, replacing oversampling and the post-hoc prior Menon et al. 2021 (ICLR), "Long-tail learning via logit adjustment": add tau*log(prior_c) to each class logit inside the training gradient. Softmax CE on adjusted logits is consistent for BALANCED error - the metric checkpoint selection already ranks on - so the loss and the deploy decision finally optimize the same thing. The engine already computed a true softmax + categorical-CE gradient and wrote it over the per-neuron sigmoid delta, so this is an offset added to three logits in the two places that gradient is built (backProp scalar path and backPropOCL). No backend, kernel or DLL change; the forward pass and every inference path are untouched, which is the point - the network learns to absorb the offset, so its raw argmax becomes the balanced-optimal decision with nothing applied at inference. Replaces rather than stacks. Minority replay is disabled while this is on, and the post-hoc inference prior is forced off. Stacking is not a theoretical worry: simulated on the measured 1118/1119/34298 distribution in the weak-signal regime, plain CE collapses to Neutral (33.4% balanced, Buy 0%), replay reaches 48.1%, logit adjustment 50.9% with better balance - and BOTH together score 45.4% with Neutral recall at 0%, worse than either alone. Buda et al. 2018 predicts exactly that. Motivation from the six-chart run: every topology took one direction to ~50% recall and abandoned the other, the direction chosen arbitrarily (the batch-norm control went Buy 1% / Sell 42%, the inverse of the other five). One era in 1,301 cleared the per-class recall floor. Fingerprinted conditionally, so the converged 60.7% models on disk keep their filenames and stay loadable as the fallback. Both builds compile 0 errors, 0 warnings. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 19:05:14 -04:00
//--- Re-install the training-time logit offsets from the priors just measured, so this
//--- era's gradient tracks the distribution the era is scored against. Runs in eval mode
//--- too: a GA candidate must train under the same loss as the real run or its score
//--- means nothing - only the PERSISTED calibration is withheld from eval mode.
ApplyLogitAdjustment();
refactor: split CExpertSignalAIBase implementation by responsibility ExpertSignalAIBase.mqh was 8216 lines: the class declaration followed by 87 method bodies covering training, labelling, feature extraction, persistence, chart drawing, online learning, the GA auto-tuner and inference, all in one file. Train() alone is 1492 lines; a change to arrow drawing meant scrolling past the era loop. Moved the bodies into Expert\AIBase\, included at the bottom of the original after the class declaration: Training.mqh 1607 era loop, plateau ladder, checkpoint select, deploy Features.mqh 1093 indicator creation + per-bar input feature vector ChartUI.mqh 634 arrows, arrow persistence, status panel, cleanup Persistence.mqh 492 .stats/.cfg sidecars, CPU-inference validation, copy OnlineLearning.mqh 461 live continual learning, EMA shadow, OOS simulator Labels.mqh 309 ZigZag pivot labels, async label-cache prebuild AutoTune.mqh 275 genetic tuner (population, crossover, halving) Inference.mqh 235 softmax, prior calibration, class priors ExpertSignalAIBase.mqh 8216 -> 3131 (declaration + topology build only) This is a pure relocation - verified mechanically, not by eye: HEAD's file reconstructed from the eight partials plus the surviving remainder is byte-identical to HEAD, span for span (scratchpad verify_split.py). No declaration moved, no signature changed, no code rewritten, so behaviour is unchanged by construction. Compiles 0 errors, 0 warnings. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 00:42:45 -04:00
m_countBuySignals = 0;
m_countSellSignals = 0;
m_countNeutralSignals = 0;
m_trueBuyCount = 0;
m_trueSellCount = 0;
m_trueNeutralCount = 0;
m_oosBuyHits = 0;
m_oosBuyTotal = 0;
m_oosSellHits = 0;
m_oosSellTotal = 0;
m_oosNeutralHits = 0;
m_oosNeutralTotal = 0;
m_oosBuyPredicted = 0;
m_oosBuyPredictedHits = 0;
m_oosSellPredicted = 0;
m_oosSellPredictedHits = 0;
m_oosNeutralPredicted = 0;
m_oosNeutralPredictedHits = 0;
m_oosBuyFired = 0;
m_oosBuyFiredHits = 0;
m_oosSellFired = 0;
m_oosSellFiredHits = 0;
m_oosConfidenceSum = 0;
// Nearest-to-present slice of this era's bars is held out as OOS and never backprop'd on;
// the rest (older bars) is the IS/training slice.
totalIter = (int)MathMax(bars - MathMax(m_historyBars, 0), 0);
oosCutoff = (int)(MathMax(0, MathMin(100, m_oosSplitPct)) / 100.0 * totalIter);
i = (int)(bars - MathMax(m_historyBars, 0) - 1);
//--- Fresh era: reset pass 2's shuffled-backprop queue (see m_isTrainQueue's declaration
//--- comment). Preallocated to a parity-shaped ESTIMATE, not a hard worst case: at full parity
//--- all 3 classes replicate to ~the majority count, so the queue lands near 3x totalIter -
//--- 4x covers that plus label drift. The queueing block below grows the arrays on demand if an
//--- era ever exceeds the estimate (it used to silently DROP overflow instead - harmless at the
//--- old totalIter*cap sizing, which could never fill, but real data loss now that the measured
//--- ratio, not a small fixed cap, decides the replica count).
ArrayResize(m_isTrainQueue, totalIter * 4);
ArrayResize(m_isTrainQueueWeightScale, totalIter * 4);
ArrayResize(m_isTrainQueuePrimary, totalIter * 4);
m_isTrainQueueCount = 0;
m_isTrainCursor = 0;
m_isPass2Active = false;
m_isPass2Done = false;
m_isPass3Active = false;
//--- Fresh per-era predicted-signal cache for the end-of-era NMS sweep (see PruneDirectionalClusters).
//--- -2 = "not scored this era" so stale bars from a longer prior era can't draw phantom arrows.
if(m_signalClusterWindow > 0)
{
ArrayResize(m_arrowSignalCache, bars);
ArrayInitialize(m_arrowSignalCache, -2.0);
}
}
else
{
//--- resuming a chunk that yielded mid-bar-loop last call - pick up exactly where it left off
bars = m_resumeBars;
totalIter = m_resumeTotalIter;
oosCutoff = m_resumeOosCutoff;
add_loop = m_resumeAddLoop;
i = m_resumeBarIndex;
m_eraResumePending = false;
}
// Restore this model's own learning-rate trajectory into the shared global right before this
// chunk's backProp() calls touch it - see m_modelEta's declaration comment.
eta = m_modelEta;
uint chunkStartTick = GetTickCount();
// Iterate over the bars - skipped entirely when resuming straight into pass 2, OR when resuming
// into a still-unfinished pass 3 (see m_isPass2Done's declaration comment for why checking
// m_isPass2Active alone isn't enough to detect the latter case): pass 1 already fully completed
// in an earlier call either way.
if(!m_isPass2Active && !m_isPass2Done)
{
for(; i >= 0 && !stop; i--)
{
//--- Build THIS bar's own feature window and feed it forward BEFORE checking/training against
//--- its label - see r's declaration comment below for why the window must end AT bar i, and
//--- why this must run before the label-check block rather than after: the label check needs
//--- this bar's own freshly-computed prediction, not the previous iteration's (see windowOk).
TempData.Clear();
TempData.Reserve((int)m_historyBars * m_neuronsCount);
//--- Window ends AT (includes) bar i itself, extending m_historyBars bars into the past - i.e.
//--- "everything known as of this bar's close." Predicting label(i) - "was THIS bar the
//--- reversal" - from a window that stops short of bar i itself would blind the model to the
//--- most recent price action, which is exactly the information a reversal call most depends
//--- on. Must match RefreshLatestSignal()'s window exactly (r=i there too, i=0), since that's
//--- what actually queries the deployed model live - training on a different window than what
//--- gets queried at inference time would teach the wrong task entirely.
int r = i;
bool windowOk = false;
double displayNeuron0 = 0, displayNeuron1 = 0, displayNeuron2 = 0;
if(r <= bars)
{
for(int b = 0; b < (int)m_historyBars; b++)
{
int bar_t = r + b;
if(!BufferTempData(bar_t))
break;
}
if(TempData.Total() >= (int)m_historyBars * m_neuronsCount)
{
windowOk = true;
add_loop = true;
}
}
//--- Determine label/queue-eligibility BEFORE running any feedForward this bar - see
//--- wouldQueue's use below for why. Mirrors the label-check condition this block used to
//--- gate on (moved earlier, unchanged).
bool haveLabel = false, buy = false, sell = false, wouldQueue = false;
if(windowOk && i < (int)(bars - MathMax(m_historyBars, 0) - 1) && i > 1 && m_Time.GetData(i) > dtStudied
&& (m_outputNeuronsCount == 1 || m_outputNeuronsCount == 3))
{
//--- The fractal/swing-confirmation/trend-context label at now-relative index i only depends
//--- on price/ATR history, never on model state, so it's identical every era until a new bar
//--- closes and shifts the index frame (see the cache invalidation check above) - cache it
//--- rather than recomputing from scratch every single era. Usually already populated by
//--- AdvanceLabelCachePrebuild() before era 0 ever starts - this is just a lazy fallback for
//--- any index it didn't cover (e.g. bars/window drifted between prebuild and era 0's start).
if(m_labelCacheHasValue[i])
{
buy = m_labelCacheBuy[i];
sell = m_labelCacheSell[i];
}
else
{
ComputeLabelForBar(i, bars, buy, sell);
m_labelCacheBuy[i] = buy;
m_labelCacheSell[i] = sell;
m_labelCacheHasValue[i] = true;
}
haveLabel = true;
bool isOOS = (i < oosCutoff);
// Embargo: a ZigZag pivot at bar `idx` is only confirmed once m_swingConfirmationBars
// MORE (more-recent) bars have closed after it (see AdvanceZigZagLabelState()'s
// declaration comment), further widened by +-LABEL_WINDOW_BARS. An IS bar within that
// distance of the OOS boundary can therefore carry a label that was only knowable using
// price action from inside the held-out OOS window - purge that narrow band from
// backprop entirely instead of training on it as ordinary IS.
int embargoBars = m_swingConfirmationBars + LABEL_WINDOW_BARS;
bool isEmbargoed = (!isOOS && i < oosCutoff + embargoBars);
wouldQueue = (!isOOS && !isEmbargoed);
}
//--- Only run this bar's feedForward (and the display/count/chart-draw work that depends on
//--- it) when pass 2 ISN'T about to redo it anyway. A queued bar gets a completely fresh
//--- feedForward moments later in pass 2 (see m_isTrainQueue's declaration comment - other
//--- queued bars ahead of it in the shuffled order may already have updated weights, so pass
//--- 2 can't reuse this scan's result even if it wanted to) - running it here too was one
//--- full forward pass per training sample thrown straight in the trash every era, on top of
//--- the one pass 2 actually needs. The book's SGD (references\neuronetworksbook.pdf, section
//--- 1.4) is one forward+backward pass per training sample, not two.
if(windowOk && !wouldQueue)
{
Net.feedForward(TempData);
Net.getResults(TempData);
if(m_outputNeuronsCount == 1)
dPrevSignal = TempData[0];
else
if(m_outputNeuronsCount == 3)
dPrevSignal = ApplyClassificationSoftmax();
//--- Snapshot the just-computed neuron output(s) for the status label display below, before
//--- the label-check block clears/refills TempData with the target label (Step A always
//--- runs after this point now) - reading TempData directly for display after that would
//--- show the TRUE LABEL of the bar just trained on, not the network's own prediction.
if(TempData.Total() > 0)
displayNeuron0 = TempData[0];
if(TempData.Total() > 1)
displayNeuron1 = TempData[1];
if(TempData.Total() > 2)
displayNeuron2 = TempData[2];
switch(DoubleToSignal(dPrevSignal))
{
case Buy:
m_countBuySignals++;
break;
case Sell:
m_countSellSignals++;
break;
default:
m_countNeutralSignals++;
break;
}
m_lastBarTime = m_Time.GetData(i);
if(i > 0)
{
// NMS on: record only - the era-end sweep is the SOLE renderer, so no raw (un-
// declustered) arrow is ever drawn mid-era. NMS off: draw inline as before.
if(m_signalClusterWindow > 0)
{
if(i < ArraySize(m_arrowSignalCache))
m_arrowSignalCache[i] = dPrevSignal;
}
else
if(DoubleToSignal(dPrevSignal) == Neutral)
DeleteObject(m_lastBarTime);
else
DrawObject(m_lastBarTime, dPrevSignal, m_High.GetData(i), m_Low.GetData(i));
}
UpdateTrainingStatusLabel(
StringFormat("Bar %d of %d -> %.2f%% (scan)", bars - i + 1, bars, (double)(bars - i + 1.0) / bars * 100),
displayNeuron0, displayNeuron1, displayNeuron2, dPrevSignal);
}
if(haveLabel)
{
// True label as an ENUM_SIGNAL, derived directly from the buy/sell bools - not read
// back from TempData, which no longer holds a target at this point at all (see above).
ENUM_SIGNAL trueSignal = buy ? Buy : (sell ? Sell : Neutral);
// Track the true class distribution this era (used below to weight IS oversampling,
// and surfaced in the status label text alongside the predicted-class counts)
switch(trueSignal)
{
case Buy:
m_trueBuyCount++;
break;
case Sell:
m_trueSellCount++;
break;
default:
m_trueNeutralCount++;
break;
}
// OOS scoring used to happen right here, against whatever weights this bar's earlier
// feedForward (this pass) happened to be using - which for era 0 is the network's
// still-untrained cold-start state (100% Neutral - see the output-layer bias seed's
// declaration comment), and for every later era is last era's END-of-training state,
// never THIS era's. That silently gave every era's OOS score a full one-era lag behind
// its own training, and made era 0's OOS score meaningless by construction. OOS scoring
// now happens in its own pass (see m_isPass3Active's declaration comment), AFTER pass 2
// has actually trained on this era's IS data, against a fresh feedForward on each OOS
// bar rather than this scan's now-stale one.
if(wouldQueue)
{
// Queue this bar for pass 2's shuffled backProp instead of training on it here,
// immediately, in strict chronological order - see m_isTrainQueue's declaration
// comment for the full rationale. Class-balance weighting and focal-loss
// modulation (m_maxClassSampleWeight/m_focalGamma), the predicted-signal counts,
// the chart-marker draw, and the dForecast/dUndefine IS-accuracy update are all
// computed in pass 2 instead now, against that bar's own freshly-recomputed
// confidence - see the matching block right after pass 2's Net.feedForward() call.
//
// Data-level oversampling: a minority-class (Buy/Sell) bar gets queued more than once
// (capped at MAX_OVERSAMPLE_REPLICAS), using the same previous-era class ratio
// m_maxClassSampleWeight's loss weight is derived from. Reason this exists alongside
// that loss weight rather than instead of it: sampleWeight only scales the gradient
// AFTER it's computed, and Adam's update (AI\Network.mqh's UpdateWeightsAdam, lt *
// mt/sqrt(vt)) is deliberately near-invariant to a constant rescaling of the gradient
// (Kingma & Ba 2015) - mt and vt both move with the weight, so the ratio largely
// cancels it back out, which is why a purely loss-weighted 5x+ imbalance was observed
// in practice to plateau at a full Neutral-only collapse for 100+ eras instead of
// correcting (era 20-115 of one run stuck at OOS accuracy ~73.6%, Buy/Sell recall 0%).
// Duplicating the bar itself changes how often Adam's moment estimates see a Buy/Sell
// gradient at all, which the mt/sqrt(vt) normalization can't cancel out the same way.
//
// repCount below is driven by the RAW previous-era class ratio (uncapped by
// m_maxClassSampleWeight) - the actual imbalance is what needs correcting, not a
// pre-shrunk version of it - bounded only by MAX_OVERSAMPLE_REPLICAS for the
// momentum-correlation safety reason above. perOccurrenceScale is pinned at 1.0:
// repCount's duplication IS the whole correction now, not one half of a split budget.
// Two earlier versions of this both stacked a SECOND multiplicative correction on top
// of repCount and both collapsed, in opposite directions: v1 let pass 2 independently
// reach for its own full m_maxClassSampleWeight on top of an uncapped repCount (up to
// 3 x 1.5 = 4.5x total), which overshot into a Buy-only (or Sell-only) collapse -
// OOS accuracy dropping to ~10%, IS error climbing 0.37->0.57 over 4 eras, the same
// Adam-overshoot signature documented at AI\Network.mqh's MAX_WEIGHT_DELTA comment,
// just from a different cause. v2 tried to "coordinate" the split by capping the ratio
// at m_maxClassSampleWeight BEFORE dividing it between repCount and perOccurrenceScale
// - but repCount x perOccurrenceScale is then mathematically forced to equal that same
// <=1.5x ceiling no matter how it's split, i.e. IDENTICAL total correction magnitude to
// the pure loss-weighting that already failed to escape the original Neutral collapse
// (era 20-115, OOS ~73.6%, Buy/Sell recall 0%) - so it just silently undershot back
// into the same failure it was meant to fix. Per Buda, Maki & Mazurowski 2018 (Neural
// Networks), combining data-level oversampling with cost-sensitive loss reweighting on
// the SAME class-balance axis doesn't reliably help and can hurt; oversampling alone is
// competitive or better. v3 let repCount alone carry the correction but capped
// MAX_OVERSAMPLE_REPLICAS at 3 - only a ~1.75-1.8x residual imbalance against the
// observed ~5.2-5.3x ratio - and still collapsed the same way (OOS accuracy climbing
// toward the ~73.6% Neutral base rate while Buy/Sell recall stayed 0% for the first 6
// eras straight, 2026-07-18). The double-counting removal was correct; the cap just
// wasn't raised far enough to reach it. MAX_OVERSAMPLE_REPLICAS is now 5, letting
// repCount reach much closer to full class parity for this ratio, per Buda et al.'s
// finding that oversampling to parity is safe.
//
// Queued BEFORE pass 2's Fisher-Yates shuffle (which runs on the whole queue, repeats
// included), so duplicates land scattered through the era's shuffled replay order, not
// back-to-back - avoiding the correlated-momentum Adam-overshoot bug a cruder
// back-to-back oversampling replay caused historically (see AI\Network.mqh's
// MAX_WEIGHT_DELTA comment).
double rawRatio = 1.0;
if(trueSignal != Neutral)
{
int prevMaxCount = (int)MathMax(m_prevEraTrueBuyCount, MathMax(m_prevEraTrueSellCount, m_prevEraTrueNeutralCount));
int prevTrueCount = (trueSignal == Buy) ? m_prevEraTrueBuyCount : m_prevEraTrueSellCount;
if(prevTrueCount > 0 && prevMaxCount > 0)
rawRatio = (double)prevMaxCount / prevTrueCount;
}
// The measured per-era ratio is the binding term; MAX_OVERSAMPLE_REPLICAS is only the
// degenerate-tally sanity ceiling - see its declaration comment. OVERSAMPLE_PARITY_
// FRACTION dials how close to full 1:1:1 parity to push (see its declaration comment) -
// <1.0 leaves Neutral proportionally heavier, trading directional recall for precision.
int repCount = 1;
feat(ai): logit-adjusted loss, replacing oversampling and the post-hoc prior Menon et al. 2021 (ICLR), "Long-tail learning via logit adjustment": add tau*log(prior_c) to each class logit inside the training gradient. Softmax CE on adjusted logits is consistent for BALANCED error - the metric checkpoint selection already ranks on - so the loss and the deploy decision finally optimize the same thing. The engine already computed a true softmax + categorical-CE gradient and wrote it over the per-neuron sigmoid delta, so this is an offset added to three logits in the two places that gradient is built (backProp scalar path and backPropOCL). No backend, kernel or DLL change; the forward pass and every inference path are untouched, which is the point - the network learns to absorb the offset, so its raw argmax becomes the balanced-optimal decision with nothing applied at inference. Replaces rather than stacks. Minority replay is disabled while this is on, and the post-hoc inference prior is forced off. Stacking is not a theoretical worry: simulated on the measured 1118/1119/34298 distribution in the weak-signal regime, plain CE collapses to Neutral (33.4% balanced, Buy 0%), replay reaches 48.1%, logit adjustment 50.9% with better balance - and BOTH together score 45.4% with Neutral recall at 0%, worse than either alone. Buda et al. 2018 predicts exactly that. Motivation from the six-chart run: every topology took one direction to ~50% recall and abandoned the other, the direction chosen arbitrarily (the batch-norm control went Buy 1% / Sell 42%, the inverse of the other five). One era in 1,301 cleared the per-class recall floor. Fingerprinted conditionally, so the converged 60.7% models on disk keep their filenames and stay loadable as the fallback. Both builds compile 0 errors, 0 warnings. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 19:05:14 -04:00
//--- Logit adjustment corrects the prior analytically in the gradient; stacking
//--- replay on top double-counts the same imbalance - the failure Buda et al. 2018
//--- warns about and the one measured here twice. Replay stays fully available as
//--- the fallback whenever the adjusted loss is off.
if(m_enableMinorityReplay && !m_useLogitAdjustedLoss && trueSignal != Neutral)
refactor: split CExpertSignalAIBase implementation by responsibility ExpertSignalAIBase.mqh was 8216 lines: the class declaration followed by 87 method bodies covering training, labelling, feature extraction, persistence, chart drawing, online learning, the GA auto-tuner and inference, all in one file. Train() alone is 1492 lines; a change to arrow drawing meant scrolling past the era loop. Moved the bodies into Expert\AIBase\, included at the bottom of the original after the class declaration: Training.mqh 1607 era loop, plateau ladder, checkpoint select, deploy Features.mqh 1093 indicator creation + per-bar input feature vector ChartUI.mqh 634 arrows, arrow persistence, status panel, cleanup Persistence.mqh 492 .stats/.cfg sidecars, CPU-inference validation, copy OnlineLearning.mqh 461 live continual learning, EMA shadow, OOS simulator Labels.mqh 309 ZigZag pivot labels, async label-cache prebuild AutoTune.mqh 275 genetic tuner (population, crossover, halving) Inference.mqh 235 softmax, prior calibration, class priors ExpertSignalAIBase.mqh 8216 -> 3131 (declaration + topology build only) This is a pure relocation - verified mechanically, not by eye: HEAD's file reconstructed from the eight partials plus the surviving remainder is byte-identical to HEAD, span for span (scratchpad verify_split.py). No declaration moved, no signature changed, no code rewritten, so behaviour is unchanged by construction. Compiles 0 errors, 0 warnings. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 00:42:45 -04:00
{
repCount = (int)MathMin(MAX_OVERSAMPLE_REPLICAS, MathMax(1, MathRound(rawRatio * m_oversampleParity)));
if(m_constrainReplay)
repCount = MathMin(repCount, 3);
}
double perOccurrenceScale = 1.0;
//--- grow on demand (reserve keeps this amortized-rare) - the prealloc above is an
//--- estimate, and dropping overflow would silently starve exactly the minority
//--- classes the replication exists to protect
if(m_isTrainQueueCount + repCount > ArraySize(m_isTrainQueue))
{
int newQueueSize = m_isTrainQueueCount + repCount;
ArrayResize(m_isTrainQueue, newQueueSize, 16384);
ArrayResize(m_isTrainQueueWeightScale, newQueueSize, 16384);
ArrayResize(m_isTrainQueuePrimary, newQueueSize, 16384);
}
for(int rep = 0; rep < repCount; rep++)
{
m_isTrainQueue[m_isTrainQueueCount] = i;
m_isTrainQueueWeightScale[m_isTrainQueueCount] = perOccurrenceScale;
//--- rep 0 is this bar's single "counts once" occurrence - see m_isTrainQueuePrimary.
//--- Every rep still trains; only the reported IS accuracy looks at this flag.
m_isTrainQueuePrimary[m_isTrainQueueCount] = (rep == 0);
m_isTrainQueueCount++;
}
}
}
stop = IsStopped() || m_trainingStopRequested;
if(!stop && i > 0 && GetTickCount() - chunkStartTick >= TRAIN_TIME_BUDGET_MS)
{
//--- yield: save exactly enough to resume this same era, mid-bar-loop, on the next call -
//--- see m_trainRunActive's declaration comment for why this must happen instead of
//--- letting one era (or the whole run) process synchronously to completion
m_resumeBars = bars;
m_resumeTotalIter = totalIter;
m_resumeOosCutoff = oosCutoff;
m_resumeAddLoop = add_loop;
m_resumeBarIndex = i - 1;
m_eraResumePending = true;
// Save this model's own learning-rate trajectory back out of the shared global before
// yielding - see m_modelEta's declaration comment.
m_modelEta = eta;
return;
}
}
} // end if(!m_isPass2Active) - pass 1
//--- Pass 2: replay the bars pass 1 queued into m_isTrainQueue for backProp, in a freshly
//--- shuffled order - see m_isTrainQueue's declaration comment for the full rationale. Runs
//--- whenever pass 1 just finished (or we resumed straight into an already-active pass 2 - see
//--- m_isPass2Active's declaration comment); skipped on a stopped run, an era with no valid window
//--- at all (add_loop still false), or - critically - a resume into a still-unfinished pass 3 (see
//--- m_isPass2Done's declaration comment): without this last check, that resume would re-shuffle
//--- and replay the ENTIRE queue again from scratch every single call.
if(!stop && add_loop && !m_isPass2Done)
{
if(!m_isPass2Active)
{
m_isPass2Active = true;
m_isTrainCursor = 0;
// 2026-07-28: a "replay-only optimizer override" was removed from here. It captured every
// neuron's optimizer and forced the whole net to SGD for the duration of pass 2, on the
// rationale that oversampled minority bars should not "exploit the same Adam-style momentum
// path as the base training pass". But pass 2 IS the base training pass - it is the only place
// Net.backProp() is called during training at all (pass 1 only feeds forward and queues) - so
// with EnableMinorityReplay on (the default) the override applied to 100% of weight updates.
// Adam's mt/vt were therefore never updated and its bias-correction step counter never
// advanced: TrainingOptimizer=ADAM was silently a no-op and the model trained purely on
// SGD+momentum at Adam's learning rate. It arrived with the DFA change set and was never part
// of any validated run. The optimizer the user selects is now the optimizer that runs.
// Fisher-Yates shuffle - a fresh random order every era, so Adam's momentum can't keep
// landing on the same short same-class label run (see LABEL_WINDOW_BARS) at the same point
// in the sequence every single era. m_isTrainQueueWeightScale is swapped in lockstep - each
// slot's stored per-occurrence weight (see the queueing block's oversampling comment) must
// stay attached to the same bar index it was computed for.
for(int sIdx = m_isTrainQueueCount - 1; sIdx > 0; sIdx--)
{
int sJ = MathRand() % (sIdx + 1);
int sTmp = m_isTrainQueue[sIdx];
m_isTrainQueue[sIdx] = m_isTrainQueue[sJ];
m_isTrainQueue[sJ] = sTmp;
double sScaleTmp = m_isTrainQueueWeightScale[sIdx];
m_isTrainQueueWeightScale[sIdx] = m_isTrainQueueWeightScale[sJ];
m_isTrainQueueWeightScale[sJ] = sScaleTmp;
//--- the primary flag must travel with its own slot too, or the "count this bar once"
//--- marker would end up attached to a different bar's occurrence - see m_isTrainQueuePrimary
bool sPrimTmp = m_isTrainQueuePrimary[sIdx];
m_isTrainQueuePrimary[sIdx] = m_isTrainQueuePrimary[sJ];
m_isTrainQueuePrimary[sJ] = sPrimTmp;
}
}
for(; m_isTrainCursor < m_isTrainQueueCount; m_isTrainCursor++)
{
int qi = m_isTrainQueue[m_isTrainCursor];
TempData.Clear();
TempData.Reserve((int)m_historyBars * m_neuronsCount);
bool qWindowOk = true;
for(int b = 0; b < (int)m_historyBars; b++)
if(!BufferTempData(qi + b))
{
qWindowOk = false;
break;
}
if(qWindowOk && TempData.Total() >= (int)m_historyBars * m_neuronsCount)
{
Net.feedForward(TempData);
Net.getResults(TempData);
// Must go through ApplyClassificationSoftmax() (3-output case) before reading the
// per-class values below - Net.getResults() returns each output neuron's own independent
// SIGMOID activation (each already in [0,1] but NOT summing to 1 across the three), not a
// true class-conditional probability distribution; ApplyClassificationSoftmax() is what
// turns that into one (and is also what pass 1/3's displayNeuron0/1/2 already go through).
double qPrevSignal = (m_outputNeuronsCount == 3) ? ApplyClassificationSoftmax() : TempData[0];
double pt0 = (TempData.Total() > 0) ? TempData[0] : 0.0;
double pt1 = (TempData.Total() > 1) ? TempData[1] : 0.0;
double pt2 = (TempData.Total() > 2) ? TempData[2] : 0.0;
bool qBuy = m_labelCacheHasValue[qi] ? m_labelCacheBuy[qi] : false;
bool qSell = m_labelCacheHasValue[qi] ? m_labelCacheSell[qi] : false;
ENUM_SIGNAL qTrueSignal = qBuy ? Buy : (qSell ? Sell : Neutral);
UpdateTrainingStatusLabel(
StringFormat("Training bar %d of %d -> %.2f%% (shuffled)", m_isTrainCursor + 1, m_isTrainQueueCount, (double)(m_isTrainCursor + 1.0) / MathMax(m_isTrainQueueCount, 1) * 100),
pt0, pt1, pt2, qPrevSignal);
//--- Predicted-signal tally, chart marker, and IS-accuracy stat that pass 1 used to compute
//--- from its own (now-removed) redundant feedForward on this same bar - see pass 1's
//--- wouldQueue comment. Uses THIS feedForward's result (the only one this bar gets), so
//--- these now reflect the model's state as of this bar's own turn in the shuffled replay
//--- (post any earlier-shuffled bar's backProp this era), not a separate pre-training
//--- snapshot - matching how a standard shuffled-epoch SGD run reports running training
//--- accuracy during the epoch rather than in a discarded pre-epoch dry run.
switch(DoubleToSignal(qPrevSignal))
{
case Buy:
m_countBuySignals++;
break;
case Sell:
m_countSellSignals++;
break;
default:
m_countNeutralSignals++;
break;
}
datetime qBarTime = m_Time.GetData(qi);
// NMS on: record only (the era-end sweep renders); off: draw inline. See pass 1's note.
if(m_signalClusterWindow > 0)
{
if(qi < ArraySize(m_arrowSignalCache))
m_arrowSignalCache[qi] = qPrevSignal;
}
else
if(DoubleToSignal(qPrevSignal) == Neutral)
DeleteObject(qBarTime);
else
DrawObject(qBarTime, qPrevSignal, m_High.GetData(qi), m_Low.GetData(qi));
bool qClassified = (DoubleToSignal(qPrevSignal) == Buy || DoubleToSignal(qPrevSignal) == Sell || DoubleToSignal(qPrevSignal) == Neutral);
if(qClassified)
{
bool isHit = (DoubleToSignal(qPrevSignal) == qTrueSignal);
if(isHit)
dForecast += (100 - dForecast) / Net.recentAverageSmoothingFactor;
else
dForecast -= dForecast / Net.recentAverageSmoothingFactor;
dUndefine -= dUndefine / Net.recentAverageSmoothingFactor;
//--- Compounded, persistent DIRECTIONAL win-rate: count only bars the model actually called
//--- Buy or Sell (a Neutral "no trade" call is neither a win nor a loss), so this tracks the
//--- accuracy of its directional signals rather than the Neutral-inflated all-class rate.
//--- Skip m_evalMode throwaways. See m_cumIsCorrect.
//--- ...and count each BAR once, not each oversampled OCCURRENCE (m_isTrainQueuePrimary):
//--- the queue duplicates minority bars up to ~21x, so counting every occurrence scored this
//--- metric over a ~58%-directional set while its OOS counterpart scored the real ~6%
//--- distribution - two numbers that look comparable, aren't, and made a healthy run read as
//--- severe overfitting. See m_isTrainQueuePrimary for the worked example.
ENUM_SIGNAL qPred = DoubleToSignal(qPrevSignal);
if(!m_evalMode && m_isTrainQueuePrimary[m_isTrainCursor] && (qPred == Buy || qPred == Sell))
{
m_cumIsTotal++;
if(isHit)
m_cumIsCorrect++;
}
}
else
if(qBuy && qSell)
dUndefine += (100 - dUndefine) / Net.recentAverageSmoothingFactor;
TempData.Clear();
if(m_outputNeuronsCount == 1)
TempData.Add(qBuy && !qSell ? 1 : !qBuy && qSell ? -1 : 0);
else
if(m_outputNeuronsCount == 3)
{
TempData.Add(qBuy ? LABEL_SMOOTH_HIGH : LABEL_SMOOTH_LOW);
TempData.Add(qSell ? LABEL_SMOOTH_HIGH : LABEL_SMOOTH_LOW);
TempData.Add((!qBuy && !qSell) ? LABEL_SMOOTH_HIGH : LABEL_SMOOTH_LOW);
}
// Class-balance component comes from m_isTrainQueueWeightScale[m_isTrainCursor] - decided
// once at queue time in pass 1 (see that block's oversampling comment). Currently always
// 1.0: repCount's duplication is the entire class-balance correction now, so this array
// carries no additional per-occurrence weight on top of it - see the pass-1 comment for
// why stacking a second correction here caused two separate collapses. Kept as a real
// per-slot value (not a literal 1.0 inline) so a future, smaller/safer supplemental weight
// can be reintroduced here without re-touching the queueing or shuffle code.
// Focal-loss modulation is still applied fresh here though - see m_focalGamma's
// declaration comment. pt is this bar's OWN just-recomputed softmax confidence in its true
// class - fresh as of THIS pass-2 step, not pass 1's now-stale snapshot, since other
// queued bars ahead of it in the shuffled order may already have updated weights.
//--- m_focalGammaRuntime, NOT m_focalGamma: the plateau ladder anneals the runtime value
//--- downward mid-run (see the PLATEAU_* constants), while the configured member stays put
//--- because it is part of this model's weights-filename fingerprint.
double qSampleWeight = m_isTrainQueueWeightScale[m_isTrainCursor];
//--- Focal-loss gamma actually applied this step. Two separate decisions, previously
//--- conflated into one condition that made focal loss silently UNREACHABLE whenever
//--- EnableMinorityReplay was off - so "plain one-pass training" also meant "no focal
//--- loss", and the two could never be isolated from each other for diagnosis:
//--- - WHETHER focal applies: purely whether a gamma was configured. It no longer
//--- depends on the replay toggle at all.
//--- - HOW STRONG it is: damped ONLY while oversampling is also running. Focal loss and
//--- minority oversampling both correct the SAME class-imbalance axis, and stacking
//--- two full-strength corrections on that axis is what caused two documented
//--- collapses (see the queueing block's oversampling comment) - Buda, Maki &
//--- Mazurowski 2018 on why combining the two is not reliably additive. With replay
//--- OFF, focal is the SOLE imbalance correction and runs at full configured gamma,
//--- which is exactly Lin et al. 2017's intended single-mechanism use.
double replayGamma = 0.0;
if(m_outputNeuronsCount == 3 && m_focalGammaRuntime > 0.0)
replayGamma = m_enableMinorityReplay
? MathMax(0.0, m_focalGammaRuntime * (m_constrainReplay ? 0.125 : 0.25))
: m_focalGammaRuntime;
if(m_outputNeuronsCount == 3 && replayGamma > 0.0)
{
double pt = (qTrueSignal == Buy) ? pt0 : (qTrueSignal == Sell) ? pt1 : pt2;
pt = MathMax(0.0, MathMin(1.0, pt));
qSampleWeight *= MathPow(1.0 - pt, replayGamma);
}
Net.backProp(TempData, qSampleWeight);
}
if(m_isTrainCursor + 1 < m_isTrainQueueCount && GetTickCount() - chunkStartTick >= TRAIN_TIME_BUDGET_MS)
{
//--- yield: save enough to resume PASS 2 mid-queue on the next call - m_isPass2Active
//--- and m_isTrainCursor (both members) carry the actual resume position; bars/oosCutoff/
//--- add_loop are stashed the same way pass 1 already does, since era-end logic just
//--- below still needs them once pass 2 finishes.
m_resumeBars = bars;
m_resumeTotalIter = totalIter;
m_resumeOosCutoff = oosCutoff;
m_resumeAddLoop = add_loop;
m_resumeBarIndex = i;
m_eraResumePending = true;
m_modelEta = eta;
return;
}
}
m_isPass2Active = false;
m_isPass2Done = true;
}
//--- Pass 3: OOS scoring, chronological, AFTER pass 2 has actually trained on this era's IS data -
//--- see m_isPass3Active's declaration comment for why this can no longer happen inline during
//--- pass 1's scan.
if(!stop && add_loop)
{
if(!m_isPass3Active)
{
m_isPass3Active = true;
m_oosScoreStartIndex = (int)MathMin(oosCutoff - 1, bars - MathMax(m_historyBars, 0) - 2);
m_oosScoreIndex = m_oosScoreStartIndex;
for(int rn = 0; rn < 3; rn++)
{
m_oosOutMin[rn] = DBL_MAX;
m_oosOutMax[rn] = -DBL_MAX;
}
m_oosOutSpreadSum = 0.0;
m_oosOutCount = 0;
}
for(; m_oosScoreIndex >= 2; m_oosScoreIndex--)
{
int oi = m_oosScoreIndex;
if(!(oi < (int)(bars - MathMax(m_historyBars, 0) - 1) && m_Time.GetData(oi) > dtStudied))
continue;
TempData.Clear();
TempData.Reserve((int)m_historyBars * m_neuronsCount);
bool oWindowOk = true;
for(int b = 0; b < (int)m_historyBars; b++)
if(!BufferTempData(oi + b))
{
oWindowOk = false;
break;
}
if(oWindowOk && TempData.Total() >= (int)m_historyBars * m_neuronsCount)
{
Net.feedForward(TempData);
Net.getResults(TempData);
// Raw output stats MUST be captured here, before ApplyClassificationSoftmax() overwrites
// TempData[0..2] in place with the softmax probabilities - see m_oosOutMin's declaration
// comment for what these feed.
if(m_outputNeuronsCount == 3 && TempData.Total() >= 3)
{
double rawHi = -DBL_MAX, rawLo = DBL_MAX;
for(int rn = 0; rn < 3; rn++)
{
double rv = TempData.At(rn);
if(rv < m_oosOutMin[rn])
m_oosOutMin[rn] = rv;
if(rv > m_oosOutMax[rn])
m_oosOutMax[rn] = rv;
rawHi = MathMax(rawHi, rv);
rawLo = MathMin(rawLo, rv);
}
m_oosOutSpreadSum += rawHi - rawLo;
m_oosOutCount++;
}
double oPrevSignal = (m_outputNeuronsCount == 3) ? ApplyClassificationSoftmax() : TempData[0];
double oDeploySignal = oPrevSignal;
if(m_outputNeuronsCount == 3)
oDeploySignal = AdjustedSignalFromSoftmax();
double oNeuron0 = (TempData.Total() > 0) ? TempData[0] : 0.0;
double oNeuron1 = (TempData.Total() > 1) ? TempData[1] : 0.0;
double oNeuron2 = (TempData.Total() > 2) ? TempData[2] : 0.0;
bool oBuy = m_labelCacheHasValue[oi] ? m_labelCacheBuy[oi] : false;
bool oSell = m_labelCacheHasValue[oi] ? m_labelCacheSell[oi] : false;
ENUM_SIGNAL oTrueSignal = oBuy ? Buy : (oSell ? Sell : Neutral);
UpdateTrainingStatusLabel(
StringFormat("Scoring OOS bar %d of %d -> %.2f%% (post-training)", m_oosScoreStartIndex - m_oosScoreIndex + 1, m_oosScoreStartIndex + 1,
(double)(m_oosScoreStartIndex - m_oosScoreIndex + 1.0) / MathMax(m_oosScoreStartIndex + 1, 1) * 100),
oNeuron0, oNeuron1, oNeuron2, oDeploySignal);
// Held-out bar: score the model's freshly-trained-this-era forecast against the actual
// outcome without learning from it - keeps the OOS accuracy an honest overfitting signal.
bool oClassified = (DoubleToSignal(oPrevSignal) == Buy || DoubleToSignal(oPrevSignal) == Sell || DoubleToSignal(oPrevSignal) == Neutral);
if(oClassified)
{
m_oosSamples++;
m_oosConfidenceSum += MathAbs(oPrevSignal);
if(dOosError < 0)
dOosError = 0;
bool hit = (DoubleToSignal(oPrevSignal) == oTrueSignal);
//--- Compounded, persistent DIRECTIONAL win-rate: count only bars the model actually called
//--- Buy or Sell (Neutral "no trade" calls aren't wins or losses), so this reflects the
//--- accuracy of its directional signals, not the Neutral-inflated all-class rate. Skip
//--- m_evalMode throwaways. See m_cumOosCorrect.
ENUM_SIGNAL oPred = DoubleToSignal(oPrevSignal);
if(!m_evalMode && (oPred == Buy || oPred == Sell))
{
m_cumOosTotal++;
if(hit)
m_cumOosCorrect++;
}
// Per-class confusion counts, used for the Buy/Sell recall convergence gate below
switch(oTrueSignal)
{
case Buy:
m_oosBuyTotal++;
if(hit)
m_oosBuyHits++;
break;
case Sell:
m_oosSellTotal++;
if(hit)
m_oosSellHits++;
break;
default:
m_oosNeutralTotal++;
if(hit)
m_oosNeutralHits++;
break;
}
// Same confusion counts keyed by what the model actually PREDICTED this bar, not the
// true label - see m_oosBuyPredicted's declaration comment for why recall alone can
// hide an over-firing class.
switch(DoubleToSignal(oPrevSignal))
{
case Buy:
m_oosBuyPredicted++;
if(hit)
m_oosBuyPredictedHits++;
break;
case Sell:
m_oosSellPredicted++;
if(hit)
m_oosSellPredictedHits++;
break;
default:
m_oosNeutralPredicted++;
if(hit)
m_oosNeutralPredictedHits++;
break;
}
// Live-decision precision: scores the bars on which the deployed EA would actually cast a
// directional vote, using the prior-corrected (logit-adjusted) posterior - see
// AdjustedSignalFromSoftmax()/RefreshLatestSignal(). The recall/argmax-precision above stay
// on the raw argmax (the model's intrinsic class separation, which the convergence gate
// needs); THIS scores what trades live, so the panel's live precision number is the
// precision a buyer gets forward. TempData still holds this bar's raw softmax probs
// (nothing overwrote them since ApplyClassificationSoftmax above), so the adjustment reads
// them directly. Neutral picks aren't counted - they cast no vote.
// No confidence-floor term any more: with the floor removed, EVERY non-Neutral adjusted
// decision casts a vote (at its tier weight), so any threshold here would score a
// different population than the one that actually votes. Whether a given vote goes on to
// OPEN a position additionally depends on Min_Vote_Open versus the AVERAGE across all
// voting filters, which this per-bar training-time scorer has no visibility of - so this
// stays the honest "would have voted, and was it right" measure rather than pretending to
// model the aggregate.
if(m_outputNeuronsCount == 3)
{
double adjSig = AdjustedSignalFromSoftmax();
ENUM_SIGNAL adjEnum = DoubleToSignal(adjSig);
if(adjEnum != Neutral)
{
bool fireHit = (adjEnum == oTrueSignal);
if(adjEnum == Buy)
{
m_oosBuyFired++;
if(fireHit)
m_oosBuyFiredHits++;
}
else
{
m_oosSellFired++;
if(fireHit)
m_oosSellFiredHits++;
}
}
}
if(hit)
{
dOosForecast += (100 - dOosForecast) / Net.recentAverageSmoothingFactor;
dOosError -= dOosError / Net.recentAverageSmoothingFactor;
}
else
{
dOosForecast -= dOosForecast / Net.recentAverageSmoothingFactor;
dOosError += (100 - dOosError) / Net.recentAverageSmoothingFactor;
}
}
// Mirror pass 1's chart annotation for this (OOS) bar, now using post-training weights
// instead of pass 1's pre-training snapshot - last write wins on the shared per-bar-time
// object, so this supersedes pass 1's earlier draw for the same bar with the correct one.
m_lastBarTime = m_Time.GetData(oi);
if(oi > 0)
{
// NMS on: record only (the era-end sweep renders); off: draw inline. See pass 1's note.
if(m_signalClusterWindow > 0)
{
if(oi < ArraySize(m_arrowSignalCache))
m_arrowSignalCache[oi] = oDeploySignal;
}
else
if(DoubleToSignal(oDeploySignal) == Neutral)
DeleteObject(m_lastBarTime);
else
DrawObject(m_lastBarTime, oDeploySignal, m_High.GetData(oi), m_Low.GetData(oi));
}
}
if(m_oosScoreIndex - 1 >= 2 && GetTickCount() - chunkStartTick >= TRAIN_TIME_BUDGET_MS)
{
//--- yield: save enough to resume PASS 3 mid-walk on the next call - m_isPass3Active and
//--- m_oosScoreIndex (both members) carry the actual resume position.
m_resumeBars = bars;
m_resumeTotalIter = totalIter;
m_resumeOosCutoff = oosCutoff;
m_resumeAddLoop = add_loop;
m_resumeBarIndex = i;
m_eraResumePending = true;
m_modelEta = eta;
return;
}
}
m_isPass3Active = false;
//--- Pass 3 done => every scored bar's prediction is now in m_arrowSignalCache. Collapse each
//--- same-direction cluster to its earliest bar so the chart shows one arrow per real turn.
PruneDirectionalClusters(bars);
}
//--- Diagnostic recall snapshot for the periodic progress log further below - populated inside
//--- the m_oosSamples>0 recall-gate block when this era actually computes it; stays -1 ("n/a"
//--- in the log) on eras that don't (era 0, or a stopped/cap-hit era).
int logBuyRecallPct = -1, logSellRecallPct = -1, logNeutralRecallPct = -1;
//--- Balanced accuracy (macro-recall) this era, surfaced in the log so the metric the checkpoint
//--- is now selected on is visible - see m_bestBalancedOos. -1 ("n/a") on eras that don't score.
int logBalancedAccPct = -1;
//--- Predicted-rate (of all OOS bars this era, how often the model called this class at all) and
//--- precision (of the calls it made, how many were right) for Buy/Sell - m_oosBuyPredicted/
//--- m_oosSellPredicted (see that member's declaration comment) were already being tracked for
//--- exactly this but never surfaced anywhere. A recall-only view can't tell "the model never once
//--- calls Sell" (predicted rate stuck at 0%) apart from "the model calls Sell plenty but always on
//--- the wrong bars" (predicted rate healthy, precision near 0%) - both show up identically as 0%
//--- Sell recall, but point at completely different problems (a suppressed/dead output vs. a
//--- miscalibrated decision boundary), so this splits them out.
int logBuyPredPct = -1, logSellPredPct = -1, logBuyPrecPct = -1, logSellPrecPct = -1;
//--- Live-fired precision (%) per direction this era - the precision on just the bars that cleared
//--- the confidence floor under the live/prior-corrected decision, i.e. what would actually trade.
int logBuyFiredPrecPct = -1, logSellFiredPrecPct = -1;
bool shouldLogProgress = false;
//--- era complete (ran out of bars) or a stop was requested mid-era
if(add_loop)
{
m_eraCount++;
m_erasSinceCooldown++;
//--- EMA shadow-weight deployment: blend the shadow a small step (SHADOW_WEIGHT_TAU) toward
//--- Net's just-updated weights, every era - see m_shadowNet's declaration comment. Must run
//--- here, inside the era loop, not just once at Train()-end: the whole point is damping the
//--- WITHIN-run oscillation (era-to-era whipsaw), which a single end-of-run blend would miss
//--- entirely.
//--- skip the deployed EMA shadow entirely while scoring a throwaway candidate (m_evalMode)
if(!m_evalMode)
{
EnsureShadowNet();
if(CheckPointer(m_shadowNet) != POINTER_INVALID)
m_shadowNet.BlendWeightsFrom(Net, SHADOW_WEIGHT_TAU);
}
//--- Status-label progress is invisible with no chart (headless/optimization runs), and even in
//--- visual mode a long training run can otherwise look "stuck" for a long time with no
//--- Journal output at all - log progress at most every ~5s (real wall-clock, not simulated
//--- time) so an operator can tell it's actively working, not hung. The actual Print() is
//--- deferred past the recall-gate block below (see logBuyRecallPct etc.) so this line can
//--- show per-class OOS recall - once OOS accuracy alone clears the target, recall is the
//--- most common thing still silently blocking convergence, and previously had no visibility
//--- outside of a regression event.
static uint lastProgressLogTick = 0;
uint nowTick = GetTickCount();
shouldLogProgress = (nowTick - lastProgressLogTick >= 5000);
if(shouldLogProgress)
lastProgressLogTick = nowTick;
//--- During a genetic-tuner candidate evaluation (m_evalMode) the cap is the rung's small era
//--- budget, and hitting it just ENDS this candidate's throwaway scoring run silently - no
//--- operator prompt (there's nothing to deploy) and no "complete" flag (it's a throwaway net).
int effectiveEraCap = m_evalMode ? m_evalEraBudget : m_maxErasPerRun;
if(m_evalMode && effectiveEraCap > 0 && m_erasSinceCooldown >= effectiveEraCap)
{
stop = true; // candidate hit its rung budget - score is already tracked in m_bestBalancedOos
}
//--- PLATEAU LADDER, terminal stage: training stopped improving and both escape attempts (warm
//--- restart, gamma anneal) failed to find anything better - see the ladder in the era-end block
//--- below, which is what raised m_plateauStage this far and already logged why. This is the
//--- normal, expected way a run finishes now that there is no absolute accuracy target to hit:
//--- it trains until it genuinely stops getting better, then deploys its best checkpoint.
//--- Same mechanism as the operator's "No" answer at the era cap (see that branch's comments for
//--- why m_trainingComplete is set here and why m_trainingStopRequested deliberately is NOT):
//--- stop ends this era loop, FinalizeTrainRun() then restores and deploys the best checkpoint.
else if(!m_evalMode && m_plateauStage >= PLATEAU_STAGE_DEPLOY && m_bestPassedRecall && m_haveOosCheckpoint)
{
stop = true;
m_trainingComplete = true;
}
else if(effectiveEraCap > 0 && m_erasSinceCooldown >= effectiveEraCap)
{
//--- Era cap reached without converging: ask the operator whether to keep training or
//--- deploy the best checkpoint and stop (see PromptContinuePastEraCap / m_maxErasPerRun).
if(PromptContinuePastEraCap(dOosForecast))
{
m_erasSinceCooldown = 0; // keep training - reset the cap window
Print(ID + ": hit the " + IntegerToString(m_maxErasPerRun) + "-era cap (best balanced accuracy " + DoubleToString(m_bestBalancedOos, 1) + "%, blended OOS " + DoubleToString(dOosForecast, 1) + "%) - CONTINUING training by operator choice.");
}
else
{
//--- stop: end THIS era loop now; FinalizeTrainRun (reached via the stop path below,
//--- because stop==true) deploys the best checkpoint. Deliberately do NOT set
//--- m_trainingStopRequested here: m_trainingComplete alone already routes every later
//--- tick to RefreshConvergedSignal (see ScheduleTrainingIfNeeded's if-branch precedence),
//--- so training never re-arms - and leaving m_trainingStopRequested false lets the deployed
//--- model run live inference AND online continual learning IN-SESSION, exactly like a
//--- normal-convergence deploy (which never sets it either). A panel Stop (StopTraining())
//--- still sets it and halts everything, including online learning - that distinction is
//--- preserved. See OnlineLearnStep()'s gate.
stop = true;
//--- Operator DELIBERATELY chose to deploy this best checkpoint as the final model. That's a
//--- terminal decision and must be PERSISTED as such: mark it complete so a later reload
//--- (chart restart OR strategy tester) runs inference instead of silently resuming a full
//--- training run. This is the terminal-deploy path; a mid-training Stop click
//--- (StopTraining()) leaves m_trainingComplete false on purpose so that genuinely-
//--- interrupted run does resume. Note the m_trainingComplete=(m_objectiveMet&&m_oosStable)
//--- line below is inside if(!stop), so it can't clobber this back to false on this path.
m_trainingComplete = true;
Print(ID + ": hit the " + IntegerToString(m_maxErasPerRun) + "-era cap before the plateau ladder finished (best balanced accuracy " + DoubleToString(m_bestBalancedOos, 1) + "%, blended OOS " + DoubleToString(dOosForecast, 1) + "%) - operator chose to DEPLOY the best checkpoint as final (marked complete; reloads will run inference, not retrain). Reaching this cap now means the run was still finding new bests, or never cleared the per-class recall floor (need >=" + IntegerToString(m_minDirectionalRecallPct) + "% each) - raise the era cap for the former, relax MinRecall/SwingConfirmationBars for the latter.");
}
}
}
if(!stop)
{
dError = Net.getRecentAverageError();
if(add_loop)
{
if(m_oosSamples > 0)
{
// Confidence calibration (classification head only - see m_confidenceCalScale's
// declaration comment): compare this era's actual OOS accuracy against the average
// confidence magnitude the model claimed, EMA-blend the resulting scale into
// m_confidenceCalScale so SignedAIConfidence() reports something closer to a real
// probability instead of the raw, uncalibrated softmax value.
if(m_outputNeuronsCount == 3 && m_oosConfidenceSum > 0.0)
{
double empiricalAccuracy = (double)(m_oosBuyHits + m_oosSellHits + m_oosNeutralHits) / m_oosSamples;
double avgClaimedConfidence = m_oosConfidenceSum / m_oosSamples;
double eraScale = MathMax(0.3, MathMin(1.5, empiricalAccuracy / avgClaimedConfidence));
m_confidenceCalScale += (eraScale - m_confidenceCalScale) / Net.recentAverageSmoothingFactor;
}
// Per-class recall gate, symmetric across all three classes: a model that "wins" on
// blended dOosForecast purely by calling everything Neutral (or, just as biased, by
// over-calling Buy/Sell at Neutral's expense) would still pass a plain accuracy check -
// require Buy, Sell, AND Neutral OOS recall to each individually clear
// m_minDirectionalRecallPct so the network can't converge while biased toward any one
// output. A class with FEWER than MIN_OOS_CLASS_SAMPLES_FOR_GATE true OOS samples this
// era doesn't block (recallPct == -1 => treated as passing) so a thin OOS window doesn't
// deadlock convergence early in a run. Computed BEFORE the checkpoint/eta-decay block
// below (not just the final m_objectiveMet gate) so "best" ranking is recall-aware too -
// see isBetterEra's comment for why that matters.
//
// The threshold matters: a bare ">0" here (the original behavior) let a run converge at
// era 44-46 with the OOS window containing exactly ZERO true Buy/Sell bars that era
// (logged as "OOS recall Buy:n/a Sell:n/a Neutral:100%") - a full Neutral-only collapse
// that the gate waved through because there was nothing to measure recall against, not
// because the model was actually unbiased. Requiring a real minimum sample count means
// an unlucky/thin OOS slice blocks convergence instead of silently passing it.
int buyRecallPct = (m_oosBuyTotal >= MIN_OOS_CLASS_SAMPLES_FOR_GATE) ? (int)MathRound(100.0 * m_oosBuyHits / m_oosBuyTotal) : -1;
int sellRecallPct = (m_oosSellTotal >= MIN_OOS_CLASS_SAMPLES_FOR_GATE) ? (int)MathRound(100.0 * m_oosSellHits / m_oosSellTotal) : -1;
int neutralRecallPct = (m_oosNeutralTotal >= MIN_OOS_CLASS_SAMPLES_FOR_GATE) ? (int)MathRound(100.0 * m_oosNeutralHits / m_oosNeutralTotal) : -1;
logBuyRecallPct = buyRecallPct;
logSellRecallPct = sellRecallPct;
logNeutralRecallPct = neutralRecallPct;
m_lastBuyRecallPct = buyRecallPct;
m_lastSellRecallPct = sellRecallPct;
// Predicted-rate (share of ALL OOS bars this era the model called this class, regardless
// of whether that call was right) and precision (of just those calls, how many were
// right) - see logBuyPredPct's declaration comment above for why this is worth logging
// alongside recall. Denominator is the per-era OOS bar count (sum of the per-class true
// totals, all tallied in the same pass-3 block and reset together each era) - NOT
// m_oosSamples, which only resets on a full model reset and so accumulates across every
// era of the run: dividing this era's calls by that all-run total diluted the logged
// rate by roughly the era number (observed: era-15 "Buy:2%" that was really ~30%),
// making a genuinely directional model read as a nearly-dead output.
int oosEraBars = m_oosBuyTotal + m_oosSellTotal + m_oosNeutralTotal;
logBuyPredPct = (oosEraBars > 0) ? (int)MathRound(100.0 * m_oosBuyPredicted / oosEraBars) : -1;
logSellPredPct = (oosEraBars > 0) ? (int)MathRound(100.0 * m_oosSellPredicted / oosEraBars) : -1;
logBuyPrecPct = (m_oosBuyPredicted > 0) ? (int)MathRound(100.0 * m_oosBuyPredictedHits / m_oosBuyPredicted) : -1;
logSellPrecPct = (m_oosSellPredicted > 0) ? (int)MathRound(100.0 * m_oosSellPredictedHits / m_oosSellPredicted) : -1;
//--- Live-fired precision (what actually trades - see m_oosBuyFired): of the directional
//--- calls that cleared the confidence floor under the live/prior-corrected rule this era,
//--- how many were right. Cached for the panel/log; -1 = the model fired none this era.
logBuyFiredPrecPct = (m_oosBuyFired > 0) ? (int)MathRound(100.0 * m_oosBuyFiredHits / m_oosBuyFired) : -1;
logSellFiredPrecPct = (m_oosSellFired > 0) ? (int)MathRound(100.0 * m_oosSellFiredHits / m_oosSellFired) : -1;
m_lastBuyFiredPrecPct = logBuyFiredPrecPct;
m_lastSellFiredPrecPct = logSellFiredPrecPct;
m_lastBuyFired = m_oosBuyFired;
m_lastSellFired = m_oosSellFired;
bool directionalRecallOK = (buyRecallPct < 0 || buyRecallPct >= m_minDirectionalRecallPct) &&
(sellRecallPct < 0 || sellRecallPct >= m_minDirectionalRecallPct) &&
(neutralRecallPct < 0 || neutralRecallPct >= m_minDirectionalRecallPct);
// Balanced accuracy (macro-recall): the mean of the three per-class recalls - the metric
// the checkpoint SELECTION ranks on (see m_bestBalancedOos). Unlike blended accuracy it
// weights Buy, Sell and Neutral equally, so it can't be inflated by the ~96%-Neutral base
// rate. Computed from the same raw per-era recalls the floor uses (not smoothed - the
// whole recall-driven side of this block is per-era-raw by design). When a directional
// class is thin/unmeasured this era (recall -1), balanced accuracy isn't meaningful, so
// fall back to the blended dOosForecast for ranking that era (prior behavior) rather than
// averaging a partial set - the recall floor + directionalRecallMeasured still guard the
// actual convergence decision separately.
double balancedOosEra = (buyRecallPct >= 0 && sellRecallPct >= 0 && neutralRecallPct >= 0)
? (buyRecallPct + sellRecallPct + neutralRecallPct) / 3.0
: dOosForecast;
logBalancedAccPct = (buyRecallPct >= 0 && sellRecallPct >= 0 && neutralRecallPct >= 0)
? (int)MathRound(balancedOosEra) : -1;
// A real (non-thin-sample, i.e. not the -1 "n/a" sentinel) 0% recall on any class means
// the model never once got that class right this era - a majority-class collapse
// (predict-everything-Neutral, or symmetrically a Buy/Sell-only collapse), not progress
// toward separating classes. Before any era has ever passed the recall floor,
// isBetterEra's fallback below is a pure blended-accuracy tiebreak, and blended accuracy
// is trivially maximized by collapsing to the majority class. Observed in practice
// (2026-07-19, SP500 H4): once a run landed on a 0%/0%/100% Buy/Sell/Neutral era, its
// accuracy kept creeping upward for 124 STRAIGHT eras purely from sharpening the
// Neutral-vs-everything boundary - each tick registered as a "new best", re-anchoring the
// checkpoint AND bumping eta back toward its ceiling (the recovery bump below), actively
// rewarding the collapse instead of remaining neutral to it. Excluding these eras from
// isBetterEra denies them that anchor/reward without touching the restore/decay branch
// below, which stays exactly as gated on m_bestPassedRecall as before - see that block's
// own comment for why loosening THAT part pre-pass caused a worse failure historically.
bool isFullyCollapsedEra = (buyRecallPct == 0 || sellRecallPct == 0 || neutralRecallPct == 0);
// Lexicographic "better than the best-so-far" ordering: passing the directional recall
// floor always outranks not passing it, regardless of blended dOosForecast; only WITHIN
// the same pass/fail category does blended accuracy break the tie. Without this, an era
// that traded a few "safe" Neutral calls for genuinely useful (recall-improving) Buy/Sell
// calls would look like a regression in blended-accuracy-only terms and get its
// checkpoint skipped / learning rate cut - fighting directly against the network learning
// to call Buy/Sell at all, since Neutral is the large majority class (~80%+ of labels) and
// a model that just calls everything Neutral already scores well on blended accuracy
// alone. isWorseEra mirrors the same ordering for the eta-decay-on-regression trigger.
// (directionalRecallOK implies !isFullyCollapsedEra already, since the floor is always
// >0%, so the first clause below needs no extra guard - only the pre-pass accuracy-only
// tiebreak in the second clause does.)
// Within the same recall-pass category the tie now breaks on BALANCED accuracy, not the
// Neutral-dominated blended dOosForecast - see m_bestBalancedOos. This is what deploys the
// most class-balanced era instead of the most Neutral-leaning one, and it also strengthens
// the pre-pass phase: a Neutral-only era scores (0+0+N)/3 in balanced terms (low) rather
// than the ~80% it scores in blended terms, so it can no longer re-anchor the checkpoint.
bool isBetterEra = (directionalRecallOK && !m_bestPassedRecall) ||
(directionalRecallOK == m_bestPassedRecall && !isFullyCollapsedEra && balancedOosEra > m_bestBalancedOos);
// The recall-pass-loss clause used to fire on ANY drop out of a full 3-way recall pass,
// even a near-miss on one class at unchanged accuracy (e.g. observed: Buy:56% Sell:41%
// Neutral:34% - Neutral alone missing the 40% floor by a few points) - treating that
// identically to a total collapse back to Neutral-only. With three classes all needing
// to simultaneously clear the floor, that made isWorseEra fire on most eras once a pass
// was ever achieved, ratcheting eta toward ETA_MIN within a handful of eras and then
// (before the recovery bump below existed) leaving it stuck there permanently - visible
// in practice as ~25 back-to-back identical "regressed from best 70.6% to 70.6%" eras.
// Now only counts as worse if accuracy ALSO dropped meaningfully (same threshold
// regardless of whether the recall-pass flag changed too) - losing the recall-pass flag
// at flat/improved accuracy is borderline variance, not a regression worth
// restoring+decaying over.
bool isWorseEra = balancedOosEra < m_bestBalancedOos - ETA_DECAY_REGRESSION_PCT;
if(isBetterEra)
{
//--- Snapshot BOTH scores at the checkpoint: m_bestBalancedOos is what ranking compares
//--- against next era; m_bestOosForecast keeps the blended value FinalizeTrainRun() and
//--- the restore branch reset dOosForecast to (see m_bestBalancedOos' declaration).
m_bestOosForecast = dOosForecast;
m_bestBalancedOos = balancedOosEra;
m_bestPassedRecall = directionalRecallOK;
//--- eval candidates are throwaway - track the score (above) but never write a checkpoint
//--- file; m_haveOosCheckpoint=false then also skips the worse-era RestoreWeights() restore.
//--- In-MEMORY weight snapshot (not a file): the file-based checkpoint re-created every
//--- neuron on restore, which the CPU-DLL backend can't do for a second live set - see
//--- CNet::CaptureWeights/RestoreWeights. eval candidates snapshot nothing.
m_haveOosCheckpoint = m_evalMode ? false : Net.CaptureWeights();
// Recovery bump: ETA_DECAY_FACTOR-only ever shrinks eta, and previously nothing ever
// grew it back - a losing streak early in a run (even a since-corrected one) would
// permanently cap how fast every later era could learn for the rest of the run, all
// the way down to ETA_MIN with no way back. A genuinely better era (new best, not
// just a tie) means the current eta is working, so nudge it back up a bit - capped at
// this model's own configured starting rate (m_etaCeiling - AdamLearningRate for
// ADAM, SgdLearningRate for SGD, see that member's declaration comment) so this
// can't runaway past the rate training was actually tuned to start at.
eta = MathMin(m_etaCeiling, eta / ETA_DECAY_FACTOR);
}
else
if(isWorseEra && m_bestOosForecast > 0)
{
// Decaying eta alone only softens FUTURE steps - it does nothing to undo the
// regression this era already baked into the weights, so a run could (and in
// practice did) spend 15+ eras compounding forward from one bad era's damage,
// each new era fighting the last one's overshoot instead of building on the best
// state found so far. Restore the last checkpointed-good weights before continuing
// (mirrors what FinalizeTrainRun() does at the END of a run, just applied live so
// the oscillation can't compound within a single run) - this is what actually turns
// "reduce LR on regression" into "step back, then retry slower", not just "drift
// slower".
//
// BOTH the restore AND the eta decay below are gated on m_bestPassedRecall: before
// ANY era has ever cleared the per-class recall floor, isBetterEra's own
// lexicographic ordering degrades to a pure blended-accuracy tiebreak
// (directionalRecallOK==false on both sides of the comparison), so "best checkpoint"
// during that phase just means "called Neutral most confidently so far" - restoring
// it would actively defend the majority-class collapse against any era that trades
// some accuracy for real Buy/Sell recall, which is exactly the bias this whole
// recall-gate mechanism exists to prevent (see isBetterEra's own comment above).
// Observed in practice: era 1-3 all "improved" on accuracy alone
// (24.9%->41.4%->52.3%) while Buy/Sell recall stayed at a flat 0% the entire time -
// restoring pre-pass would have locked training into that trajectory instead of
// letting it explore past it. Decaying eta has the same bias one step removed:
// every regression relative to a Neutral-collapse "best" shrinks eta a little more,
// steadily strangling the exploration needed to escape that collapse until eta
// bottoms out at ETA_MIN with no real solution ever found and no checkpoint to fall
// back on either - observed in practice as a run whose best-ever blended accuracy
// kept landing on 0%/0%/100% Buy/Sell/Neutral recall eras, each one triggering
// another decay on the very next era, until eta floored out around era 20 and the
// remaining eras just oscillated between collapse states with no way to make a
// large-enough move to escape and no way to reset. Once m_bestPassedRecall is true,
// there IS a genuinely good state worth protecting, and both restoring the
// checkpoint and decaying eta on regression are safe/correct again.
fix(training): escape the recall-gate catch-22 that let runs decay unchecked Evidence (MQL5\Logs, SP500 H1, 2026-07-29): Perceptron era 61 Buy 32% Sell 27% Neut 94% bal 51% LSTM era 160 Buy 16% Sell 11% Neut 98% bal 42% (peaked 49% @ era 44) Hybrid era 179 Buy 5% Sell 2% Neut 99% bal 35% (peaked 41%) CONV era 228 Buy 2% Sell 4% Neut 99% bal 35% (peaked 40% @ era 122) Every model peaks early then decays monotonically toward Neutral, and nothing stops it: the restore-best-weights + decay-eta handler is gated on m_bestPassedRecall, which stays false forever when no checkpoint ever clears the per-class floor. CONV ran 228 eras with eta pinned at its 0.000300 start. The plateau ladder cannot end such a run either (stage 3 refuses to deploy without a recall pass, so it resets ~27 times), making it a 1000-era one-way trip. The gate's own justification had expired. It was written when the pre-pass tiebreak was blended-accuracy-only, where "best" really did mean "called Neutral most confidently". The balanced-selection change replaced that with `balancedOosEra > m_bestBalancedOos` plus an isFullyCollapsedEra exclusion, so a Neutral-only era now scores ~33% - the FLOOR of the balanced metric - and cannot anchor the checkpoint at all. Pre-pass "best" now means "most class-balanced so far", which is worth defending; and isWorseEra is itself a balanced-accuracy regression, so it cannot fire merely for trading Neutral calls for Buy/Sell. The original concern still holds while the best-so-far IS near-collapse, so the escape is margin-guarded: defend the checkpoint only once balanced accuracy sits more than BALANCED_WORTH_DEFENDING_MARGIN_PCT (5pp) above the one-class floor of 100/3. Against the run above that engages for all three stuck topologies (42.3/41.3/50.0 vs a 38.3 threshold) while a genuinely collapsed run still explores freely. Two inputs restored to the regime that actually produced a deploy: - MinRecall 60 -> 40. The one successful auto-deploy in the logs (Hybrid, 28th 00:50, best balanced 66.0%) ran against a 40% floor. 60 has never been shown reachable here - a floor above what the config can reach is the same "target set too high" failure the surrounding comment already warns about. - OversampleParity 60 -> 90. 60 overcorrected. Runs now START Neutral-dominant (Buy 0-11% recall at era 1) and call Buy/Sell on 0-4% of bars against a ~6% true base rate - under-calling, with no headroom to converge down from. The deploying run began at Buy 90% / Sell 36%, 24% of bars called, and settled into the floor from above. Raw over-calling is the intended starting condition; live calls are base-rate-calibrated by AILogitPriorStrength, which is why the input's own note says to judge over-calling by live-fired precision, not raw counts. Compiles 0 errors, 0 warnings. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 07:51:08 -04:00
// 2026-07-29: the m_bestPassedRecall gate above has an escape now, because its
// stated premise expired. It was written when the pre-pass tiebreak really was
// blended-accuracy-only; the balanced-selection change (m_bestBalancedOos) replaced
// that with `balancedOosEra > m_bestBalancedOos` AND an isFullyCollapsedEra
// exclusion, so a Neutral-only era now scores ~33% (the FLOOR of the balanced
// metric) and cannot anchor the checkpoint at all. "Best checkpoint" pre-pass
// therefore no longer means "called Neutral most confidently" - it means "most
// class-balanced state found so far", which is worth defending, and isWorseEra is
// itself a balanced-accuracy regression, so it cannot fire merely for trading
// Neutral calls for Buy/Sell.
//
// Leaving the gate absolute had a failure mode of its own, and it is not
// hypothetical: if NO checkpoint ever clears the recall floor, m_bestPassedRecall
// stays false forever, so there is never any restore and never any eta decay.
// Observed on SP500 H1 2026-07-29 across three topologies - CONV ran 228 eras with
// eta pinned at its 0.000300 start while balanced accuracy slid 40% -> 35% and Buy
// recall 11% -> 2%. The run had no regression control whatsoever, and the plateau
// ladder could not end it either (stage 3 refuses to deploy without a recall pass),
// so it was a 1000-era one-way trip into a Neutral collapse.
//
// The original concern still applies while the best-so-far IS near-collapse:
// decaying eta against such a "best" strangles the exploration needed to escape it.
// So the escape is margin-guarded - defend the checkpoint only once it sits clearly
// above the one-class floor, which is exactly when there is something real to lose.
bool bestWorthDefending = (m_bestBalancedOos >
BALANCED_COLLAPSE_PCT + BALANCED_WORTH_DEFENDING_MARGIN_PCT);
if(m_bestPassedRecall || bestWorthDefending)
refactor: split CExpertSignalAIBase implementation by responsibility ExpertSignalAIBase.mqh was 8216 lines: the class declaration followed by 87 method bodies covering training, labelling, feature extraction, persistence, chart drawing, online learning, the GA auto-tuner and inference, all in one file. Train() alone is 1492 lines; a change to arrow drawing meant scrolling past the era loop. Moved the bodies into Expert\AIBase\, included at the bottom of the original after the class declaration: Training.mqh 1607 era loop, plateau ladder, checkpoint select, deploy Features.mqh 1093 indicator creation + per-bar input feature vector ChartUI.mqh 634 arrows, arrow persistence, status panel, cleanup Persistence.mqh 492 .stats/.cfg sidecars, CPU-inference validation, copy OnlineLearning.mqh 461 live continual learning, EMA shadow, OOS simulator Labels.mqh 309 ZigZag pivot labels, async label-cache prebuild AutoTune.mqh 275 genetic tuner (population, crossover, halving) Inference.mqh 235 softmax, prior calibration, class priors ExpertSignalAIBase.mqh 8216 -> 3131 (declaration + topology build only) This is a pure relocation - verified mechanically, not by eye: HEAD's file reconstructed from the eight partials plus the surviving remainder is byte-identical to HEAD, span for span (scratchpad verify_split.py). No declaration moved, no signature changed, no code rewritten, so behaviour is unchanged by construction. Compiles 0 errors, 0 warnings. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 00:42:45 -04:00
{
if(m_haveOosCheckpoint && Net.RestoreWeights())
dOosForecast = m_bestOosForecast;
if(eta > ETA_MIN)
eta = MathMax(ETA_MIN, eta * ETA_DECAY_FACTOR);
Print(ID + ": OOS balanced accuracy regressed from best " + DoubleToString(m_bestBalancedOos, 1) +
"% to " + DoubleToString(balancedOosEra, 1) + "% (blended " + DoubleToString(m_bestOosForecast, 1) +
"%->" + DoubleToString(dOosForecast, 1) + "%) - restoring best checkpoint and decaying learning rate to " + DoubleToString(eta, 6));
}
else
Print(ID + ": OOS balanced accuracy regressed from best " + DoubleToString(m_bestBalancedOos, 1) +
"% to " + DoubleToString(balancedOosEra, 1) + "% (blended " + DoubleToString(m_bestOosForecast, 1) +
fix(training): escape the recall-gate catch-22 that let runs decay unchecked Evidence (MQL5\Logs, SP500 H1, 2026-07-29): Perceptron era 61 Buy 32% Sell 27% Neut 94% bal 51% LSTM era 160 Buy 16% Sell 11% Neut 98% bal 42% (peaked 49% @ era 44) Hybrid era 179 Buy 5% Sell 2% Neut 99% bal 35% (peaked 41%) CONV era 228 Buy 2% Sell 4% Neut 99% bal 35% (peaked 40% @ era 122) Every model peaks early then decays monotonically toward Neutral, and nothing stops it: the restore-best-weights + decay-eta handler is gated on m_bestPassedRecall, which stays false forever when no checkpoint ever clears the per-class floor. CONV ran 228 eras with eta pinned at its 0.000300 start. The plateau ladder cannot end such a run either (stage 3 refuses to deploy without a recall pass, so it resets ~27 times), making it a 1000-era one-way trip. The gate's own justification had expired. It was written when the pre-pass tiebreak was blended-accuracy-only, where "best" really did mean "called Neutral most confidently". The balanced-selection change replaced that with `balancedOosEra > m_bestBalancedOos` plus an isFullyCollapsedEra exclusion, so a Neutral-only era now scores ~33% - the FLOOR of the balanced metric - and cannot anchor the checkpoint at all. Pre-pass "best" now means "most class-balanced so far", which is worth defending; and isWorseEra is itself a balanced-accuracy regression, so it cannot fire merely for trading Neutral calls for Buy/Sell. The original concern still holds while the best-so-far IS near-collapse, so the escape is margin-guarded: defend the checkpoint only once balanced accuracy sits more than BALANCED_WORTH_DEFENDING_MARGIN_PCT (5pp) above the one-class floor of 100/3. Against the run above that engages for all three stuck topologies (42.3/41.3/50.0 vs a 38.3 threshold) while a genuinely collapsed run still explores freely. Two inputs restored to the regime that actually produced a deploy: - MinRecall 60 -> 40. The one successful auto-deploy in the logs (Hybrid, 28th 00:50, best balanced 66.0%) ran against a 40% floor. 60 has never been shown reachable here - a floor above what the config can reach is the same "target set too high" failure the surrounding comment already warns about. - OversampleParity 60 -> 90. 60 overcorrected. Runs now START Neutral-dominant (Buy 0-11% recall at era 1) and call Buy/Sell on 0-4% of bars against a ~6% true base rate - under-calling, with no headroom to converge down from. The deploying run began at Buy 90% / Sell 36%, 24% of bars called, and settled into the floor from above. Raw over-calling is the intended starting condition; live calls are base-rate-calibrated by AILogitPriorStrength, which is why the input's own note says to judge over-calling by live-fired precision, not raw counts. Compiles 0 errors, 0 warnings. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 07:51:08 -04:00
"%->" + DoubleToString(dOosForecast, 1) + "%) - best so far is still within " +
DoubleToString(BALANCED_WORTH_DEFENDING_MARGIN_PCT, 1) + "pp of the " +
DoubleToString(BALANCED_COLLAPSE_PCT, 1) + "% one-class floor, so there is nothing worth" +
" restoring yet - continuing to explore without decaying eta (still " + DoubleToString(eta, 6) + ")");
refactor: split CExpertSignalAIBase implementation by responsibility ExpertSignalAIBase.mqh was 8216 lines: the class declaration followed by 87 method bodies covering training, labelling, feature extraction, persistence, chart drawing, online learning, the GA auto-tuner and inference, all in one file. Train() alone is 1492 lines; a change to arrow drawing meant scrolling past the era loop. Moved the bodies into Expert\AIBase\, included at the bottom of the original after the class declaration: Training.mqh 1607 era loop, plateau ladder, checkpoint select, deploy Features.mqh 1093 indicator creation + per-bar input feature vector ChartUI.mqh 634 arrows, arrow persistence, status panel, cleanup Persistence.mqh 492 .stats/.cfg sidecars, CPU-inference validation, copy OnlineLearning.mqh 461 live continual learning, EMA shadow, OOS simulator Labels.mqh 309 ZigZag pivot labels, async label-cache prebuild AutoTune.mqh 275 genetic tuner (population, crossover, halving) Inference.mqh 235 softmax, prior calibration, class priors ExpertSignalAIBase.mqh 8216 -> 3131 (declaration + topology build only) This is a pure relocation - verified mechanically, not by eye: HEAD's file reconstructed from the eight partials plus the surviving remainder is byte-identical to HEAD, span for span (scratchpad verify_split.py). No declaration moved, no signature changed, no code rewritten, so behaviour is unchanged by construction. Compiles 0 errors, 0 warnings. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 00:42:45 -04:00
}
//=== PLATEAU LADDER ====================================================================
//--- Neither branch above fires in the dead zone between "new best" and "regressed by more
//--- than ETA_DECAY_REGRESSION_PCT". This is the response to sitting in it: count eras since
//--- the last new best and escalate. See the PLATEAU_* constants for the full rationale and
//--- the two papers behind the two escape moves. m_evalMode candidates are throwaway runs
//--- scored over a handful of eras (a rung budget) - escalating inside one would corrupt the
//--- comparison between candidates, so they only ever run stage 0.
if(!m_evalMode)
{
if(isBetterEra)
{
//--- Moving again: retire the ladder. Deliberately does NOT restore the annealed
//--- gamma - scheduled-gamma annealing is monotone within a run (see PLATEAU_*), and
//--- if a lower gamma is what produced this new best, putting it back would undo the
//--- very change that worked.
if(m_plateauStage > 0)
Print(ID + ": new best balanced accuracy " + DoubleToString(m_bestBalancedOos, 1) +
"% - plateau escape worked, clearing plateau stage " + IntegerToString(m_plateauStage) +
" (focal gamma stays at " + DoubleToString(m_focalGammaRuntime, 2) + ")");
m_erasSinceBestBalanced = 0;
m_plateauStage = 0;
}
else
{
m_erasSinceBestBalanced++;
int dueStage = m_erasSinceBestBalanced / PLATEAU_PATIENCE_ERAS;
if(dueStage > m_plateauStage)
{
m_plateauStage = dueStage;
string stageNote = IntegerToString(m_erasSinceBestBalanced) + " eras with no new best balanced accuracy (best " +
DoubleToString(m_bestBalancedOos, 1) + "%)";
if(m_plateauStage == PLATEAU_STAGE_RESTART || m_plateauStage == PLATEAU_STAGE_ANNEAL)
{
//--- WARM RESTART: jump eta back to this model's configured starting rate. A
//--- plateau needs a bigger step to climb out of its basin, not a smaller one.
double etaBefore = eta;
eta = m_etaCeiling;
//--- GAMMA ANNEAL (classification head only - gamma is meaningless for the
//--- single-output regression head). Monotone: only ever steps down.
double gammaBefore = m_focalGammaRuntime;
if(m_outputNeuronsCount == 3)
m_focalGammaRuntime = (m_plateauStage == PLATEAU_STAGE_RESTART)
? MathMin(m_focalGammaRuntime, PLATEAU_GAMMA_STEP)
: 0.0;
Print(ID + ": PLATEAU stage " + IntegerToString(m_plateauStage) + " - " + stageNote +
". Warm restart: learning rate " + DoubleToString(etaBefore, 6) + "->" + DoubleToString(eta, 6) +
", focal gamma " + DoubleToString(gammaBefore, 2) + "->" + DoubleToString(m_focalGammaRuntime, 2) +
(m_focalGammaRuntime <= 0.0 ? " (plain class-balanced weighting)" : "") +
". Best checkpoint is safe - this only changes how the NEXT eras train.");
}
else
if(m_plateauStage >= PLATEAU_STAGE_DEPLOY)
{
//--- Exhausted: both escapes were tried and neither found a better model, so
//--- this IS the best this configuration reaches. The deploy itself happens in
//--- the era-cap/plateau branch at the TOP of the next era, which reuses the
//--- proven "stop + mark complete -> FinalizeTrainRun restores and deploys the
//--- best checkpoint" path rather than duplicating it here.
//--- Safety: only ever auto-deploys a checkpoint that CLEARED the per-class
//--- recall floor (m_bestPassedRecall). If nothing ever did, there is no model
//--- worth deploying - so the ladder resets and keeps trying instead, leaving
//--- the era cap as the ultimate backstop. That is what stops "train to the
//--- best possible result" from degenerating into "deploy a Neutral collapse".
if(m_bestPassedRecall && m_haveOosCheckpoint)
Print(ID + ": PLATEAU stage " + IntegerToString(PLATEAU_STAGE_DEPLOY) + " - " + stageNote +
" after a warm restart AND full gamma anneal. Training has converged on what this"
+ " configuration can reach - deploying the best checkpoint (balanced accuracy "
+ DoubleToString(m_bestBalancedOos, 1) + "%, blended " + DoubleToString(m_bestOosForecast, 1) + "%).");
else
{
Print(ID + ": PLATEAU stage " + IntegerToString(PLATEAU_STAGE_DEPLOY) + " - " + stageNote +
", but no checkpoint has ever cleared the per-class recall floor (need >=" +
IntegerToString(m_minDirectionalRecallPct) + "% Buy/Sell/Neutral each), so there is nothing safe to"
+ " deploy. Restarting the plateau ladder and continuing to train rather than deploying a"
+ " one-class model; the " + IntegerToString(m_maxErasPerRun) + "-era cap remains the backstop.");
m_erasSinceBestBalanced = 0;
m_plateauStage = 0;
}
}
}
}
}
m_oosWindow.Add(dOosForecast);
while(m_oosWindow.Total() > STABILITY_WINDOW)
m_oosWindow.Delete(0);
m_oosStable = false;
if(m_oosWindow.Total() >= STABILITY_WINDOW)
{
double oosMin = m_oosWindow.At(0), oosMax = m_oosWindow.At(0);
for(int w = 1; w < m_oosWindow.Total(); w++)
{
oosMin = MathMin(oosMin, m_oosWindow.At(w));
oosMax = MathMax(oosMax, m_oosWindow.At(w));
}
m_oosStable = (oosMax - oosMin) <= STABILITY_TOLERANCE;
}
// The dError<0.1 RMS-error floor is meaningful for the single-neuron regression head
// (m_outputNeuronsCount==1), where it's the only convergence signal available. For the
// 3-neuron one-hot classification head it's redundant with, and far stricter than,
// dOosForecast/directionalRecallOK: reaching RMS error 0.1 across 3 one-hot targets
// needs every output neuron within ~0.17 of its target on average, i.e. near-perfect
// confident calibration on EVERY bar, not just correct argmax calls - unreachable in
// practice under normal market label noise, so classification runs would oscillate
// forever (era after era hitting good OOS accuracy and passing recall, but never
// satisfying this) without this carve-out.
bool errorGateOK = (m_outputNeuronsCount == 3) ? true : (dError < 0.1);
// Convergence (unlike isBetterEra's ranking) FINALIZES the model, so both directional
// classes must have actually been MEASURED this era. An n/a (-1, thin-sample) Buy or
// Sell recall passing directionalRecallOK is deliberate for ranking (early thin
// windows shouldn't deadlock "best" tracking), but letting it pass HERE converges on
// a window that contained no directional bars to disprove the model. Observed
// 2026-07-19: a mid-run label-cache wipe relabeled the whole window Neutral,
// "accuracy" hit 84.9% with Buy/Sell recall both n/a - without this gate a Neutral-only
// model finalizes as a certified success.
bool directionalRecallMeasured = (m_outputNeuronsCount != 3) || (buyRecallPct >= 0 && sellRecallPct >= 0);
//--- VALIDITY of this era's model, no longer "did it hit a target accuracy". The absolute
//--- OOS-accuracy target (the old MinWR input) is gone: an accuracy number typed in ahead of
//--- time is either unreachable for the symbol/timeframe - in which case the run never
//--- converges and burns to the era cap - or set low enough to stop a run that was still
//--- improving. Neither is what "train to the best result" means. Quality is now enforced by
//--- WHICH era gets deployed (balanced-accuracy checkpoint ranking + this per-class recall
//--- floor) and WHEN a run ends (the plateau ladder), not by an accuracy threshold. Note the
//--- recall floor deliberately stays: it is not a performance target but the anti-collapse
//--- gate that makes auto-deploy safe.
m_objectiveMet = errorGateOK && directionalRecallOK && directionalRecallMeasured;
}
// Only mark the persisted model "complete" once it actually converged this era -
// an interruption (stop) or an ordinary in-progress era must stay flagged incomplete
// so a restart resumes training instead of quietly treating a partial run as done.
// m_evalMode throwaway candidates persist NOTHING (no weights, no complete flag, no shadow) -
// only the deployed final retrain writes files; the score lives in m_bestBalancedOos.
if(!m_evalMode)
{
// Convergence = "the plateau ladder is exhausted AND there is a recall-passing checkpoint
// to deploy" - the exact same condition the deploy branch beside the era-cap check uses, so
// the flag written into the .nnw here can never disagree with the decision to stop. While a
// run is still improving (or still has an escape stage left to try) this stays false and the
// per-era save correctly records an in-progress run. Previously this was
// (m_objectiveMet && m_oosStable), which needed the removed absolute accuracy target to mean
// anything: with that target gone m_oosStable alone - just 3 eras inside a 2pp band, which
// is true constantly - would have converged the run at the first flat spot.
m_trainingComplete = (m_plateauStage >= PLATEAU_STAGE_DEPLOY) && m_bestPassedRecall && m_haveOosCheckpoint;
double currentIndicatorParams[];
m_indicatorTuner.Flatten(currentIndicatorParams);
if(!Net.Save(m_activeFileName + ".nnw", dError, dUndefine, dForecast, dtStudied, m_activeFileCommon, m_eraCount, m_trainingComplete, currentIndicatorParams))
Print(__FUNCTION__ + ": ERROR - era-end Net.Save failed for " + m_activeFileName + ".nnw (era " + IntegerToString(m_eraCount) + "). Training continues but this era's checkpoint was NOT persisted - a crash/restart now would resume from an older era.");
if(!SaveModelStats(m_activeFileName, m_activeFileCommon)) // keep calibration state paired with the just-saved weights
Print(__FUNCTION__ + ": ERROR - SaveModelStats failed for " + m_activeFileName + " (era " + IntegerToString(m_eraCount) + "). Calibration/online-learning state not persisted this era.");
SaveShadowNet(currentIndicatorParams);
}
}
}
if(shouldLogProgress)
{
string recallInfo = (logBuyRecallPct < 0 && logSellRecallPct < 0 && logNeutralRecallPct < 0) ? "" :
(" | OOS recall Buy:" + (logBuyRecallPct < 0 ? "n/a" : IntegerToString(logBuyRecallPct) + "%") +
" Sell:" + (logSellRecallPct < 0 ? "n/a" : IntegerToString(logSellRecallPct) + "%") +
" Neutral:" + (logNeutralRecallPct < 0 ? "n/a" : IntegerToString(logNeutralRecallPct) + "%") +
" (need >=" + IntegerToString(m_minDirectionalRecallPct) + "% each)");
//--- Balanced accuracy = the checkpoint-selection metric (see m_bestBalancedOos). Shown so the
//--- number the deployed model is actually chosen on is visible next to the recalls it averages.
string balancedInfo = (logBalancedAccPct < 0) ? "" : (" | OOS balanced acc " + IntegerToString(logBalancedAccPct) + "%");
// See logBuyPredPct's declaration comment for why this is worth logging alongside recall -
// it's what tells apart a suppressed/dead output (predicted rate stuck at 0%) from a
// miscalibrated boundary (predicted rate healthy, precision poor), which look identical from
// recall alone.
string predictedInfo = (logBuyPredPct < 0 && logSellPredPct < 0) ? "" :
(" | OOS calls Buy:" + (logBuyPredPct < 0 ? "n/a" : IntegerToString(logBuyPredPct) + "%") +
" (win rate " + (logBuyPrecPct < 0 ? "n/a" : IntegerToString(logBuyPrecPct) + "%") + ")" +
" Sell:" + (logSellPredPct < 0 ? "n/a" : IntegerToString(logSellPredPct) + "%") +
" (win rate " + (logSellPrecPct < 0 ? "n/a" : IntegerToString(logSellPrecPct) + "%") + ")");
//--- Live-fired precision: the number that actually predicts forward-trading performance - only
//--- the directional calls that cleared the confidence floor under the live/prior-corrected rule
//--- (see AdjustedSignalFromSoftmax). Count in parentheses = how many bars the model would have
//--- traded this era. "0" fires = the calibration is (this era) suppressing all directional trades.
string liveInfo = (m_lastBuyFired <= 0 && m_lastSellFired <= 0) ? " | live fires 0 this era" :
(" | live win rate Buy:" + (logBuyFiredPrecPct < 0 ? "n/a" : IntegerToString(logBuyFiredPrecPct) + "%") +
" (" + IntegerToString(m_lastBuyFired) + ")" +
" Sell:" + (logSellFiredPrecPct < 0 ? "n/a" : IntegerToString(logSellFiredPrecPct) + "%") +
" (" + IntegerToString(m_lastSellFired) + ")");
// Raw-output saturation diagnostic - see m_oosOutMin's declaration comment. Spread ~0 with
// all six min/max values pinned together = the collapsed constant-classifier state.
string rawOutInfo = (m_oosOutCount <= 0) ? "" :
StringFormat(" | OOS raw out B:%.3f..%.3f S:%.3f..%.3f N:%.3f..%.3f spread avg %.4f",
m_oosOutMin[0], m_oosOutMax[0], m_oosOutMin[1], m_oosOutMax[1],
m_oosOutMin[2], m_oosOutMax[2], m_oosOutSpreadSum / m_oosOutCount);
//--- No "(target X%)" any more - there is no absolute accuracy target. What replaces it as the
//--- progress indicator is the plateau counter: how many eras since the last new best, and how
//--- close that is to ending the run (see the PLATEAU_* ladder).
string plateauInfo = (m_evalMode || m_bestBalancedOos < 0) ? "" :
(" | best bal " + DoubleToString(m_bestBalancedOos, 1) + "%, " + IntegerToString(m_erasSinceBestBalanced) +
" eras since (stage " + IntegerToString(m_plateauStage) + "/" + IntegerToString(PLATEAU_STAGE_DEPLOY) +
", gamma " + DoubleToString(m_focalGammaRuntime, 2) + ")");
Print(ID + ": training in progress - era " + IntegerToString(m_eraCount) + ", OOS accuracy " + DoubleToString(dOosForecast, 1) + "%, IS error " + DoubleToString(dError, 2) + recallInfo + balancedInfo + predictedInfo + liveInfo + plateauInfo + rawOutInfo);
// Forced (unthrottled) panel refresh, right here alongside the console line above, using this
// era's own just-finalized m_eraCount/dOosForecast - see UpdateTrainingStatusLabel's
// declaration comment for why this can't just rely on the next throttled bar-scan call to
// catch up (it would, but a full era later than the console already reported it).
UpdateTrainingStatusLabel("Era complete", m_lastDisplayNeuron0, m_lastDisplayNeuron1, m_lastDisplayNeuron2, m_lastDisplaySignal, true);
}
//--- Genuine convergence THIS era (not a stale m_trainingComplete carried over from a previous
//--- run) - (re)start the evaluation-only continual-learning OOS walk. Always rebuilt fresh from
//--- the just-converged weights; never resumes a stale walk from a superseded model.
//--- m_trainingComplete is the plateau ladder's verdict now (see where it is assigned): "stopped
//--- improving after both escape attempts, and there is a recall-passing checkpoint to deploy".
//--- It replaces the old (m_objectiveMet && m_oosStable) test, which depended on the removed
//--- absolute accuracy target to mean anything - without it, m_oosStable alone (3 eras inside a 2pp
//--- band) would have declared convergence at the first flat spot in every run.
if(!stop && m_trainingComplete && !m_evalMode)
{
Print(ID + ": training CONVERGED at era " + IntegerToString(m_eraCount) + " - this is the best this configuration reached: balanced accuracy " +
DoubleToString(m_bestBalancedOos, 1) + "%, blended OOS " + DoubleToString(dOosForecast, 1) + "%, IS error " + DoubleToString(dError, 2) +
". No new best for " + IntegerToString(m_erasSinceBestBalanced) + " eras across a warm restart and a full focal-gamma anneal." +
" Weights saved, switching to live inference.");
StartOosContinualSimulation(bars, oosCutoff);
}
if(stop || m_trainingComplete)
FinalizeTrainRun();
//--- else: this era is done but the run continues - the next Train() call (re-triggered via
//--- ScheduleTrainingIfNeeded()'s custom event, same mechanism as always) starts the next era
//--- fresh, since m_eraResumePending is false while m_trainRunActive stays true
//--- Save this model's own learning-rate trajectory back out of the shared global before
//--- returning - see m_modelEta's declaration comment. Covers every path that reaches here
//--- (natural era completion, whether or not the run itself just finalized).
m_modelEta = eta;
}
//+------------------------------------------------------------------+
//| Ends the current Train() run: restores the best-scoring era's |
//| checkpointed weights (if any beat the era the loop happened to |
//| end on), persists final state, and clears the resumable-run |
//| flags. Called both from Train() itself (natural stop/converge) |
//| and from StopTraining() (a mid-chunk Stop click won't get |
//| another "New Bar" event to resume into, since |
//| ScheduleTrainingIfNeeded() refuses to schedule while |
//| m_trainingStopRequested is set, so it must finalize synchronously |
//| there instead of being left dangling). |
//+------------------------------------------------------------------+
//| Era-cap decision: keep training (true) or deploy best + stop |
//| (false). Live chart -> operator dialog; headless -> stop. |
//+------------------------------------------------------------------+
bool CExpertSignalAIBase::PromptContinuePastEraCap(double bestOos)
{
//--- No GUI in the Strategy Tester/optimizer - MessageBox() is unavailable there and would just
//--- stall a headless run, so deploy the best checkpoint found so far and stop (the safe default).
if(MQLInfoInteger(MQL_TESTER) || MQLInfoInteger(MQL_OPTIMIZATION) || MQLInfoInteger(MQL_FORWARD))
return false;
//--- Reaching this cap is now the UNUSUAL outcome: a run normally ends itself when the plateau ladder
//--- runs out of escapes (see the PLATEAU_* constants), which is a statement about the run having
//--- stopped improving rather than about any accuracy number. So the interesting question here is why
//--- the ladder had not finished yet - either the run was still finding new bests (just needs more
//--- eras), or nothing has ever cleared the per-class recall floor, which blocks auto-deploy on
//--- purpose so a one-class model can never ship. Spell out which.
bool recallMet = (m_lastBuyRecallPct < 0 || m_lastBuyRecallPct >= m_minDirectionalRecallPct) &&
(m_lastSellRecallPct < 0 || m_lastSellRecallPct >= m_minDirectionalRecallPct);
string neutralNote = (m_priorNeutral > 0.0)
? ("inflated by the ~" + IntegerToString((int)MathRound(m_priorNeutral * 100.0)) + "% Neutral base rate")
: "inflated by the dominant Neutral class";
string reasons = "";
if(!m_bestPassedRecall)
reasons += " - No era has ever cleared the per-class recall floor, so there is no model safe to\n" +
" auto-deploy yet (a model that ignores Buy or Sell must never ship)\n";
else
reasons += " - Still improving: " + IntegerToString(m_erasSinceBestBalanced) + " eras since the last new best, plateau stage " +
IntegerToString(m_plateauStage) + " of " + IntegerToString(PLATEAU_STAGE_DEPLOY) + " (the run ends itself at stage " +
IntegerToString(PLATEAU_STAGE_DEPLOY) + ")\n";
if(!recallMet)
reasons += " - Latest era's per-class recall below the floor: Buy " +
(m_lastBuyRecallPct < 0 ? "n/a" : IntegerToString(m_lastBuyRecallPct) + "%") + " / Sell " +
(m_lastSellRecallPct < 0 ? "n/a" : IntegerToString(m_lastSellRecallPct) + "%") +
" (need >=" + IntegerToString(m_minDirectionalRecallPct) + "% each)\n";
if(!m_objectiveMet)
reasons += " - The latest era did not produce a valid model (recall floor not met/not measured)\n";
string balancedStr = (m_bestBalancedOos > 0.0)
? ("\nBest balanced accuracy (Buy/Sell/Neutral averaged, the metric the deployed\ncheckpoint is chosen on): " + DoubleToString(m_bestBalancedOos, 1) +
"%\nBest blended OOS accuracy: " + DoubleToString(bestOos, 1) + "% (" + neutralNote + ")\n")
: "";
string msg = ID + ": training reached the " + IntegerToString(m_maxErasPerRun) +
"-era cap before it finished on its own.\n\n" +
"Training now runs until it stops improving, then deploys its best model. Status:\n" +
reasons +
balancedStr +
"\nContinue training?\n\n" +
"Yes = keep training for another " + IntegerToString(m_maxErasPerRun) + " eras\n" +
"No = deploy the best checkpoint so far and stop training";
int res = MessageBox(msg, "Warrior EA - training", MB_YESNO | MB_ICONQUESTION);
return (res == IDYES);
}
//+------------------------------------------------------------------+
//+------------------------------------------------------------------+
//| See the declaration comment - the single deploy-persistence path. |
//+------------------------------------------------------------------+
void CExpertSignalAIBase::PersistDeployedModel(void)
{
if(CheckPointer(Net) == POINTER_INVALID)
return;
double currentIndicatorParams[];
m_indicatorTuner.Flatten(currentIndicatorParams);
if(!Net.Save(m_activeFileName + ".nnw", dError, dUndefine, dForecast, dtStudied, m_activeFileCommon, m_eraCount, m_trainingComplete, currentIndicatorParams))
Print(__FUNCTION__ + ": ERROR - Net.Save failed for " + m_activeFileName + ".nnw. The deployed model was NOT persisted to disk.");
//--- Deploy-time gate: does this model's pure-MQL5 forward pass match the backend? If so, an
//--- inference-only backtest can run DLL-free (see ValidateCpuInference / CNet::SetCpuInference).
//--- Persisted into the .stats written next. Chart-only; safe-false everywhere else.
m_mqlInferenceValidated = ValidateCpuInference();
if(!SaveModelStats(m_activeFileName, m_activeFileCommon)) // keep calibration state paired with the just-saved weights
Print(__FUNCTION__ + ": ERROR - SaveModelStats failed for " + m_activeFileName + ". Calibration state not persisted.");
SaveShadowNet(currentIndicatorParams);
}
//+------------------------------------------------------------------+
void CExpertSignalAIBase::FinalizeTrainRun(void)
{
//--- deploy the most stable/best-scoring era's weights rather than whatever the run happened to
//--- end on (which may reflect drift after the objective was first hit, or an aborted run). Restore
//--- is now the in-MEMORY snapshot (CNet::RestoreWeights) - see CaptureWeights' note for why the old
//--- file-based restore couldn't work on the CPU-DLL backend.
if(m_haveOosCheckpoint)
{
if(Net.RestoreWeights())
{
dOosForecast = m_bestOosForecast;
RefreshLatestSignal();
PersistDeployedModel();
}
}
//--- Clean up any legacy on-disk checkpoint from an older (file-based) build so it can't linger.
int checkpointFlags = m_activeFileCommon ? FILE_COMMON : 0;
if(FileIsExist(m_activeFileName + "_ckpt.tmp", checkpointFlags))
FileDelete(m_activeFileName + "_ckpt.tmp", checkpointFlags);
//--- Do NOT advance dtStudied while scoring a throwaway candidate (m_evalMode) - that marker belongs
//--- to the DEPLOYED model's "studied up to" state; a candidate eval must leave it untouched. The
//--- checkpoint block above is already inert in eval mode (m_haveOosCheckpoint stays false).
if(m_eraCount > 0 && !m_evalMode)
dtStudied = m_lastBarTime;
m_trainRunActive = false;
m_eraResumePending = false;
m_haveOosCheckpoint = false;
//--- Persist the arrows now drawn on the chart so a deploy/stop survives a later re-add/recompile
//--- without a retrain (durable even if the terminal never gets a clean OnDeinit). A throwaway
//--- genetic-tuner candidate (m_evalMode) draws nothing worth keeping - skip it.
if(!m_evalMode)
SaveChartSignals();
}
#endif // WARRIOR_AIBASE_TRAINING_MQH