Commit graph Warrior_EA/Expert/AIBase/Training.mqh
Author SHA1 Message Date
AnimateDread
9756e2b64f fix(deinit): O(n^2) arrow prune blew the shutdown budget and littered 3 charts
Reported as "the perceptron correctly cleaned its chart on deinit, the
other 3 did not, abnormal termination". Measured from the 2026-08-01 log,
time from "OnDeinit: shutting down" to MetaTrader force-terminating:

  PAI     3.75 s  -> survived, chart cleaned
  CONV    4.71 s  -> Abnormal termination
  LSTM    4.28 s  -> Abnormal termination
  HYBRID  4.16 s  -> Abnormal termination

In all four the last line printed is the inference census, which is the
end of StopTraining() - so the overrun is inside ShutdownChartCleanup(),
i.e. between saving the arrows and purging them.

The cost is the prune loop at the end of SaveChartSignals():

    for(int i = 0; i < prunedCount; i++)
       ObjectDelete(0, SIG_ARROW_PREFIX + TimeToString(pruned[i]));

ObjectDelete is O(objects) on a crowded chart, so this is O(n^2). It was
harmless while the model called a direction on ~6% of bars. After the
triple-barrier relabel the models call on 83-94% of bars, the chart
carries many thousands of arrows, and the loop overran MetaTrader's
OnDeinit budget - so PurgeChart() never ran and the arrows stayed on
screen. The slow tidy-up starved the fast one.

The work was pure waste at that moment: ShutdownChartCleanup purges every
arrow with a single bulk ObjectsDeleteAll immediately afterwards.
Deleting them one at a time first has no effect except to prevent the
bulk delete from happening at all.

SaveChartSignals takes a pruneChartObjects flag, and the two shutdown
call sites pass false:

 - ShutdownChartCleanup passes `preserveChartArrows`, which is exactly
   right: prune when the arrows are STAYING (chart and sidecar must
   agree), skip when they are about to be purged wholesale.
 - FinalizeTrainRun passes !m_trainingStopRequested. Removing a chart
   MID-ERA reaches StopTraining -> FinalizeTrainRun, which took the
   expensive path a second time, even earlier, before anything had been
   cleared. Same defect one call site up; it only escaped notice because
   the observed removals happened to land between eras.

Normal convergence and the live per-era path are unchanged - they still
prune, which is what keeps the chart object count bounded.

This also restores the invariant the 2026-07 fix intended ("chart cleanup
runs BEFORE the heavy weight save so a stall cannot leave the chart
littered"). That fix moved cleanup ahead of the WEIGHT save, but cleanup
had since grown its own slow step ahead of its own fast one.

Both builds compile 0 errors / 0 warnings.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 10:38:36 -04:00
AnimateDread
6db0519472 perf(autotune): replace the genetic search with a filter score - hours to seconds
MEASURED COST OF THE GA, which is what retired it. Per generation:
  rung 0: 8 cand x 3 seeds x  3 eras =  72 eras
  rung 1: 4 cand x 3 seeds x  8 eras =  96
  rung 2: 2 cand x 3 seeds x 20 eras = 120
  = 288 eras/generation x 4 generations = 1152 eras BEFORE the winner's
real training began. Against the observed era times on SP500 H1:

  PAI     29.1 s/era  ->   9.3 h   (matches the observed 00:37 -> 09:22)
  CONV    41.3 s/era  ->  13.2 h
  LSTM   150.4 s/era  ->  48.1 h
  HYBRID 154.6 s/era  ->  49.5 h

Two days to tune is not a first-run experience, and it is the phase in
which the panel goes quiet, which is what made it look like a hang.

It also bought nothing. The space is 90 points (10 MA periods x 9 MA
types), so 1152 evaluations revisited each point ~13 times; and rungs of
3 and 8 eras cannot separate two MA periods at all. The 2026-08-01 run
proves it: every finalist scored 25.0-25.9% balanced accuracy - below the
33.3% one-class floor, i.e. indistinguishable noise - and the search then
"deployed the winner" of that.

THE ERROR WAS THE SCORING FUNCTION, not its constants. Using a full
training run to choose a feature's period is a wrapper method paying
wrapper prices for a decision that does not need one. The reference book
does not do this: ch. 3.3 selects inputs by measuring each candidate
indicator's CORRELATION with the target and dropping the ones with none,
with no network involved.

So: rank candidates by the MUTUAL INFORMATION between the resulting
feature vector and the triple-barrier label. MI rather than correlation
because the label is 3-class categorical and the features are not
monotonically related to it. Equal-FREQUENCY binning (rank-based),
because these features are ATR-normalised and heavy-tailed - fixed-width
bins put nearly everything in one bucket and report ~0 information for a
genuinely useful feature.

Scoring is arithmetic over the feature cache, so it costs seconds and its
cost is independent of topology: LSTM now tunes as fast as the MLP.
Coordinate sweep, not product sweep - cost is the SUM of per-parameter
candidate counts, so enabling every indicator stays affordable - with a
second pass that breaks early once nothing moves.

Sampling is IS-ONLY. Letting the OOS window influence which indicator
settings ship would mean the holdout had been used for selection and had
stopped being a holdout.

HONEST LIMIT, recorded because it is the price: MI is marginal, so a
parameter that only pays off in combination with another can be missed
(Guyon & Elisseeff 2003, filter vs wrapper). Given the wrapper it
replaces was ranking pure noise at 48 h a run, this is strictly better.

Deleted with it: GaRungEras/GaExtract/GaStore/GaMutate/GaRandomCandidate/
GaBlockCrossover/GaSortAliveByScoreDesc/GaBreedNextGeneration, 14 m_ga*
members, the GA_*/TUNE_POP_* constants, and ComputeTuneTrialBudget.

AND m_evalMode/m_evalEraBudget, because nothing set them any more - 28
read sites all permanently inert. That is not a tidy-up: the `if
(!m_evalMode)` guard on UpdateClassPriors is exactly what silently
disabled the imbalance correction for entire runs two commits ago. Dead
machinery that still reads like live machinery is this codebase's most
expensive recurring bug, and leaving 28 more instances of it would have
been indefensible.

The panel's tuning-progress state goes too - tuning no longer takes long
enough to need one.

Both builds compile 0 errors / 0 warnings.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 10:29:31 -04:00
AnimateDread
3bae2f9254 fix: the imbalance correction never ran during the auto-tune search
Neutral collapse on all four topologies by era 5 with a 2:6 barrier
(recall Buy 0% / Sell 0% / Neutral 100%), and the panel stuck on
"measuring...". One root cause, and it was not the barrier.

The labels were fine: Buy 25.4% / Sell 22.0% / Neutral 52.5%, which is
exactly gambler's ruin for m=2,k=6 (2/8 = 25% per side), with only 0.1%
of Neutral coming from the vertical barrier - so the new m*k horizon
scaling is right, arguably generous.

What was broken: Train()'s era-start block wrapped UpdateClassPriors() in
`if(!m_evalMode)`. The auto-tune GA scores every candidate in eval mode,
and AutoTuneIndicators ships ON, so on a default configuration EVERY era
of the search ran with unmeasured priors. ApplyLogitAdjustment() requires
measured priors; without them it calls ClearLogitAdjustment() and returns.

So the entire search trained under PLAIN cross-entropy. With a 52.5%
majority class the optimum of plain CE is "always predict Neutral", and
that is precisely what all four models found. The panel followed: its
counters only advance on bars the model CALLED Buy or Sell, so a
collapsed model leaves them at zero and the line reads "measuring..."
forever.

This was latent, not new. It has been true for every auto-tuned run, but
it was invisible while the labels were near-balanced - last night's
accidental 1:1 barrier gave 43/40/17, where plain CE has no majority to
collapse into. Widening the stop to 2*ATR (correctly - 1*ATR is too tight
to survive noise) moved Neutral to the majority and exposed it.

The guard's stated fear cannot happen. These priors are measured from the
LABEL distribution, and the tuner only perturbs indicator periods
(MA/RSI/MACD/Ichimoku/AD). The barrier label depends on ATR, SL_Mode and
TP_Mode - none of which the search touches - so every candidate sees
byte-identical labels and identical priors. There is nothing to
contaminate. What the guard actually protected was the .stats write, and
that is gated separately: eval candidates never checkpoint and never
persist.

Also, because this is the THIRD quiet no-op to cost a run in this
codebase (after the fictional oversampling log line and the shadow-blend
skip):

- ApplyLogitAdjustment() now WARNS when it declines to install, instead
  of silently clearing. A mechanism that cannot announce it is not
  running is indistinguishable from one that is.
- The panel distinguishes "measuring..." (before era 1, nothing scored
  yet - an honest warm-up) from "no directional calls yet" (eras trained,
  zero calls - a finding, not a wait).

Both builds compile 0 errors / 0 warnings. No retrain forced by this
commit itself, but the collapsed models must be discarded.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 00:46:24 -04:00
AnimateDread
25813523d3 fix: refuse invalid SL/TP, fix the unreachable deploy floor, scale the horizon
Three defects found by reading the 2026-08-01 training logs, all of which
only became visible because the relabel made the numbers mean something.

1. A STALE ENUM TRAINED FOUR MODELS ON THE WRONG TARGET.

   `OnInit: trade settings snapshot - SL_Mode=1 TP_Mode=-101`

   -101 was TP_PREV_SWING, deleted from TAKE_PROFIT_MODE on 2026-07-31 in
   7eb48f5. MetaTrader does not validate a saved enum input against the
   enum's current members, so charts saved before that kept the old
   integer. BarrierMultiples()'s `if(tpMult <= 0.0) tpMult = slMult;`
   then quietly turned it into a 1:1 barrier, and all four topologies
   trained ~250 eras against a strategy nobody selected - while the log
   reported "target 1.00*ATR" as though it were configured.

   Since the relabel these two inputs ARE the label definition, so this
   is not a bad trade setting, it is a wrong dataset. ValidateBarrier-
   Inputs() now refuses to start (INIT_FAILED + Alert + an explicit fix)
   on any value that is not an enum member. Members are enumerated rather
   than range-checked because both enums are sparse and carry negative
   sentinels, so no min/max test can tell a legal value from a deleted
   one - which is the entire failure mode. The fallback survives as
   belt-and-braces but now announces itself: a fallback that cannot say
   it fired is indistinguishable from correct behaviour.

2. THE DEPLOYABILITY FLOOR BECAME MATHEMATICALLY UNREACHABLE.

   `tradeableOK` required `dirPrecPct >= baseRatePct`, where baseRatePct
   is Buy+Sell as a share of all bars. At the old exact-pivot target that
   was ~6%, so "beat the base rate" read as "beat chance" and the test
   looked sound. Triple-barrier labels put it at ~83%, so the gate now
   demanded 83% directional precision - impossible by construction.
   Observed live: all four topologies cycling "PLATEAU stage 3 ... nothing
   safe to deploy" at a perfectly healthy 43-45% precision, with no
   checkpoint able to ship however good it got.

   Replaced with ZERO-SKILL precision, max(Buy,Sell)/allBars: exactly the
   score of the degenerate always-call-one-direction model this floor
   exists to reject. Correct at any base rate - ~43% on the current
   labels, ~3% on the old rare-pivot ones. The era line now prints
   "(chance N%, edge +Mpp)" beside the selection score, because 44%
   precision is excellent against a 3% chance level and worthless against
   a 43% one, and reading the first as the second is what made tonight's
   run look better than it was.

3. THE HORIZON IGNORED THE BARRIER GEOMETRY.

   ComputeBarrierHorizonBars() returned the median ZigZag leg, which
   measures how long a ~1 ATR move takes and says nothing about how long
   the CONFIGURED barrier needs. First-passage time out of [-m,+k] scales
   with m*k, so a 1:3 barrier takes ~3x as long as 1:1; the unscaled
   horizon would have timed out most 1:3 trades and pushed Neutral
   straight back up, re-creating the imbalance the relabel removes.
   Now multiplied by slMult*tpMult, calibrated against a real measurement
   rather than assumed: the accidental 1:1 run resolved at horizon 12 with
   only 16.7% timeouts, so the swing median is the right scale at m*k=1.

   Verifiable, not just asserted: the prebuild now counts barriers that
   ended on the VERTICAL barrier and reports them as a share of Neutral.
   Neutral conflates "timed out" with "stopped out" and only the first
   indicts the horizon.

Both builds compile 0 errors / 0 warnings. Forces a retrain - correcting
TP_Mode re-keys the fingerprint (|TB:1:-101 -> |TB:1:3), which is right:
no existing model was trained on the intended target.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 00:30:49 -04:00
AnimateDread
b4a704d309 feat(ai): triple-barrier labels replace exact-pivot ZigZag targets
The 31:1 class imbalance was self-inflicted by the TARGET, not a property
of the market. Labelling only the exact bar where a ZigZag pivot confirms
gave Buy 1164 / Sell 1164 / Neutral 35841, and every correction mechanism
this codebase accumulated sits downstream of that one choice: the
logit-adjusted loss and its range cap, the prior EMA, the +-3.0 output-bias
seed, balanced-accuracy-then-precision selection with its coverage floor,
the recall floor and its catch-22, the alternation gate, NMS, and the four
oversampling designs that collapsed before them.

The reference this engine is built on (references/neuronetworksbook.pdf
ch. 3.1/3.3) also uses ZigZag, but targets the DIRECTION TO THE NEXT
EXTREMUM on every bar - ~50/50 by construction, with no imbalance to
correct at all. It never had this problem because it never asked "is this
the pivot bar".

Labels are now the triple barrier (Lopez de Prado ch. 3), using the EA's
OWN SL_Mode/TP_Mode: does a trade opened at this bar's close reach its
target before its stop, within a horizon. Buy = long resolves, Sell =
short resolves, Neutral = neither. Consequences:

- dir-precision in the era line stops being a proxy and becomes the win
  rate of the strategy under its own exit rules.
- Expected balance ~25/25/50 at the shipped 1:3 (gambler's ruin), i.e.
  ~2:1 instead of 31:1. Measured and logged at the end of the prebuild.
- Spread is charged on both legs, so it is a NET win rate.
- Intrabar ambiguity resolves to the STOP. OHLC cannot order two touches
  inside one bar and the optimistic reading is how a backtested edge
  becomes a live loss.

ZigZag stays as input features (EnableSwingContext) and now also supplies
the vertical barrier: the horizon is the median confirmed leg length,
snapped to a coarse ladder. Derived, not configured, and deliberately kept
out of the filename fingerprint - a filename keyed on a measured quantity
orphans a trained model the moment the measurement moves.

Removed, because the premise died with the old target:
- the alternation gate. Correct for pivot labels (a ZigZag cannot emit two
  same-type pivots in a row, so a repeat was provably a false fire), and
  wrong for barrier labels, which answer each bar independently. It also
  took its worst consequence with it: a one-sided model previously got ONE
  trade per backtest, a hard blocker on marketplace validation.
- SignalClusterWindow now defaults off - it de-duplicated repeats that are
  now real trades. Kept as an opt-in display control.
- LABEL_WINDOW_BARS, the pivot-widening pass, ConfirmedZigZagLabel.
- the era-0 output-bias seed now needs a genuinely dominant class (0.70)
  rather than 0.40; at ~50% Neutral a +-3.0 seed is a distortion, not a
  correction.

Also fixed, both found while wiring the above:

1. RefreshConvergedSignal sized its buffers from a date delta
   (Bars(sym, period, dtStudied, TimeCurrent())). dtStudied is a training
   watermark; in the tester it is loaded from a live-chart save AHEAD of
   the simulated date, so the interval inverted, Bars() returned ~0, and
   the buffer came out at exactly m_historyBars - deep enough for the OHLC
   window and far too shallow for the Donchian-50 / 20-bar-return / SMA
   extension behind it. Inference silently computed DIFFERENT features
   from the ones training learned on, live as well as in the tester. Now
   sized from what the feature builder actually needs.

2. The barrier horizon is resolved on the deployed path too. A deployed
   model never enters Train(), so it never reached the prebuild, and
   OnlineLearnStep reads the horizon as its confirmation delay - left at
   the fallback it would have backpropped bars whose barriers had not
   resolved. Silent lookahead in the one place that writes to a live model.

SL_Mode/TP_Mode join the weights fingerprint: they define the labels now,
so a model trained at 1:3 must never be silently reused at 1:1. This
re-keys every pre-existing model by design - none were trained on this task.

Inference census extended with the vote gate. LongCondition/ShortCondition
open with a readiness check the refresh counters never see; in the tester it
reduces to "the seeded _optcache.nnw must have LOADED", and if it did not,
every vote is hard-zeroed while the model still answers Buy. The old three
counters would have read that as "the model says Neutral" - false, and a
completely different fix. This is the leading candidate for the
zero-direction backtest and the census can now name it in one run.

Both builds compile 0 errors / 0 warnings. Forces a full retrain.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 20:39:49 -04:00
AnimateDread
251e9711dd diag: report |dW| alongside d|W| per layer, and drop two stale log claims
The per-era `dW/W` line measured the change in each layer's weight NORM.
That statistic cannot separate "this layer only shrank under weight decay"
from "this layer moved somewhere useful" - a rotation at constant norm and
pure decay can print the same number.

It matters right now: on SP500 H1 the LSTM layers print a near-constant
~1.05%/era that exactly equals their geometric norm decay over 318 eras
(HYB lstm2 12.966 -> 0.755, monotone, never once up), while a sibling conv
oscillates around a much slower drift. Norm-change can only hint at that.

Now prints norm(d|W|% / |dW|%). Under decay alone the two are equal; any
gradient component adds in quadrature to the second, so a learning layer
shows the second clearly larger. Diagnostic only - no training behaviour
changes, fingerprint untouched.

Also removes two log lines that described machinery deleted in 397b0ea:
the plateau ladder's terminal stage still claimed "after a warm restart
AND full gamma anneal", and the CONVERGED line "across a warm restart and
a full focal-gamma anneal". Both printed on every CONV/LSTM convergence
today. Same failure mode as the oversampling line 397b0ea fixed: a log
that describes what an older version would have done is confirming
evidence for a false hypothesis.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 14:33:29 -04:00
AnimateDread
397b0eac1f refactor(ai): nine class-imbalance inputs down to two
The imbalance section offered nine controls for one job. Audited against the
code, five of them did not do what their names said at the shipped defaults:

  AILogitPriorStrength  DEAD - Inference.mqh's post-hoc prior early-returns
                        whenever the adjusted loss is on, which is default.
  OversampleParity      DEAD in training - Training.mqh gated the replay loop
                        on !useLogitAdjustedLoss (correctly, citing Buda et
                        al. 2018). Live only in the online-learning path.
  EnableMinorityReplay  DEAD as replay. It survived ONLY as a focal-gamma
                        damper - "replay minority bars through pass-2
                        oversampling" was a focal-loss switch.
  ConstrainReplay       DEAD as a cap; it only chose damper 0.125 vs 0.25.
  UseStaticPrior        An exact duplicate of FreezePriorCalibration - the two
                        were OR'd together in the single place either is read.

So they were not five mechanisms fighting; they were one mechanism plus eight
knobs that mostly described machinery that no longer ran. That is worse than
a real conflict, because the log agreed with the names: the label-cache line
printed "reps up to 28x (90% parity) (seeding era 0's class-balance
oversampling)" on every run, describing an oversampling pass that had been
switched off. It is fixed here too - it cost this session a wrong diagnosis.

The one genuine redundancy was focal loss, running at gamma*0.125 alongside
the adjusted loss: two corrections on the same axis, the exact stacking
failure this file already cited Buda et al. for in two other places, damped
by a replay flag whose replay path was itself dead. Removed rather than
re-tuned. The plateau ladder is unaffected - its escape is the learning-rate
warm restart; the gamma anneal beside it only ever stepped toward zero.

WHAT REMAINS is logit-adjusted loss (Menon et al. 2021) plus a prior freeze:

  LogitAdjustTau         0 = off; replaces the separate EnableLogitAdjusted-
                         Loss boolean, since a strength dial where 0 already
                         means off does not need an on/off switch beside it.
  FreezePriorCalibration unchanged.

It is the only one of the six corrections with a consistency guarantee, and
it is consistent for exactly the balanced-error metric checkpoint selection
already ranks on - so the loss and the deploy decision optimize one thing.

The online continual-learning path keeps its own alpha-balanced focal weight,
now as constants pinned to the removed inputs' shipped defaults, so its
behaviour is unchanged. It legitimately needs its own correction:
ApplyLogitAdjustment() only runs inside a training run, so a deployed model
that was reloaded carries no logit offsets and would otherwise stream 31:1
data into itself uncorrected.

The weights-filename fingerprint is BYTE-IDENTICAL. The focal slot was a
double fed to a %d conversion and had always emitted a literal 0; the |MR:
segment is written as the constant its shipped defaults produced. Dropping
either would have re-keyed every model and forced a from-scratch retrain of
the one topology currently converged and trading.

Also removed as orphans: FOCAL_GAMMA_PRESET, MAX_OVERSAMPLE_REPLICAS,
OVERSAMPLE_PARITY_FRACTION, PLATEAU_GAMMA_STEP, and the now-unreachable
"neutralized by prior correction" diagnostic.

Both builds compile 0 errors, 0 warnings. No retrain forced.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 11:46:57 -04:00
AnimateDread
18f63c7a52 feat(ai): report per-layer weight movement each era
Adds "dW/W dense1:0.412(0.31%) conv1:0.088(0.000%) ..." to the era line:
each layer's weight L2 norm and its relative change since the previous era.

Why: a frozen stage and a badly-suited architecture look identical from the
outside. Both give a flat metric and a retreat to the majority class, and
neither the loss, the accuracy nor the per-class recall can tell them apart.
This session cost two full retrain cycles guessing between them - a forget-
gate bias (a real bug, measured, but not the cause of the observed failure)
and a conv receptive field (which turned out to be a regression, not a fix).

A layer sitting at ~0.000% era after era while its neighbours move is
receiving no gradient, and no amount of retraining or hyperparameter work
will change that. A net where every layer moves and the output still
collapses is a genuine architecture or objective problem. The distinction is
one glance at the log instead of a redeploy-and-wait cycle per hypothesis.

Costs one host-side buffer read per layer per era, off the training path.

Both builds compile 0 errors, 0 warnings.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 07:10:09 -04:00
AnimateDread
4eae763849 fix(ai): report the metric actually compared; surface the derived front-end
The plateau/regression line printed balancedOosEra as the current value while
comparing against m_bestBalancedOos, which has held the SELECTION score since
a142749. Two different metrics in one sentence, so HYBRID logged "regressed
from best 14.4% to 34.0%" a hundred times - a regression to a higher number,
which is not a thing. The comparison itself was right (selectionScore, coverage
weighted, genuinely below best); only the print was wrong. 1039ad9 relabelled
these strings but missed that this site passes the wrong variable.

The startup config line had the same shape of gap: it printed the dense taper
and called itself self-verifying while the DERIVED conv and recurrent stages -
the ones that dominate CONV/LSTM/HYBRID - were invisible. It now shows the
width into and out of each front-end stage, and flags the case where the dense
stack is wider than the vector reaching it (a linear fan-out cannot recover
what the bottleneck discarded; it only adds parameters). Flagged, not silently
reshaped - that would re-key trained topologies mid-comparison.

UsesConvStage()/UsesLstmStage() replace HasConvBeforeLstm() as the primitive,
so each subclass declares its composition once and both the capacity budget and
the config line derive from it rather than restating it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-30 15:20:30 -04:00
AnimateDread
1039ad936f feat(ai): measure precision per confidence tier; fix stale metric labels
Two things the 2026-07-30 run exposed.

1. Every user-facing message still called the selection metric "balanced
   accuracy". It has ranked on directional precision since a142749, so
   "CONVERGED ... balanced accuracy 32.5%" was reporting a 32.5%
   PRECISION as if it were macro-recall, while the same era logged an
   actual balanced accuracy of 49%. Two different numbers under one
   name, in the line that announces a deploy. Relabelled at every site,
   including the stage-3 refusal, which still described the per-class
   recall floor that stopped being the gate.

2. Precision is now bucketed by confidence tier and logged per era,
   both per-tier and cumulatively from each tier upward:

     | tier prec T0:19%(410)[>=28%/1204] T1:31%(520)[>=34%/794] ...

   The per-tier number says whether confidence is calibrated to
   correctness at all; if it does not rise T0->T3, raising the floor
   buys nothing and that is the finding. The ">=" number is what a floor
   would actually deliver, with its fire count, so the coverage cost is
   visible in the same line. Tier weights are 25/50/75/100, so for an
   AI-only config Min_Vote_Open maps straight across: 50 = ">=T1",
   75 = ">=T2", 100 = ">=T3".

Bucketing happens at the existing live-fired accounting site, so it
measures exactly the population that trades - not the raw argmax.

Both builds compile 0 errors, 0 warnings. No retrain needed for either.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-30 11:47:15 -04:00
AnimateDread
ce90fc74b6 fix(ai): discount selection precision by coverage shortfall
The 2026-07-30 run caught a bug in the precision-led selection metric
within 8 eras. HYBRID made exactly ONE directional call in era 7, got it
right, scored 100% precision, and locked that in as best-ever. Nothing
can beat 100%, so the checkpoint froze on a single sample and the run
could only burn to the era cap deploying it.

The coverage floor already existed and already blocked that era from
being DEPLOYABLE - but the ranking ignored coverage entirely whenever no
era had qualified yet, which is precisely the phase where the ranking is
the only thing steering the run.

Precision is now discounted by coverage/floor, capped at 1.0. Continuous
rather than a threshold: an era at half the floor scores half its
precision, so coverage and precision both improve rank and neither can
be traded away. Above the floor the credit saturates, so ranking among
genuinely deployable eras is unchanged pure precision.

Also: the startup config line printed "tau 1.00" while every chart was
actually running the capped 0.35 - the effective value depends on the
measured class priors and is not knowable at init. Now reads
"1.00 requested"; ApplyLogitAdjustment still logs the real figure.

Both builds compile 0 errors, 0 warnings.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-30 10:40:36 -04:00
AnimateDread
45b35b3d1d feat(nn): derive dense depth, train on all history, pin the shape in .cfg
Completes the derived-topology work. Three inputs removed.

AIType loses its depth suffix - AI_MLP/AI_CONV/AI_LSTM/AI_HYBRID, five
entries instead of eight. Depth is now derived from the two endpoints
the taper already has to connect (derived first-layer width, output-tied
final width) at a 2x per-layer compression target, clamped [2..5].
Asking a user to pick a layer count while the code derives the widths
those layers taper between was asking for half a decision: at 64 units
tapering to 12, four layers compress by 1.4x per step and five by 1.3x,
so the extra depth bought no abstraction. On the shipping H1/10y default
the derivation lands on 3 layers - the depth that actually won Run 2.

StudyPeriods removed. There is no case for training on less data than
the broker provides at a ~6% directional base rate; the honest
generalization read comes from the OOS holdout, not from withholding
history. Training now starts at the earliest available bar, floored by
MinTrainYear, which answers a different question (excluding dubious
pre-history) and stays.

That required closing the hazard the old code documented: the capacity
budget now MEASURES the symbol's real bar count, and a topology derived
from a measurement would widen as history downloads. Both ends are now
pinned. Every derived value left the weights-filename fingerprint -
keying a filename on a measured quantity means the EA looks for a file
that does not exist, starts from era 0 and orphans a trained model,
silently, because a missing cache is the normal first-run state. The
shape lives in the .cfg instead, where LoadAndCompare now ADOPTS the
four derived fields rather than diffing them; a mismatch there would
discard a fully-trained model over nothing the user did. Two fields
appended to the .cfg for the conv/LSTM stages, length-guarded on read
because FileReadInteger past EOF returns 0 with no error.

ForceHiddenLayers, a compile-time constant like DebuggingMode, pins
depth for diagnostic comparisons. It joins the fingerprint only when
non-zero, so forced depths get their own files - sequential comparisons
only, not simultaneous from one .ex5.

Derived shape, H1/10y defaults (21 features x 20 bars): first layer 64,
3 dense, 8 conv filters, 16 LSTM units. The LSTM block halves from
~58k to ~28k weights.

Both builds compile 0 errors, 0 warnings. Re-keys existing models.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-30 10:05:40 -04:00
AnimateDread
ebf2e73667 fix(ui): unique chart tag, product-grade panel, responsive under load
Three separate reports from one deploy.

1. CONV, LSTM and HYBRID all came back tagged [4109]. The weights
   fingerprint omits the topology type on purpose - the file path already
   separates it (State\CONV\ vs State\LSTM\ vs State\HYB\) and hashing a
   value that is constant within a folder buys nothing while re-keying
   every trained model into a forced retrain. So the files were never at
   risk, but the tag could not do its one job. Prefixing the short id
   makes it unique on the display side only; the hex half still greps
   straight to the .nnw inside the folder the prefix names.

2. The default panel read like a training console. Six lines down to
   three, each answering a question an owner actually has. The deploy
   internals (best score, eras-since-best, ladder stage) were developer
   diagnostics describing a recall floor that no longer decides anything,
   and were already in the era-end journal line. In-sample accuracy left
   the panel too: it grades the model on bars it trained on, so it always
   flatters, and showing it beside the honest number invites reading the
   wrong one. New compile-time DebuggingMode constant - deliberately not
   an input - carries the IS/OOS pair and the resolved model path into
   the journal instead. No extra Inputs row, no extra Market description
   line, no user-reachable firehose.

3. Panel drag and buttons stuttered under training load, exactly as the
   2026-07-26 note raising the chunk budget to 200ms warned they might.
   Backed off to the documented 120ms - worst-case click latency is that
   budget - and the derived topology (~292k weights to ~29k) makes the
   throughput this costs far cheaper than when that note was written.
   Also halved the panel redraw rate to 2.5 Hz: ChartRedraw repaints the
   whole chart, so its cost scales with accumulated arrows, and 5 Hz was
   the larger half of the stutter. Era-end still force-refreshes.

Both builds compile 0 errors, 0 warnings.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-30 09:05:58 -04:00
AnimateDread
a142749a87 feat(ai): rank checkpoints on directional precision, not balanced accuracy
Balanced accuracy is maximized by exactly the model this system must never
deploy. Measured frontier at fixed signal strength, base rate 6.1%:

    tau 0.00 -> calls  0.2% of bars at 27.3% precision, balanced 34.0%
    tau 0.35 -> calls  2.0% of bars at 15.5% precision, balanced 36.3%
    tau 1.00 -> calls 49.6% of bars at  6.4% precision, balanced 53.5%

It rises monotonically as the model calls MORE and is right LESS, because
two of its three terms are directional recalls that a call-everything model
drives to ~95%, while the Neutral term it sacrifices counts for only a
third. The 2026-07-29 run landed exactly there: balanced 58-64% while
calling a direction on ~100% of bars at a 5-7% win rate against a ~6% base
rate. Only the per-class recall floor stopped those deploying - a guard
doing the job the objective should have been doing - and that same guard
also rejected the genuinely useful sparse-but-precise checkpoints.

Ranking is now DIRECTIONAL PRECISION: of the bars called Buy or Sell, how
many were right. That is what a trading edge is. Two anti-degenerate floors
bracket it, since precision alone is trivially maximized by calling almost
nothing: coverage must reach a fraction of the true directional base rate
(derived, not configured - it adapts to any symbol/timeframe/label rule),
and precision must at least beat that base rate.

Against the same frontier the deploy order inverts from
  tau 1.00 > 0.50 > 0.35 > 0.15   (old, worst model first)
to
  tau 0.35 > 0.50 > 1.00          (new; 0.00/0.15 rejected on coverage)

Balanced accuracy is kept in the log as a diagnostic and marked as such, so
a run where the two disagree - the signature of an over-caller - is visible
at a glance. MinRecall no longer decides what ships; it now only drives the
diagnostic recall line and is a candidate for removal.

Both builds compile 0 errors, 0 warnings.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-30 07:13:08 -04:00
AnimateDread
f2ec1edf84 feat(ai): logit-adjusted loss, replacing oversampling and the post-hoc prior
Menon et al. 2021 (ICLR), "Long-tail learning via logit adjustment": add
tau*log(prior_c) to each class logit inside the training gradient. Softmax
CE on adjusted logits is consistent for BALANCED error - the metric
checkpoint selection already ranks on - so the loss and the deploy decision
finally optimize the same thing.

The engine already computed a true softmax + categorical-CE gradient and
wrote it over the per-neuron sigmoid delta, so this is an offset added to
three logits in the two places that gradient is built (backProp scalar path
and backPropOCL). No backend, kernel or DLL change; the forward pass and
every inference path are untouched, which is the point - the network learns
to absorb the offset, so its raw argmax becomes the balanced-optimal
decision with nothing applied at inference.

Replaces rather than stacks. Minority replay is disabled while this is on,
and the post-hoc inference prior is forced off. Stacking is not a
theoretical worry: simulated on the measured 1118/1119/34298 distribution
in the weak-signal regime, plain CE collapses to Neutral (33.4% balanced,
Buy 0%), replay reaches 48.1%, logit adjustment 50.9% with better balance -
and BOTH together score 45.4% with Neutral recall at 0%, worse than either
alone. Buda et al. 2018 predicts exactly that.

Motivation from the six-chart run: every topology took one direction to
~50% recall and abandoned the other, the direction chosen arbitrarily (the
batch-norm control went Buy 1% / Sell 42%, the inverse of the other five).
One era in 1,301 cleared the per-class recall floor.

Fingerprinted conditionally, so the converged 60.7% models on disk keep
their filenames and stay loadable as the fallback.

Both builds compile 0 errors, 0 warnings.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 19:05:14 -04:00
AnimateDread
cc625c827e fix(training): escape the recall-gate catch-22 that let runs decay unchecked
Evidence (MQL5\Logs, SP500 H1, 2026-07-29):

  Perceptron  era  61  Buy 32% Sell 27% Neut 94%  bal 51%
  LSTM        era 160  Buy 16% Sell 11% Neut 98%  bal 42%  (peaked 49% @ era 44)
  Hybrid      era 179  Buy  5% Sell  2% Neut 99%  bal 35%  (peaked 41%)
  CONV        era 228  Buy  2% Sell  4% Neut 99%  bal 35%  (peaked 40% @ era 122)

Every model peaks early then decays monotonically toward Neutral, and nothing
stops it: the restore-best-weights + decay-eta handler is gated on
m_bestPassedRecall, which stays false forever when no checkpoint ever clears the
per-class floor. CONV ran 228 eras with eta pinned at its 0.000300 start. The
plateau ladder cannot end such a run either (stage 3 refuses to deploy without a
recall pass, so it resets ~27 times), making it a 1000-era one-way trip.

The gate's own justification had expired. It was written when the pre-pass
tiebreak was blended-accuracy-only, where "best" really did mean "called Neutral
most confidently". The balanced-selection change replaced that with
`balancedOosEra > m_bestBalancedOos` plus an isFullyCollapsedEra exclusion, so a
Neutral-only era now scores ~33% - the FLOOR of the balanced metric - and cannot
anchor the checkpoint at all. Pre-pass "best" now means "most class-balanced so
far", which is worth defending; and isWorseEra is itself a balanced-accuracy
regression, so it cannot fire merely for trading Neutral calls for Buy/Sell.

The original concern still holds while the best-so-far IS near-collapse, so the
escape is margin-guarded: defend the checkpoint only once balanced accuracy sits
more than BALANCED_WORTH_DEFENDING_MARGIN_PCT (5pp) above the one-class floor of
100/3. Against the run above that engages for all three stuck topologies
(42.3/41.3/50.0 vs a 38.3 threshold) while a genuinely collapsed run still
explores freely.

Two inputs restored to the regime that actually produced a deploy:

- MinRecall 60 -> 40. The one successful auto-deploy in the logs (Hybrid, 28th
  00:50, best balanced 66.0%) ran against a 40% floor. 60 has never been shown
  reachable here - a floor above what the config can reach is the same "target
  set too high" failure the surrounding comment already warns about.

- OversampleParity 60 -> 90. 60 overcorrected. Runs now START Neutral-dominant
  (Buy 0-11% recall at era 1) and call Buy/Sell on 0-4% of bars against a ~6%
  true base rate - under-calling, with no headroom to converge down from. The
  deploying run began at Buy 90% / Sell 36%, 24% of bars called, and settled into
  the floor from above. Raw over-calling is the intended starting condition; live
  calls are base-rate-calibrated by AILogitPriorStrength, which is why the input's
  own note says to judge over-calling by live-fired precision, not raw counts.

Compiles 0 errors, 0 warnings.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 07:51:08 -04:00
AnimateDread
2de93539d4 refactor: split CExpertSignalAIBase implementation by responsibility
ExpertSignalAIBase.mqh was 8216 lines: the class declaration followed by 87
method bodies covering training, labelling, feature extraction, persistence,
chart drawing, online learning, the GA auto-tuner and inference, all in one
file. Train() alone is 1492 lines; a change to arrow drawing meant scrolling
past the era loop.

Moved the bodies into Expert\AIBase\, included at the bottom of the original
after the class declaration:

  Training.mqh        1607  era loop, plateau ladder, checkpoint select, deploy
  Features.mqh        1093  indicator creation + per-bar input feature vector
  ChartUI.mqh          634  arrows, arrow persistence, status panel, cleanup
  Persistence.mqh      492  .stats/.cfg sidecars, CPU-inference validation, copy
  OnlineLearning.mqh   461  live continual learning, EMA shadow, OOS simulator
  Labels.mqh           309  ZigZag pivot labels, async label-cache prebuild
  AutoTune.mqh         275  genetic tuner (population, crossover, halving)
  Inference.mqh        235  softmax, prior calibration, class priors

  ExpertSignalAIBase.mqh  8216 -> 3131 (declaration + topology build only)

This is a pure relocation - verified mechanically, not by eye: HEAD's file
reconstructed from the eight partials plus the surviving remainder is
byte-identical to HEAD, span for span (scratchpad verify_split.py). No
declaration moved, no signature changed, no code rewritten, so behaviour is
unchanged by construction.

Compiles 0 errors, 0 warnings.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 00:42:45 -04:00