Commit graph Warrior_EA/Expert/AIBase/AutoTune.mqh
Author SHA1 Message Date
AnimateDread
40af4a4b5b fix(labels): the geometry scan rewarded the labels it should reject
First run named 3:10 on all four charts, at 2.3x the configured 2:6. That
answer was wrong and the fault was the ranking statistic.

3:10 wants a horizon of ~swingMedian*30 (~320 bars) and gets
BARRIER_HORIZON_MAX. Clamped, most trades never resolve, the unresolved
remainder all lands in Neutral, and H(Y) collapses. The old statistic
divided the excess BY H(Y) - so a collapsing denominator made the most
degenerate label look like the most predictable one. Every geometry from
2:6 upward was already showing the clamped h128, and the two widest
scored highest, which is the fingerprint of the artefact rather than of
signal.

Two fixes:

Rank on the raw excess in nats. Subtracting each geometry's OWN measured
null already removes the class-balance bias, which is the only thing the
normalisation was ever needed for.

Disqualify clamped geometries outright rather than ranking them down. The
deployed EA holds until SL or TP with no bar limit, so a truncated label
trains the model on a question the strategy never asks. They are still
printed, marked '!', so the disqualification is visible instead of a
silent omission - and the scan now says so explicitly when nothing
eligible is left, because "the limit is the feature set, not the target"
is itself the finding in that case.

The scan also reports each geometry's directional share and timeout share
now. A label nobody can trade is not a candidate however well it scores,
and that has to be visible in the same line as the score.

Compiles 0 errors / 0 warnings. Build tag geometry-scan-v2.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 14:46:25 -04:00
AnimateDread
f97ab9f1d6 feat(labels): measure which barrier is predictable at entry, don't guess
The alignment scan settled the shape of the problem: 4.7x more is
knowable 5 bars into a 128-bar window than at the entry the model
actually trades. A 6xATR target reached over 128 bars is decided
overwhelmingly by what happens DURING the window, so whatever the entry
state knows is buried under 128 bars of later noise. That is a property
of the TARGET, and it is why four different architectures all landed on
precision exactly equal to the base rate - no topology can undo it.

So measure the target. For each SL/TP pairing a user can actually select,
relabel the same sampled bars and score how much the SAME features say
about THAT outcome at entry. Seconds, no training, no topology, and it
runs on the diagnostic path that already exists.

Ranked on excess over its OWN null as a share of its OWN H(Y), never on
raw nats: each geometry has a different class balance, hence a different
finite-sample bias and a different amount of information there to find,
so raw MI would rank the most BALANCED label rather than the most
PREDICTABLE one. The break-even win rate m/(m+k) is printed beside each
so the ranking is read next to the bar the model must clear.

Stated in the output because it is the easy thing to get wrong: chance
precision EQUALS break-even at every geometry, so a tighter target does
not hand you expectancy. It buys predictability - less noise piled on top
of what the entry state knows - which is the one thing changing topology
cannot do.

Read-only by construction: it relabels a sampled copy via
TripleBarrierLabel(), never writes the label cache (which belongs to the
configured geometry), and restores the horizon and overrides it borrowed.
The overrides apply only when BOTH are positive, so a half-set pair can
never silently relabel a live run.

Compiles 0 errors / 0 warnings, standard and Market. Build tag
geometry-scan-v1. Redeploy only - no retrain to READ the ranking.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 14:32:04 -04:00
AnimateDread
4443ce85c1 fix(diag): the alignment scan cried misalignment at its own arithmetic
First run came back "WARNING - peak at k=+5, NOT 0 ... a feature/label
misalignment upstream of every topology". That was a false alarm produced
by the diagnostic's own design, and exactly the kind of plausible-looking
output this project has lost days to.

Bar indices are MQL5 SERIES indices - HIGHER index = OLDER bar
(TripleBarrierLabel walks its window as `for(t = idx-1; t >= idx-horizon;
t--)`, decreasing index = forward in time). The two directions therefore
mean opposite things and the scan treated them as symmetric:

  k < 0  label belongs to a NEWER bar, its barrier window opens AFTER the
         features exist. Nothing at bar i can legitimately know it, so a
         peak here is real lookahead and a bug.
  k > 0  label belongs to an OLDER bar, already k bars into its window by
         the time bar i happens - so the features hold the realised first
         k bars of that outcome. MI MUST rise with k. Arithmetic.

Only the k<0 side can indict the pipeline, and on the observed data it is
clean: -5/-3/-2/-1 all sit at or below the k=0 value and the noise floor,
so there is no lookahead - a real negative result, not an absence of
evidence.

The k>0 side is now reported as what it is, a second positive control,
with its gradient as the finding: 0.01881 at k=+5 against 0.00401 at k=0
means ~4.7x more is knowable 5 bars into a 128-bar window than at the
entry the model actually trades on.

Compiles 0 errors / 0 warnings. Build tag mi-align-v2. Redeploy only.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 14:25:26 -04:00
AnimateDread
87c8656b53 diag(autotune): a positive control, and a scan that separates "no signal"
from "signal knocked out of step"

Four architecturally different networks landed on the same precision -
Buy 23-25% against a 25.4% base rate, Sell 19-22% against 22.0% - while
making completely different calls (HYBRID votes Sell on 69% of bars, PAI
on 41%). Precision equal to the base rate is what INDEPENDENCE looks
like, and precision under independence is fixed by the label
distribution, not by the architecture, so all four converging on it is
arithmetic rather than coincidence. Accuracy meanwhile tracks coverage
exactly as independence predicts (31.1/30.3/25.0 predicted vs
31.8/28.9/24.6 observed for PAI/CONV/HYB).

But "no information in the data" and "information destroyed upstream of
every topology" produce that identical picture, and the MI test alone
cannot tell them apart either. Two additions:

POSITIVE CONTROL. Three "measurements" in this codebase have turned out
to be silent no-ops that produced plausible numbers - the MI scorer
reading an array nobody filled, the eval-mode guard that switched off the
imbalance correction, the alternation gate whose premise was never true.
So the estimator now has to prove it responds to a signal known to be
present before any floor reading is believed: the label of a neighbouring
sample row, ~19 bars away and far inside the 128-bar barrier horizon, so
the two outcome windows overlap heavily and MUST be associated. Same
binning, same estimator. Near the floor => every MI figure is void.

ALIGNMENT SCAN. Re-scores against the label taken from bar i+k for k in
-5..+5. A peak at k != 0 is a feature/label misalignment - an off-by-one
in the label index, a horizon applied to the wrong bar, a feature window
that lags what it claims - which would destroy the information before any
topology saw it and would look identical in every accuracy number this EA
prints. A flat profile says the features simply do not carry this target.
The sampled range is trimmed by |k| at both ends so a shift is measured
rather than an edge effect, and both bars must carry a real label.

Also: BuildMiSample publishes its stride instead of the report
recomputing that arithmetic (it would drift), and the control sizes its
buffers from its own sample count rather than the caller's.

Compiles 0 errors / 0 warnings, standard and Market.
Build tag mi-control-align-v1. Redeploy only - no retrain, no model
deletion; the diagnostic runs on resumed models.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 14:12:10 -04:00
AnimateDread
9a5f645dc3 diag(autotune): stop making the feature test cost a trained model
The permutation test lived inside TuneIndicatorsByFilter, which is gated
on era 0 - correctly, because re-running the SWEEP would change the input
vector out from under weights already fitted to the old one. But the test
itself reads cached features and writes nothing, so none of that applies
to it, and the gate meant the only way to see the answer on a running
model was to delete the model. Today that price was PAI's 45 trained eras
and CONV's 31, spent to re-ask a read-only question.

Split into ReportFeatureLabelInformation(), called from the sweep when it
runs and directly when it does not - a resumed model, a disabled tuner,
nothing tunable. Once per attach either way.

Compiles 0 errors / 0 warnings. Build tag permtest-v2.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 14:01:32 -04:00
AnimateDread
9920754dec diag(autotune): five permutations was still a coin flip - use a real test
The 5-draw z-score shipped an hour ago disproved itself on its first run.
All four charts scored the IDENTICAL 0.00401 nats on identical features
and identical labels - and reported z of +1.3, +2.0, +4.0 and +4.7. Two
"AT THE NOISE FLOOR", two "a real association", same data. The entire
swing came from estimating the null's spread from five draws, where the
standard deviation of the standard-deviation estimate is ~35%: the
denominator was noisier than the effect it was judging.

Replaced with an empirical permutation test. 200 draws, p counted by rank
with the +1/(B+1) correction (Phipson & Smyth 2010) so p is never
reported as exactly zero - no normality assumption and no spread to
estimate. The strongest single column is tested against the null
distribution OF THE MAXIMUM, which corrects for scoring 26 features at
once by construction and is far less conservative than Bonferroni.

Affordable because BuildMiSample is now split out of ScoreCurrentParamsByMI
and runs ONCE for the whole test - every draw reuses that sample and costs
a relabel plus 26 histogram passes, not 2000 feature extractions. The
coordinate sweep still calls the combined form, which is correct there:
each candidate changes the indicator settings, so its features really do
have to be re-extracted.

The verdict line keeps both questions apart and prints both answers: the
p-value for "is it real", the excess as a percentage of H(Y) for "is it
big enough to trade". At n=2000 those can disagree, and collapsing them
into one word is how a worthless effect gets called a discovery.

Compiles 0 errors / 0 warnings. Build tag permtest-v1.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 13:45:46 -04:00
AnimateDread
12a1fbd133 diag(autotune): one label shuffle cannot settle the no-edge question
The permutation baseline added in 018afb1 came back on all four charts as
0.00401 nats against floors of 0.00267 / 0.00298 / 0.00318 - three draws
whose spread is as wide as the excess being judged, because one shuffle
is one sample from the null, not the null. That is not enough to retire a
topology on.

Now MI_NOISE_PERMUTATIONS draws, reported as mean +/- sd with a z-score,
plus two numbers the mean over 26 columns cannot express:

  - the STRONGEST single feature's MI, against its own shuffled value.
    One informative column among 25 useless ones is precisely the case
    the mean hides, and precisely the case worth finding.
  - the excess as a percentage of H(Y). At these sample sizes a z-score
    can be comfortably significant while the effect is worthless, so
    "is it real" and "is it big enough to matter" are asked separately
    and answered separately.

The verdict line also now states the measure's limit every time rather
than only when the news is bad: this is a MARGINAL, PER-BAR statistic and
the network reads m_historyBars bars jointly, so it can prove signal
exists but never that it does not. It rules out a per-feature edge - and
therefore any indicator retuning - not an edge that lives in a
combination or across time.

Compiles 0 errors / 0 warnings, standard and Market.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 13:32:12 -04:00
AnimateDread
018afb1ba9 fix(autotune): MI scorer read an array nobody filled; add the permutation floor
THE TUNER WAS A SILENT NO-OP. Every chart logged

  auto-tune complete - 17 candidate settings scored in ~139s,
  feature/label mutual information 0.0000 -> 0.0000 nats (no improvement)

0.0000 is not a weak result, it is a broken measurement: finite-sample MI
is biased UPWARD, so even pure noise scores above zero. Cause:
ScoreCurrentParamsByMI called BufferTempDataCompute(), which APPENDS the
bar's features to TempData and never touches m_featureCache - only the
caching wrapper BufferTempData() writes that array. It then read
m_featureCache, which ReInitADIndicators had just invalidated. Every
column came back constant, FeatureColumnMI returned 0 for all of them,
and all 17 candidates tied at exactly zero. 139 s per chart to return the
settings it started with.

Now reads the values back out of TempData, where they actually land. And
an exactly-zero best score is called out as a fault rather than reported
as "no improvement", because that is what it is.

ADDED: a PERMUTATION BASELINE, which is the diagnostic this project has
been missing. MI's finite-sample bias is ~(bins-1)(classes-1)/(2n) nats -
at these sample sizes the same order as any real edge in this domain - so
a raw MI figure is uninterpretable on its own. Shuffling the labels
destroys every genuine association while leaving sample size, binning and
class proportions intact, so the score it produces IS this dataset's
noise floor, measured rather than approximated. The log now reads

  feature/label information - X nats against a shuffled-label floor of Y

and says outright whether the features carry usable information about the
target. It needs no training, no topology and no convergence, so unlike
every accuracy number in this codebase it cannot be confounded by an
optimizer or an objective. If the score sits on the floor, no change of
architecture can help - which is the question the last three days of
zero-edge results have been circling.

DEPLOY FLOOR: `dirPrecPct > chancePrecPct` passed anything above chance by
any amount. At ~11,000 directional calls the standard error of the
precision estimate is ~0.4pp, so that gate was accepting sub-one-sigma
noise - the perceptron deployed at edge +0pp on 2026-08-01. Now requires
EDGE_MIN_SIGMAS (2.0) standard errors above chance, computed from the
actual call count, so the bar scales with the evidence instead of needing
a hand-picked constant.

Recorded with it, because it is why chance is the right reference at all:
under a driftless random walk P(touch +k*ATR before -m*ATR) = m/(m+k),
and the break-even win rate for a k:m reward:risk trade is ALSO m/(m+k).
The label's own base rate IS the break-even rate, at every SL/TP setting.
So "beats chance" and "is profitable" are the same test, and no choice of
SL/TP can manufacture an edge - only prediction can.

Both builds compile 0 errors / 0 warnings.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 13:05:50 -04:00
AnimateDread
6db0519472 perf(autotune): replace the genetic search with a filter score - hours to seconds
MEASURED COST OF THE GA, which is what retired it. Per generation:
  rung 0: 8 cand x 3 seeds x  3 eras =  72 eras
  rung 1: 4 cand x 3 seeds x  8 eras =  96
  rung 2: 2 cand x 3 seeds x 20 eras = 120
  = 288 eras/generation x 4 generations = 1152 eras BEFORE the winner's
real training began. Against the observed era times on SP500 H1:

  PAI     29.1 s/era  ->   9.3 h   (matches the observed 00:37 -> 09:22)
  CONV    41.3 s/era  ->  13.2 h
  LSTM   150.4 s/era  ->  48.1 h
  HYBRID 154.6 s/era  ->  49.5 h

Two days to tune is not a first-run experience, and it is the phase in
which the panel goes quiet, which is what made it look like a hang.

It also bought nothing. The space is 90 points (10 MA periods x 9 MA
types), so 1152 evaluations revisited each point ~13 times; and rungs of
3 and 8 eras cannot separate two MA periods at all. The 2026-08-01 run
proves it: every finalist scored 25.0-25.9% balanced accuracy - below the
33.3% one-class floor, i.e. indistinguishable noise - and the search then
"deployed the winner" of that.

THE ERROR WAS THE SCORING FUNCTION, not its constants. Using a full
training run to choose a feature's period is a wrapper method paying
wrapper prices for a decision that does not need one. The reference book
does not do this: ch. 3.3 selects inputs by measuring each candidate
indicator's CORRELATION with the target and dropping the ones with none,
with no network involved.

So: rank candidates by the MUTUAL INFORMATION between the resulting
feature vector and the triple-barrier label. MI rather than correlation
because the label is 3-class categorical and the features are not
monotonically related to it. Equal-FREQUENCY binning (rank-based),
because these features are ATR-normalised and heavy-tailed - fixed-width
bins put nearly everything in one bucket and report ~0 information for a
genuinely useful feature.

Scoring is arithmetic over the feature cache, so it costs seconds and its
cost is independent of topology: LSTM now tunes as fast as the MLP.
Coordinate sweep, not product sweep - cost is the SUM of per-parameter
candidate counts, so enabling every indicator stays affordable - with a
second pass that breaks early once nothing moves.

Sampling is IS-ONLY. Letting the OOS window influence which indicator
settings ship would mean the holdout had been used for selection and had
stopped being a holdout.

HONEST LIMIT, recorded because it is the price: MI is marginal, so a
parameter that only pays off in combination with another can be missed
(Guyon & Elisseeff 2003, filter vs wrapper). Given the wrapper it
replaces was ranking pure noise at 48 h a run, this is strictly better.

Deleted with it: GaRungEras/GaExtract/GaStore/GaMutate/GaRandomCandidate/
GaBlockCrossover/GaSortAliveByScoreDesc/GaBreedNextGeneration, 14 m_ga*
members, the GA_*/TUNE_POP_* constants, and ComputeTuneTrialBudget.

AND m_evalMode/m_evalEraBudget, because nothing set them any more - 28
read sites all permanently inert. That is not a tidy-up: the `if
(!m_evalMode)` guard on UpdateClassPriors is exactly what silently
disabled the imbalance correction for entire runs two commits ago. Dead
machinery that still reads like live machinery is this codebase's most
expensive recurring bug, and leaving 28 more instances of it would have
been indefensible.

The panel's tuning-progress state goes too - tuning no longer takes long
enough to need one.

Both builds compile 0 errors / 0 warnings.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 10:29:31 -04:00
AnimateDread
f48bc93f9b refactor(inputs): 96 -> 70 inputs; remove two untested/unusable filter modules
Every removal below is FINGERPRINT-NEUTRAL by construction: each retired
input is pinned to the exact value it already shipped with, so running
models keep their filenames and resume rather than restarting at era 0.
Verified field by field against BuildConfigFingerprint.

Removed as inputs, kept as pinned constants (the value was never a
preference the user had a basis to change):

- OutputNeuronsCount. The regression head predicts a continuous quantity
  the triple-barrier label does not contain; the target is an EVENT, so
  the right output is its probability. The regression code paths stay
  implemented and dormant - they cost nothing and removing them would
  touch every scoring path at once.
- MinRecall. A safety floor, not a preference, and the only direction a
  user can move it is the harmful one: raising it past what the config
  reaches yields NO model, not a better one (observed repeatedly at 60).
- SwingConfirmationBars. Stopped gating the labels with the relabel, but
  is STILL load-bearing for the swing-context input features - it is the
  ZigZag repainting embargo, and without it those 9 features read a leg
  the live bar could not have had yet. Pinned, not deleted.
- MaxErasPerRun (runaway backstop, never reached in a healthy run),
  FreezePriorCalibration (unanswerable by a user; near-balanced labels
  make the priors stable anyway), VerboseMode (developer view, joins
  DebuggingMode), MACD/Ichimoku periods x6 (both indicators ship
  disabled, and as optimizer dimensions they are pure overfitting
  surface - the AI auto-tuner is the supported way to move them).
- SignalClusterWindow -> 3, no longer an input. Barrier labels make
  consecutive setups real, which argued for 0; it is not 0 because on D1+
  a 6-bar window spans over a week and two arrows a day apart on a
  weekly-scale move are one event. 3 splits it correctly by timeframe.
- EnableOnlineLearning -> ON. Adapting to a changing market is what keeps
  a months-attached model from going stale, and the rolling-accuracy
  freeze is what makes it safe. See the caveat noted in the handoff: it
  had not been forward-tested on a live feed when this became default.

Removed entirely:

- Intraday Time Filter (5 inputs + Signals/SignalITF.mqh). Two of its
  five inputs were raw BITMASKS, which is an implementation detail
  exposed as a control. The job is covered three times over by things
  that are declarative or that learn: the session filter, the
  time-of-day/day-of-week input features (the network discovers which
  hours are good rather than being told), and the journal's time buckets.
- Market Depth Filter (5 inputs + Signals/SignalMarketDepth.mqh, plus
  its OnInit probe and OnDeinit release). It needs real level-2 data
  that this broker - and most retail MT5 brokers - do not provide, so
  the module has never once executed against real data. Shipping four
  tuning dropdowns for an untested path is worse than shipping nothing:
  the only users who could enable it would be its first-ever testers,
  live. If DOM returns it should be a FEATURE fed to the network, not a
  rule-based veto with hand-tuned thresholds - imbalance is data.
- IndicatorTuneTrials, replaced by ComputeTuneTrialBudget(). The useful
  budget depends on how many parameters are actually being searched,
  which depends on which features are enabled - so one number meant
  wildly different things run to run. The shipped 32 was ~10 candidates
  per dimension against one enabled indicator (wasteful: each costs
  GA_SEEDS full training runs) and under one per dimension against all
  nine (blind). Now population ~ 4 x active dimensions, clamped [8,64],
  with CADIndicatorTuner::ActiveDimensions() defined immediately above
  PerturbRandom() so the two cannot drift apart.
- Six orphaned enums (TUNE_TRIALS_PRESET, DOM_*, ENTRY_HOUR_OF_DAY,
  TIME_FILTER_DAY_OF_WEEK), 81 lines.

Other UX:

- SL_ATR_x1 / TP_ATR_x3 now carry the "(classic)" default marker every
  other preset enum in the file already used. Nothing in the SL/TP
  dropdowns previously told a user which pair was the shipped default -
  which matters far more since the relabel, because those two define the
  labels and changing either forces a retrain.
- Neural Network section moved directly ABOVE AI Input Features: choose
  the architecture, then choose what it sees. NN Optimizer / Performance
  stays last - the Adam/Sgd inputs are declared in AI/Network.mqh and
  render immediately after that divider.
- News feature + window moved to the end of the AI feature list, below
  Wyckoff Bar Inversion.
- Dropped "(0-100)" from Min vote to open - it is an enum, not a number.

Both builds compile 0 errors / 0 warnings. No retrain forced.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 21:22:02 -04:00
AnimateDread
2de93539d4 refactor: split CExpertSignalAIBase implementation by responsibility
ExpertSignalAIBase.mqh was 8216 lines: the class declaration followed by 87
method bodies covering training, labelling, feature extraction, persistence,
chart drawing, online learning, the GA auto-tuner and inference, all in one
file. Train() alone is 1492 lines; a change to arrow drawing meant scrolling
past the era loop.

Moved the bodies into Expert\AIBase\, included at the bottom of the original
after the class declaration:

  Training.mqh        1607  era loop, plateau ladder, checkpoint select, deploy
  Features.mqh        1093  indicator creation + per-bar input feature vector
  ChartUI.mqh          634  arrows, arrow persistence, status panel, cleanup
  Persistence.mqh      492  .stats/.cfg sidecars, CPU-inference validation, copy
  OnlineLearning.mqh   461  live continual learning, EMA shadow, OOS simulator
  Labels.mqh           309  ZigZag pivot labels, async label-cache prebuild
  AutoTune.mqh         275  genetic tuner (population, crossover, halving)
  Inference.mqh        235  softmax, prior calibration, class priors

  ExpertSignalAIBase.mqh  8216 -> 3131 (declaration + topology build only)

This is a pure relocation - verified mechanically, not by eye: HEAD's file
reconstructed from the eight partials plus the surviving remainder is
byte-identical to HEAD, span for span (scratchpad verify_split.py). No
declaration moved, no signature changed, no code rewritten, so behaviour is
unchanged by construction.

Compiles 0 errors, 0 warnings.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 00:42:45 -04:00