Commit graph Warrior_EA/Expert/AIBase/AutoTune.mqh
Author SHA1 Message Date
AnimateDread
b3b7e7bceb fix: excursion window must not depend on the barrier it sizes
DIRECTION IS NOT THERE, and this run is what establishes it. Three symbols:

  raw ASYMMETRY   clears on all three (p=0.0199 / 0.0050 / 0.0050)
  norm ASYMMETRY  collapses on all three (p=0.3433 / 0.5075 / 0.2736),
                  USDCAD landing BELOW its own null
  RANGE control   strengthens to 3-5x its null everywhere

Divide sigma out and the apparent directional signal vanishes entirely. What
cleared was volatility leaking through an unnormalised difference. Note this
would have passed any replication test: three instruments at p=0.005 is exactly
the evidence one would accept before committing to a rebuild, and the confound
reproduces perfectly. Replication was never going to catch it - only the
normalisation could.

Two defects of mine, both surfaced by the same run.

1. THE GEOMETRY DERIVATION WAS DIVERGING, NOT CONVERGING. It produced a
   14.57*ATR stop and a 29.14*ATR target that only 5.7% of bars ever reach.
   Excursions were measured over the barrier horizon; the horizon scales with
   the target; the target is a quantile of the excursions - so target ->
   horizon -> excursions -> target ran away, and "settled" only because the
   horizon ladder caps at 384 bars. A saturated runaway, which the iteration
   guard could not catch because it watches for OSCILLATION.
   Fixed at the root: excursions now accumulate only over m_swingMedianBars -
   the UNSCALED median ZigZag leg, a property of the instrument that owes
   nothing to the barrier. The barrier walk still runs the full horizon,
   because that is how long the trade is held; only the MEASUREMENT used to
   size the barrier is confined to a geometry-independent window.
   (The Min_Risk_Reward_Ratio warning fired correctly and is what flagged it -
   the diagnostic worked while the derivation behind it did not.)

2. THE CONFOUND VERDICT WAS UNREACHABLE. `sizeCleared && !asymCleared` was
   tested first and is true whenever size clears - i.e. always - so the branch
   that NAMES the volatility confound never printed; all three symbols showed
   the generic size-not-direction message instead. Verdict chain rewritten with
   the specific case first, and the dangling elses my first patch introduced
   removed.

FORCES A FULL RETRAIN (the excursion window changes every derived barrier).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-07 13:57:23 -04:00
AnimateDread
32ffeb99f3 fix: normalise the asymmetry target - the raw one is confounded by volatility
Three symbols ran the excursion test. RANGE/UP/DOWN cleared on all three;
raw ASYMMETRY cleared on EURUSD and USDCAD at p=0.0050 and not on SP500
(p=0.1045). That looked like the first directional signal this project has
found. It probably is not, and the test as built could not tell.

(up-dn) IS NOT SCALE-FREE. If sigma is predictable - and RANGE clears at ~4x its
null on every instrument - and the directional part is symmetric noise eps, then
up-dn ~ sigma*eps, so a large sigma pushes the value into BOTH outer terciles. A
pure volatility predictor scores positive MI against a 3-bin (up-dn) while
carrying no directional information at all. Crucially that confound REPLICATES,
so reproducing on two instruments is not evidence against it - and the effect
sizes fit it: asymmetry runs 1.3-1.6x its null where RANGE runs ~4x, and carries
~0.1% of the target's entropy against RANGE's ~0.9%. That is the shape of a
leaked fraction of the volatility signal, not an independent one.

So add (up-dn)/(up+dn): bounded in [-1,+1], volatility divided out, and the only
target a directional claim may rest on. The verdict now separates the cases and
NAMES the confound when raw clears while normalised does not, instead of
reporting the raw line as a finding.

Two bugs of mine in the same block, both caught by output rather than review:

  - The derived-geometry line had a MISORDERED argument list: it printed
    "stop 25.00*ATR (q3 of adverse travel)" - the quantile percentage as the
    multiple and the multiple as the quantile. Real values were 2.61 stop /
    8.03 target. A 25*ATR stop is absurd on its face, which is why it was seen.
  - THE STOP QUANTILE WAS BACKWARDS, and this one changes labels. It was 0.25
    "so ordinary noise does not reach it", but q25 means 75% of bars EXCEED the
    stop - hit three times in four. The printed reachability said exactly that
    ("stop on 75.0% of bars"). Now 0.75. A quantile is a threshold, not a rate.
    This is the entire reason reachability is measured and printed rather than
    assumed.

Also raises BARRIER_DERIVE_MAX_PASSES 3 -> 5: SP500 did not settle in 3 (stop
still moving ~14% per pass) while EURUSD and USDCAD converged on pass 2. And
bounds both quantile indices with MathMin(..., n-1) so q=1.0 cannot run off the
end of the sorted array.

The geometry from the previous run is NOT usable and the asymmetry result is
unresolved, not established. Both are decided by the next run.

FORCES A FULL RETRAIN (the stop quantile changes every label).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-07 13:04:13 -04:00
AnimateDread
a7701f032b feat: derive the ATR multiples from measured excursions - no hardcoded geometry
The barrier was still two constants. SL_Mode/TP_Mode left the Inputs tab in
3482b6c, but the fallback was a hardcoded 2:6 and the geometry scan only ever
chose from a hardcoded grid {2,3} x {2,3,4,6,8,10}. Picking the least-bad of
eleven guesses is not deriving anything.

WHY THE SCAN WAS THE WRONG INSTRUMENT, now measurable rather than argued. It
ranks pairings by how predictable their OUTCOME is - a question about direction.
The excursion test (2c78f3b) ran on SP500 H1 and direction is the one thing
absent: ASYMMETRY p=0.0846, against RANGE/UP/DOWN all at p=0.0050, with RANGE
scoring 0.01345 vs a 0.00343 null - 4x, where the barrier label sits at 1.01x.
Hence the scan failing its own gate on every run, and its "winner" wandering
2:8 -> 3:8 -> 2:8 -> 2:4 across four runs of the same data. Excursion SIZE is
strongly measurable, so derive the geometry from that instead.

  stop   = q25 of measured ADVERSE travel   (ordinary noise does not reach it)
  target = q50 of measured FAVOURABLE travel (reached ~half the time, by
           construction, inside the horizon)

Continuous, in ATR units, superseding the enum multiples. Reachability ("target
on X% of bars, stop on Y%") and the implied break-even are printed so the choice
is auditable rather than trusted.

FIXED-POINT ITERATION, not one-shot. ComputeBarrierHorizonBars scales the
horizon with the target (first-passage time grows with the band) and the
excursions are measured OVER the horizon, so target -> horizon -> excursions ->
target is a real loop - deriving once sizes the target from travel measured
under the PREVIOUS horizon. Re-measures until the multiples move <5%, capped at
3 passes, and says so if it does not settle.

Does NOT create expectancy, and the log says as much: chance precision equals
break-even at every geometry (m/(m+k) on both sides). It buys a target the
market reaches and a stop that survives noise. Where Min_Risk_Reward_Ratio
forces a target the market rarely reaches, it WARNS rather than overriding -
the ratio is the user's risk policy, so the honest move is to state its cost.
That is the collision that once rejected 100% of setups.

Pinned in the .cfg as doubles appended AFTER this morning's two ints, so .cfg
files written earlier today still load (their length guard finds no doubles) and
a model that carries them was trained on them and never re-derives.

Also fixes a message from e5ceed6 that claimed "this model resumed from disk"
unconditionally - it printed above a "seeding era 0" line on a brand-new model,
because the branch fires whenever the cache is not built, which is equally true
before a fresh model's first prebuild. A diagnostic that misreports its own
trigger is worse than one that says nothing: it gets quoted back as evidence.

FORCES A FULL RETRAIN (labels change).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-07 12:06:25 -04:00
AnimateDread
2c78f3b90d diag: is "optimal SL/TP" learnable? Score the features against excursions
Proposed direction: train the net to predict entry/SL/TP that maximise return
and minimise drawdown, rather than to classify direction. Before rebuilding a
head, measure whether the target is learnable at all.

That question splits into two that behave nothing alike:
  HOW FAR price travels (MFE/MAE) - essentially volatility, and volatility
    clustering is about the most robust regularity in markets.
  WHICH WAY it goes first (the asymmetry) - direction, which is what every
    noise-floor verdict in this project has been about.
Expectancy comes ONLY from the second. The first buys position sizing and
drawdown control - worth having under prop-firm limits, but not an edge: exit
management on RANDOM entries already moved the payoff ratio 0.92 -> 5.72 with
expectancy FLAT.

Crucially this is NOT already answered. Every MI figure here scored the
triple-barrier label, i.e. one specific question at one fixed geometry. A
noise-floor result there says nothing about whether excursion MAGNITUDE is
learnable - different target, different answer.

Four targets, and the verdict is the CONTRAST, printed explicitly because the
dangerous misreading of "UP clears" is "we can predict profitable trades":
  RANGE (up+dn)  - realised volatility, included as a POSITIVE CONTROL that
                   SHOULD clear. Every prior verdict here lacked a control
                   expected to pass; a range target at the floor indicts the
                   measurement, not the market.
  UP / DOWN      - MFE / MAE.
  ASYMMETRY      - up-dn, the only one that can pay.

Collected inside the walk the label already does (one max, one min per bar).
The early-out when both barriers resolved is GONE: it would have truncated the
excursions at whichever bar tripped the last barrier, making the measurement a
function of the CURRENT SL/TP - the circularity this is trying to escape. The
loop was already bounded by the horizon, so only the average cost moves.

Discretised into 3 EQUAL-FREQUENCY bins, so every downstream piece (block
permutation, null, p-value) is reused unchanged. Equal-frequency because MFE is
fat-tailed and fixed-width bins would put nearly every row in bin 0; it also
pins H(Y) at ln(3)=1.099 for all four, making them comparable to each other and
to the barrier label's ~1.02 instead of confounded by class balance.

Two bugs fixed in this code before it ever ran, both of which would have
produced a plausible quiet wrong answer rather than an error:
  - TripleBarrierLabel early-returns on invalid ATR/close BEFORE the point the
    accumulators were reset, so one bar's excursions would be cached under
    another bar's index. Cleared at the top now, ahead of every return.
  - An unresolvable bar is still flagged as labelled but carries excursions of
    exactly 0. Under equal-frequency binning a block of identical zeros drags
    the lowest cut onto zero and a third of the sample lands in one
    uninformative bin - a depressed score that reads as "not predictable", a
    false negative in the direction that would wrongly kill the idea. Rows
    where both excursions are zero are dropped; price cannot travel zero both
    ways over a whole horizon.

Read-only diagnostic. No topology or label change: no retrain of its own.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-07 10:22:41 -04:00
AnimateDread
3482b6c238 feat: entry/SL/TP stop being inputs - the barrier geometry is measured
Three enums left the Inputs tab. They were three things a user had to pick and,
in the tester, three more axes for a genetic optimization to overfit.

Entry_Multiplier is pinned to MARKET. Its pending modes place the entry at a
LEVEL while the rest of the pipeline measures from the bar open - the exact
mismatch that manufactured the +0.097 R "retail fade" result later retracted as
a fill artifact. This codebase's fill model cannot honestly simulate a pending
entry, so it is no longer offered.

SL_Mode/TP_Mode become a STARTING pair. ReportBarrierGeometryScan now ADOPTS its
winner instead of printing "set SL_Mode/TP_Mode to X and retrain":

  - only when it clears the family-wise gate from 04ee2e1 (beat the null of the
    MAXIMUM, not merely the incumbent). This is why that gate had to land first:
    without it, removing the inputs would hand a noise-picked geometry direct
    control over the training target with no human in the loop - strictly worse
    than the input it replaced. On SP500 H1 today it does NOT clear (p=0.1463),
    so 2:6 is what you get - now chosen by measurement rather than assumed.
  - only at m_eraCount == 0. Relabelling a partly-trained net moves the target
    out from under weights already fitted to the old one.

THE GEOMETRY LEFT THE WEIGHTS-FILENAME HASH, because it is now measured. Same
rule that moved the horizon and the derived topology values out: a filename
keyed on a measured quantity changes the moment the measurement does - a few
more bars shift which pairing wins - and the EA then looks for a file that does
not exist, starts from era 0 and orphans a trained model silently. It is PINNED
IN THE .cfg instead: appended at the end (the only backward-safe change),
length-guarded like the 2026-07-30 derived pair, and ADOPTED on load rather than
compared, so a trained model keeps the barriers it actually learned and never
re-measures.

Two traps closed while wiring it, neither of which announces itself:

  - m_barrierHorizonResolved latches the horizon ONCE PER PROCESS. Adopting 2:8
    (wants ~192 bars) after it settled for 2:6 (128) would label the new target
    against the old ceiling - the truncation fixed in 168422f, where every model
    learned "target within 128 bars" while the EA holds to SL/TP. It lands in
    Neutral, not in the timeout counter watching for it. Unlatched on adoption,
    along with the label cache the old barriers filled.
  - the .cfg adopt runs at init, before the horizon latches and before any label
    is computed, so a resumed model has its pinned pair in place first. Verified,
    not assumed.

FORCES A FULL RETRAIN: the fingerprint change orphans every existing .nnw.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-07 09:39:30 -04:00
AnimateDread
9e1c72aacc fix: make the indicator tuner actually measure, and gate what it installs
ROOT CAUSE of the zero spread measured on SP500 H1 2026-08-07 (all 17 candidates
returned exactly 0.00359 nats): the tune loop re-inits the indicators and then
scores, with no RefreshData() between.

ReInitADIndicators() does its part - Create() builds a NEW handle carrying the
new parameters, and the feature cache is flagged stale so features really are
recomputed. But BufferTempDataCompute() reads the CIndicatorBuffer objects, and
only Refresh() copies data out of a handle into those. So every candidate was
scored on values still held from the PREVIOUS handle. My earlier guess in the
diagnostic ("suspect the feature cache") was wrong: the cache invalidation works.

Two things land together, because neither is safe alone:

1. RefreshData() after the re-init, so a candidate is scored on its own features.
2. A SELECTION GATE on the install. bestScore is a MAXIMUM over candidates, and
   the maximum of N draws from a null beats its incumbent almost every time - so
   "it beat the incumbent" installs noise. This selector is the highest-stakes of
   the three found in this audit because it ACTS: it overwrites the user's
   configured indicator settings and forces BuildFreshTopology(), so the network
   then trains on whatever the noise picked. Fixing (1) without (2) would have
   made a dormant bug actively harmful.

The gate draws the winner's own permutation null once, then corrects the p-value
for having chosen it out of N with Sidak: p_family = 1 - (1-p)^N. Sidak rather
than the max-of-N resample used by the geometry scan because each candidate here
has a DIFFERENT feature set, so their draws cannot be pooled; Sidak needs only
the one null. Exact under independence, mildly anti-conservative under positive
dependence - stated in the comment rather than hidden. A rejected winner restores
the configured settings, which best[] cannot do since the descent mutates it.

Also reports the least-ready tunable handle's BarsCalculated(). IndicatorCreate()
calculates asynchronously, so if the spread is STILL zero the handles simply are
not done and the tuner needs to yield between candidates rather than score them
back to back - a state machine like the label prebuild. That distinction is now
readable from the log instead of requiring another guess.

No input, topology or label change: no retrain.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-07 09:31:06 -04:00
AnimateDread
e5ceed6466 fix: MI diagnostics never ran on a resumed model - the stated intent was never achieved
A comment above the diagnostic branch says it "runs even when the sweep does
not: on a resumed model ... tying it to that gate meant the only way to see the
answer on a running model was to delete the model."

It does not. Moving the diagnostic out of the tuner's gate left it behind
m_labelCachePrebuilt, which has the same effect: the eager label pre-scan runs
only on a FRESH start, because a net loaded from disk labels lazily per bar. So
on a resumed model the flag is false forever and the whole MI block - headline,
positive control, alignment scan, lag profile, geometry scan, winner test, and
the auto-tune line - silently never runs.

Measured on SP500 H1 2026-08-07: attached at era 271, still nothing by era 314,
zero MI lines in the day's log, and the only "label cache pre-built" entry
predates the attach. It also explains the shape of every capture on 08-05/06:
each one came directly after a weights reset. The situation the comment was
written to eliminate is exactly the situation that persisted.

So drive the pre-scan when it is the only thing missing. Safe on a trained net:
its one fresh-net side effect, pushing the output-layer bias toward the dominant
class, is already gated on m_eraCount == 0, and the advance gate in Train() sits
ABOVE if(!m_trainRunActive), so the era loop keeps its state - training pauses
for the scan (~1s at 38k bars) and continues from where it was, not from 0.
Announced only on a start that actually armed, since StartLabelCachePrebuild()
returns unarmed when history is not ready and is retried per bar event.

NOT sampled from the lazily-filled cache instead: BuildMiSample skips bars with
no cached label, so that would score whichever subset training happened to have
visited - a biased subsample presented as a measurement, which is the failure
this diagnostic exists to catch.

Also corrects a claim in 0d58923's comment. It argued four consecutive "no
improvement" runs were ~1-in-100,000 evidence the indicator tuner is inert, by
multiplying 5.6% across four runs. They are not independent trials: the MI
scorer is deterministic and all four covered nearly the same bars, so an
incumbent that is the maximum on this data is the maximum on every run. One
~1-in-18 observation with three correlated repeats, ~5.6% - unremarkable. The
same independence assumption that made the uncorrected lag profile star four
lags. The candidate-spread line stands: it settles inert-vs-live directly.

No input, topology or label change: no retrain. Training in flight stays valid.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-07 09:08:01 -04:00
AnimateDread
0d5892357b diag: report the indicator tuner's candidate spread - "no improvement" is ambiguous
Auditing the other best-of-N scans after cccf94f turned up a third instance of
the same pattern, and this one is worse than the two already fixed: the geometry
scan and the lag profile PRINT a row, whereas TuneIndicatorsByFilter INSTALLS
its winner (Unflatten + ReInitADIndicators) and the caller then calls
BuildFreshTopology(), so an unguarded maximum changes the feature vector the
network trains on.

It has no null of any kind. But before adding one, the logs say something a
noise-driven best-of-N cannot: 2026-08-05/06, four consecutive runs, 17
candidates each, every one "no improvement" with start and best identical to
4dp. The maximum of 17 draws from a noise distribution beats its incumbent
about 94% of the time, so 4/4 is on the order of 1 in 100,000.

Two readings fit and they want opposite responses:
  - INERT: trial scores come back identical to the incumbent because the
    parameter change never reaches the scored features (suspect the feature
    cache surviving ReInitADIndicators), so `sc > bestScore` can never fire.
    That is a dead code path, and gating it would be decorating a corpse.
  - LIVE and correctly finding nothing: then it needs the family-wise gate.

The current log line cannot separate them, so add the number that can: the span
of the candidate scores, with an explicit ZERO SPREAD callout naming the likely
cause. Also widened the MI figures from 4dp to 5dp - at this scale 4dp rounds
the entire effect away.

No gate yet, deliberately: measure which failure this is, then fix that one.

Read-only diagnostic. No input, topology or label change: no retrain.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 21:59:10 -04:00
AnimateDread
cccf94f9ca fix: correct the lag profile across lags too - it contradicted itself
3271f1e tested each of ~21 lags against its OWN null at alpha 0.05 and starred
whatever cleared. That is about one false positive per run before any signal
exists, and because neighbouring lags share nearly their entire feature window
the false positives arrive in CLUSTERS that read like a hump.

It did exactly that on SP500 H1, twice in one afternoon on identical data:

  13:55  nothing clears at any lag       headline MI p=0.4478
  16:22  k6/k10/k12/k16 starred,         headline MI p=0.8756, observed
         "information survives to lag 16"   BELOW its own null mean

Same 31 features, same 2009 samples, same 287 blocks, cross-asset absent in
both - so this was not two different measurements. Non-replication on identical
data is the signature of an uncorrected multiple comparison, and acting on the
second run would have pinned the lookback to 17 off noise.

Galling detail: 04ee2e1 had just added exactly this correction to the
barrier-geometry scan one function below. The rigorous bar went on the report
with 6 candidates and the naive one stayed on the report with 21.

So the lag profile now uses the same construction as the geometry winner test:
one draw from every lag, keep the largest, repeat; a lag clears only by beating
that distribution. Draws centred leave-one-out to match how the observed excess
is centred. Independence across lags overstates the spread of the maximum
(neighbours share their window), so it errs toward rejecting.

Also: the positive branch now says to re-run before acting, because one run of
this report has demonstrably not been a result; and MI_LAG_MAX_PROFILE caps the
retained-draw matrix rather than trusting a derived m_historyBars.

Read-only diagnostic. No input, topology or label change: no retrain, and a
training run already in flight stays valid.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 16:28:57 -04:00
AnimateDread
04ee2e113a fix: gate the barrier-geometry winner on a family-wise null, not its own
The scan ends by printing "set SL_Mode/TP_Mode to <winner> and retrain".
That advisory fired on `bestExcess > cfgExcess * 1.5` - a ratio between two
numbers, with no test that either is distinguishable from zero.

bestExcess is a MAXIMUM over the eligible candidates. The maximum of several
draws from a null sits well above any single draw from it, so a max-shaped
statistic tested against a single-candidate null crowns a winner on noise
almost every time. On SP500 H1 the winner is 2:8 at +0.00081 nats - and the
lag profile committed in 3271f1e measures the pure-noise swing on this exact
data at +/-0.0004, peaking at +0.00042 with nothing clearing its own null at
any lag. The advisory was one ratio away from talking us into relabelling and
retraining all four topologies to chase that.

So build the null OF THE MAXIMUM: retain every candidate's permutation draws,
take one draw from each candidate, keep the largest, repeat. The winner must
beat that distribution.

- draws centred LEAVE-ONE-OUT, so a draw is centred by a mean excluding it -
  exactly how the observed score is centred. Centring a draw by a mean that
  contains it shrinks it toward zero and would deflate the null.
- only ELIGIBLE candidates enrol: the family the max was taken over is the
  family to correct for, and a clamped or sub-minRR pairing can never win.
  rrOK hoisted above the draws for this.
- draws per candidate 20 -> MI_GEOMETRY_PERMUTATIONS (40): they now have to
  resolve an upper tail, which is where 20 draws are thinnest.
- MI_GEOMETRY_ALPHA 0.05, stricter than the lag profile's: a wrong lookback
  costs input width, a wrong geometry costs a full retrain from era 0.

Independence across candidates overstates the spread of the max (the real
candidates share features and overlapping label windows), so the gate errs
toward rejecting - the safe direction when passing costs a retrain.

Read-only diagnostic. No input, topology or label change: no retrain.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 14:04:16 -04:00
AnimateDread
3271f1ea93 diag: MI feature-lag profile - close the blind spot in every MI verdict so far
BuildMiSample samples features from ONE bar. So every "MI is at the noise
floor" result this codebase has produced - including yesterday's p=0.18 on
SP500 H1 - described the ENTRY BAR's 31 features only, while the network is
fed 20 bars of them. If information lived at lag 7 and not lag 0, the report
would have said "no signal" while the model could still learn. The diagnostic
we have been making decisions on had a blind spot exactly the width of the
input vector.

Adds a FEATURE-side offset to BuildMiSample, which is not the same thing as
the existing labelBarOffset and is not interchangeable with it. Shifting the
LABEL changes which trade is predicted, so at any non-zero offset the
features sit inside the labelled window and the score is lookahead - that is
precisely what the alignment scan measures and correctly reports (4.7x more
knowable 5 bars into a 128-bar window). Shifting the FEATURES keeps the label
pinned to the entry bar, so every row stays causal.

ReportFeatureLagProfile() then scores k = 0..historyBars against the same
block-permutation null and reports the deepest lag that clears it - the
lookback the data supports, versus the 20 that was picked by hand and never
measured. The null is redrawn PER LAG: finite-sample MI bias moves with the
realised class counts and bin occupancy, and different rows survive the
validity checks at each lag, so one shared floor would be right for lag 0 and
wrong everywhere else. Draw count is reduced accordingly (40, not 200) since
cost is draws x historyBars; this figure decides a lookback, never a trade.

MiShiftPad now also covers historyBars, keeping the fixed-pad invariant that
makes two builds comparable row by row.

Read-only - no input, topology or label change, so no retrain. Both builds
0/0. Build tag lag-profile-v1.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 13:51:39 -04:00
AnimateDread
8ccbddb051 Add new research scripts for trading strategy analysis
- Implemented sqx_audit.py to audit StrategyQuant X trade lists, focusing on performance metrics and cost analysis.
- Created sqx_portfolio.py to evaluate portfolio performance based on uncorrelated components and their impact on risk and return.
- Developed swing.py to analyze cost ratios across different holding periods and assess swing trading structures.
- Introduced test_management.py to investigate the effectiveness of exit rules on random entries and their impact on expectancy.
2026-08-02 12:25:20 -04:00
AnimateDread
f1b7dcf7f3 fix: correct MI sample alignment and improve BN weight diagnostic report
The MI sample builder used `MathAbs(labelBarOffset)` as a padding, causing rows from offset and non-offset builds to be paired with a double shift. This broke the positive control, failed the 5× gate, and voided all reported mutual‑information figures. Replace with the fixed `MiShiftPad` constant to ensure builds enumerate the same set of bars and row-k alignment is preserved.

Add `BatchOptionsTotal()` to `CNeuronBatchNormOCL` and split the packed BN weight array in the learning report into separate norms for the outgoing dense matrix, gamma, beta, running statistics, and Adam moment buffers. This turns an ambiguous single‑norm reading into precise diagnostics that distinguish weight divergence from scaling issues.
2026-08-02 08:12:47 -04:00
AnimateDread
7d038df749 research: export the feature matrix and a raw OHLCV grid for offline work
The bottleneck on this project has never been the modelling - it is that
every hypothesis costs a compile, a deploy, an attach and a log read, and
answers exactly one question. Days have gone into questions that are
seconds of arithmetic once the data is in hand.

Adds a RESEARCH-ONLY build, gated behind WARRIOR_EXPORT_FEATURES and
never compiled into a shipped binary, which writes two things to
Common\Files\Warrior_EA\Research\ and then does nothing at all:

  <symbol>_<tf>_features.csv - one row per bar: index, time, OHLC, ATR,
  and the m_neuronsCount feature values. Exactly what the network sees.
  The raw bars ride along on purpose: with OHLC and ATR offline, every
  barrier geometry, horizon and in-trade target is recomputable without
  MetaTrader in the loop.

  <symbol>_<tf>_rates.csv - raw OHLCV across a grid of 8 symbols x 5
  timeframes. The 26 engineered features only exist for the attached
  chart (indicator handles bind to PERIOD_CURRENT); raw rates do not, so
  ONE attach yields the whole research grid. The bar time also makes
  session/hour/day-of-week derivable - the only inputs in play that are
  not a transform of the same OHLCV series.

Safety, because this binary gets attached to a chart on a LIVE ACCOUNT to
reach real history:
  - OnTick returns immediately, so Expert.OnTick() - the entire trading
    path - is unreachable regardless of the AlgoTrading toggle, the
    signal state or the inputs. Structurally incapable of sending an
    order, not merely unlikely to.
  - No config lock. It never trains and never saves a model, so it has
    nothing to protect against a concurrent chart - and taking the lock
    would make it refuse to start exactly when the config it wants to
    read is already open, which is when it is most useful.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 15:49:57 -04:00
AnimateDread
004f2a04f7 fix(diag): the symbol sweep was measuring its own sampling, not the market
Twelve cells came back with higher-timeframe "signal" 5-9x anything on
H1, at p=0.005. It was an artifact, and the sweep's own columns gave it
away: excess tracked the sampling STRIDE almost monotonically, and the
three D1 cells - stride collapsed to 1-5 bars against a 128-bar horizon,
i.e. ~99% window overlap - were the three highest. Three flaws, all the
same family: comparing numbers without the spread that belongs to them.

1. THE NULL ASSUMED INDEPENDENCE THE LABELS DO NOT HAVE. Triple-barrier
labels overlap; two rows less than one horizon apart share most of their
outcome window. A free Fisher-Yates shuffle destroys that dependence
along with the association, making the null far narrower than the truth
and handing out significance that isn't there - Lopez de Prado ch. 4
arriving through the back door of the significance test. Now permutes
contiguous BLOCKS of at least one horizon, so the null keeps the
autocorrelation and the p-value means what it says. It degrades honestly:
severe overlap leaves few blocks, the null widens, nothing reaches
significance. The block count is now printed, because THAT - not the row
count - is the sample size a p-value rests on, and a warning fires under
30 blocks so "not significant" is not misread as "no signal" when it
means "not enough independent history to tell".

2. THE POSITIVE CONTROL'S STRENGTH DEPENDED ON THE DATASET. It paired
each row's label with the NEXT SAMPLE ROW's, whose distance is the
stride - so on M5, where stride ran 160-717 bars against a 128-bar
horizon, it was pairing two windows that never overlap. All three M5
cells duly reported a FAILED estimator and voided their own results with
nothing wrong. A control whose strength varies with the cell cannot
certify the cell. Now pinned to a quarter of the horizon, where ~75%
overlap is guaranteed by construction.

3. THE LOOKAHEAD VERDICT HAD NO MARGIN. It flagged 7 of 12 cells on gaps
of 0.00008-0.00040 nats against a measured null sd of ~0.00030 - noise,
every one. Now requires 3 sd, the same discipline the deploy floor
applies to precision.

Compiles 0 errors / 0 warnings, standard and Market. Build tag
blockperm-v1. Supersedes every number from the sweep.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 15:11:40 -04:00
AnimateDread
168422ff7a fix(labels): the 128-bar horizon ceiling was truncating the shipped label
The corrected geometry scan exposed something bigger than the geometry
question it was asked. Every pairing from 2:6 upward came back CLAMPED -
including 2:6, the SHIPPED configuration.

First-passage time for a driftless walk leaving [-m,+k] goes as m*k, and
the measured swing median here is ~12 bars at m*k=1, so 2:6 wants ~144
bars and 3:10 wants ~360. The ladder stopped at 128. A clamped label
stops meaning "does the target come before the stop" and quietly becomes
"...within 128 bars", while the deployed EA holds until SL or TP with no
bar limit. So the target the models have been trained on all along was
not the strategy the EA executes, and the trades it silently reclassified
as Neutral were the SLOW WINNERS - precisely the ones a 1:3 barrier
exists to capture. Timeout share stayed ~0% throughout, which is why this
never showed up: the truncation lands in Neutral, not in the timeout
counter that was watching for it.

Ladder extended to 384 (12..128, 192, 256, 384) so every selectable
geometry gets an honest horizon. Cost is one embargo of at most 384 bars
out of ~38k.

Second fix, same class of error as the H(Y) one: the scan's "best
eligible" was 2:2, a 1:1 barrier, against a shipped Min_Risk_Reward_Ratio
of 1:2. Training four topologies on that target would have produced a
model whose every setup is rejected at the door - the exact failure
behind four consecutive Market rejections for "no trading operations".
Sub-minRR geometries are now ineligible and marked [<minRR], printed
rather than hidden.

Also drops the dense-depth tag from the display name ("Perceptron 3L" ->
"Perceptron"). Depth is derived, so it names nothing a user chose; the
config tag [PAI-0be2] already disambiguates concurrent charts and does it
for every input rather than one. Full topology still logged by "config -".

Compiles 0 errors / 0 warnings, standard and Market. Build tag
horizon-384-v1. Changes the LABEL for every geometry, so the next scan
supersedes the previous numbers - and a retrain is required before any
model trained under the truncated target means anything.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 14:54:45 -04:00
AnimateDread
40af4a4b5b fix(labels): the geometry scan rewarded the labels it should reject
First run named 3:10 on all four charts, at 2.3x the configured 2:6. That
answer was wrong and the fault was the ranking statistic.

3:10 wants a horizon of ~swingMedian*30 (~320 bars) and gets
BARRIER_HORIZON_MAX. Clamped, most trades never resolve, the unresolved
remainder all lands in Neutral, and H(Y) collapses. The old statistic
divided the excess BY H(Y) - so a collapsing denominator made the most
degenerate label look like the most predictable one. Every geometry from
2:6 upward was already showing the clamped h128, and the two widest
scored highest, which is the fingerprint of the artefact rather than of
signal.

Two fixes:

Rank on the raw excess in nats. Subtracting each geometry's OWN measured
null already removes the class-balance bias, which is the only thing the
normalisation was ever needed for.

Disqualify clamped geometries outright rather than ranking them down. The
deployed EA holds until SL or TP with no bar limit, so a truncated label
trains the model on a question the strategy never asks. They are still
printed, marked '!', so the disqualification is visible instead of a
silent omission - and the scan now says so explicitly when nothing
eligible is left, because "the limit is the feature set, not the target"
is itself the finding in that case.

The scan also reports each geometry's directional share and timeout share
now. A label nobody can trade is not a candidate however well it scores,
and that has to be visible in the same line as the score.

Compiles 0 errors / 0 warnings. Build tag geometry-scan-v2.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 14:46:25 -04:00
AnimateDread
f97ab9f1d6 feat(labels): measure which barrier is predictable at entry, don't guess
The alignment scan settled the shape of the problem: 4.7x more is
knowable 5 bars into a 128-bar window than at the entry the model
actually trades. A 6xATR target reached over 128 bars is decided
overwhelmingly by what happens DURING the window, so whatever the entry
state knows is buried under 128 bars of later noise. That is a property
of the TARGET, and it is why four different architectures all landed on
precision exactly equal to the base rate - no topology can undo it.

So measure the target. For each SL/TP pairing a user can actually select,
relabel the same sampled bars and score how much the SAME features say
about THAT outcome at entry. Seconds, no training, no topology, and it
runs on the diagnostic path that already exists.

Ranked on excess over its OWN null as a share of its OWN H(Y), never on
raw nats: each geometry has a different class balance, hence a different
finite-sample bias and a different amount of information there to find,
so raw MI would rank the most BALANCED label rather than the most
PREDICTABLE one. The break-even win rate m/(m+k) is printed beside each
so the ranking is read next to the bar the model must clear.

Stated in the output because it is the easy thing to get wrong: chance
precision EQUALS break-even at every geometry, so a tighter target does
not hand you expectancy. It buys predictability - less noise piled on top
of what the entry state knows - which is the one thing changing topology
cannot do.

Read-only by construction: it relabels a sampled copy via
TripleBarrierLabel(), never writes the label cache (which belongs to the
configured geometry), and restores the horizon and overrides it borrowed.
The overrides apply only when BOTH are positive, so a half-set pair can
never silently relabel a live run.

Compiles 0 errors / 0 warnings, standard and Market. Build tag
geometry-scan-v1. Redeploy only - no retrain to READ the ranking.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 14:32:04 -04:00
AnimateDread
4443ce85c1 fix(diag): the alignment scan cried misalignment at its own arithmetic
First run came back "WARNING - peak at k=+5, NOT 0 ... a feature/label
misalignment upstream of every topology". That was a false alarm produced
by the diagnostic's own design, and exactly the kind of plausible-looking
output this project has lost days to.

Bar indices are MQL5 SERIES indices - HIGHER index = OLDER bar
(TripleBarrierLabel walks its window as `for(t = idx-1; t >= idx-horizon;
t--)`, decreasing index = forward in time). The two directions therefore
mean opposite things and the scan treated them as symmetric:

  k < 0  label belongs to a NEWER bar, its barrier window opens AFTER the
         features exist. Nothing at bar i can legitimately know it, so a
         peak here is real lookahead and a bug.
  k > 0  label belongs to an OLDER bar, already k bars into its window by
         the time bar i happens - so the features hold the realised first
         k bars of that outcome. MI MUST rise with k. Arithmetic.

Only the k<0 side can indict the pipeline, and on the observed data it is
clean: -5/-3/-2/-1 all sit at or below the k=0 value and the noise floor,
so there is no lookahead - a real negative result, not an absence of
evidence.

The k>0 side is now reported as what it is, a second positive control,
with its gradient as the finding: 0.01881 at k=+5 against 0.00401 at k=0
means ~4.7x more is knowable 5 bars into a 128-bar window than at the
entry the model actually trades on.

Compiles 0 errors / 0 warnings. Build tag mi-align-v2. Redeploy only.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 14:25:26 -04:00
AnimateDread
87c8656b53 diag(autotune): a positive control, and a scan that separates "no signal"
from "signal knocked out of step"

Four architecturally different networks landed on the same precision -
Buy 23-25% against a 25.4% base rate, Sell 19-22% against 22.0% - while
making completely different calls (HYBRID votes Sell on 69% of bars, PAI
on 41%). Precision equal to the base rate is what INDEPENDENCE looks
like, and precision under independence is fixed by the label
distribution, not by the architecture, so all four converging on it is
arithmetic rather than coincidence. Accuracy meanwhile tracks coverage
exactly as independence predicts (31.1/30.3/25.0 predicted vs
31.8/28.9/24.6 observed for PAI/CONV/HYB).

But "no information in the data" and "information destroyed upstream of
every topology" produce that identical picture, and the MI test alone
cannot tell them apart either. Two additions:

POSITIVE CONTROL. Three "measurements" in this codebase have turned out
to be silent no-ops that produced plausible numbers - the MI scorer
reading an array nobody filled, the eval-mode guard that switched off the
imbalance correction, the alternation gate whose premise was never true.
So the estimator now has to prove it responds to a signal known to be
present before any floor reading is believed: the label of a neighbouring
sample row, ~19 bars away and far inside the 128-bar barrier horizon, so
the two outcome windows overlap heavily and MUST be associated. Same
binning, same estimator. Near the floor => every MI figure is void.

ALIGNMENT SCAN. Re-scores against the label taken from bar i+k for k in
-5..+5. A peak at k != 0 is a feature/label misalignment - an off-by-one
in the label index, a horizon applied to the wrong bar, a feature window
that lags what it claims - which would destroy the information before any
topology saw it and would look identical in every accuracy number this EA
prints. A flat profile says the features simply do not carry this target.
The sampled range is trimmed by |k| at both ends so a shift is measured
rather than an edge effect, and both bars must carry a real label.

Also: BuildMiSample publishes its stride instead of the report
recomputing that arithmetic (it would drift), and the control sizes its
buffers from its own sample count rather than the caller's.

Compiles 0 errors / 0 warnings, standard and Market.
Build tag mi-control-align-v1. Redeploy only - no retrain, no model
deletion; the diagnostic runs on resumed models.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 14:12:10 -04:00
AnimateDread
9a5f645dc3 diag(autotune): stop making the feature test cost a trained model
The permutation test lived inside TuneIndicatorsByFilter, which is gated
on era 0 - correctly, because re-running the SWEEP would change the input
vector out from under weights already fitted to the old one. But the test
itself reads cached features and writes nothing, so none of that applies
to it, and the gate meant the only way to see the answer on a running
model was to delete the model. Today that price was PAI's 45 trained eras
and CONV's 31, spent to re-ask a read-only question.

Split into ReportFeatureLabelInformation(), called from the sweep when it
runs and directly when it does not - a resumed model, a disabled tuner,
nothing tunable. Once per attach either way.

Compiles 0 errors / 0 warnings. Build tag permtest-v2.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 14:01:32 -04:00
AnimateDread
9920754dec diag(autotune): five permutations was still a coin flip - use a real test
The 5-draw z-score shipped an hour ago disproved itself on its first run.
All four charts scored the IDENTICAL 0.00401 nats on identical features
and identical labels - and reported z of +1.3, +2.0, +4.0 and +4.7. Two
"AT THE NOISE FLOOR", two "a real association", same data. The entire
swing came from estimating the null's spread from five draws, where the
standard deviation of the standard-deviation estimate is ~35%: the
denominator was noisier than the effect it was judging.

Replaced with an empirical permutation test. 200 draws, p counted by rank
with the +1/(B+1) correction (Phipson & Smyth 2010) so p is never
reported as exactly zero - no normality assumption and no spread to
estimate. The strongest single column is tested against the null
distribution OF THE MAXIMUM, which corrects for scoring 26 features at
once by construction and is far less conservative than Bonferroni.

Affordable because BuildMiSample is now split out of ScoreCurrentParamsByMI
and runs ONCE for the whole test - every draw reuses that sample and costs
a relabel plus 26 histogram passes, not 2000 feature extractions. The
coordinate sweep still calls the combined form, which is correct there:
each candidate changes the indicator settings, so its features really do
have to be re-extracted.

The verdict line keeps both questions apart and prints both answers: the
p-value for "is it real", the excess as a percentage of H(Y) for "is it
big enough to trade". At n=2000 those can disagree, and collapsing them
into one word is how a worthless effect gets called a discovery.

Compiles 0 errors / 0 warnings. Build tag permtest-v1.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 13:45:46 -04:00
AnimateDread
12a1fbd133 diag(autotune): one label shuffle cannot settle the no-edge question
The permutation baseline added in 018afb1 came back on all four charts as
0.00401 nats against floors of 0.00267 / 0.00298 / 0.00318 - three draws
whose spread is as wide as the excess being judged, because one shuffle
is one sample from the null, not the null. That is not enough to retire a
topology on.

Now MI_NOISE_PERMUTATIONS draws, reported as mean +/- sd with a z-score,
plus two numbers the mean over 26 columns cannot express:

  - the STRONGEST single feature's MI, against its own shuffled value.
    One informative column among 25 useless ones is precisely the case
    the mean hides, and precisely the case worth finding.
  - the excess as a percentage of H(Y). At these sample sizes a z-score
    can be comfortably significant while the effect is worthless, so
    "is it real" and "is it big enough to matter" are asked separately
    and answered separately.

The verdict line also now states the measure's limit every time rather
than only when the news is bad: this is a MARGINAL, PER-BAR statistic and
the network reads m_historyBars bars jointly, so it can prove signal
exists but never that it does not. It rules out a per-feature edge - and
therefore any indicator retuning - not an edge that lives in a
combination or across time.

Compiles 0 errors / 0 warnings, standard and Market.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 13:32:12 -04:00
AnimateDread
018afb1ba9 fix(autotune): MI scorer read an array nobody filled; add the permutation floor
THE TUNER WAS A SILENT NO-OP. Every chart logged

  auto-tune complete - 17 candidate settings scored in ~139s,
  feature/label mutual information 0.0000 -> 0.0000 nats (no improvement)

0.0000 is not a weak result, it is a broken measurement: finite-sample MI
is biased UPWARD, so even pure noise scores above zero. Cause:
ScoreCurrentParamsByMI called BufferTempDataCompute(), which APPENDS the
bar's features to TempData and never touches m_featureCache - only the
caching wrapper BufferTempData() writes that array. It then read
m_featureCache, which ReInitADIndicators had just invalidated. Every
column came back constant, FeatureColumnMI returned 0 for all of them,
and all 17 candidates tied at exactly zero. 139 s per chart to return the
settings it started with.

Now reads the values back out of TempData, where they actually land. And
an exactly-zero best score is called out as a fault rather than reported
as "no improvement", because that is what it is.

ADDED: a PERMUTATION BASELINE, which is the diagnostic this project has
been missing. MI's finite-sample bias is ~(bins-1)(classes-1)/(2n) nats -
at these sample sizes the same order as any real edge in this domain - so
a raw MI figure is uninterpretable on its own. Shuffling the labels
destroys every genuine association while leaving sample size, binning and
class proportions intact, so the score it produces IS this dataset's
noise floor, measured rather than approximated. The log now reads

  feature/label information - X nats against a shuffled-label floor of Y

and says outright whether the features carry usable information about the
target. It needs no training, no topology and no convergence, so unlike
every accuracy number in this codebase it cannot be confounded by an
optimizer or an objective. If the score sits on the floor, no change of
architecture can help - which is the question the last three days of
zero-edge results have been circling.

DEPLOY FLOOR: `dirPrecPct > chancePrecPct` passed anything above chance by
any amount. At ~11,000 directional calls the standard error of the
precision estimate is ~0.4pp, so that gate was accepting sub-one-sigma
noise - the perceptron deployed at edge +0pp on 2026-08-01. Now requires
EDGE_MIN_SIGMAS (2.0) standard errors above chance, computed from the
actual call count, so the bar scales with the evidence instead of needing
a hand-picked constant.

Recorded with it, because it is why chance is the right reference at all:
under a driftless random walk P(touch +k*ATR before -m*ATR) = m/(m+k),
and the break-even win rate for a k:m reward:risk trade is ALSO m/(m+k).
The label's own base rate IS the break-even rate, at every SL/TP setting.
So "beats chance" and "is profitable" are the same test, and no choice of
SL/TP can manufacture an edge - only prediction can.

Both builds compile 0 errors / 0 warnings.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 13:05:50 -04:00
AnimateDread
6db0519472 perf(autotune): replace the genetic search with a filter score - hours to seconds
MEASURED COST OF THE GA, which is what retired it. Per generation:
  rung 0: 8 cand x 3 seeds x  3 eras =  72 eras
  rung 1: 4 cand x 3 seeds x  8 eras =  96
  rung 2: 2 cand x 3 seeds x 20 eras = 120
  = 288 eras/generation x 4 generations = 1152 eras BEFORE the winner's
real training began. Against the observed era times on SP500 H1:

  PAI     29.1 s/era  ->   9.3 h   (matches the observed 00:37 -> 09:22)
  CONV    41.3 s/era  ->  13.2 h
  LSTM   150.4 s/era  ->  48.1 h
  HYBRID 154.6 s/era  ->  49.5 h

Two days to tune is not a first-run experience, and it is the phase in
which the panel goes quiet, which is what made it look like a hang.

It also bought nothing. The space is 90 points (10 MA periods x 9 MA
types), so 1152 evaluations revisited each point ~13 times; and rungs of
3 and 8 eras cannot separate two MA periods at all. The 2026-08-01 run
proves it: every finalist scored 25.0-25.9% balanced accuracy - below the
33.3% one-class floor, i.e. indistinguishable noise - and the search then
"deployed the winner" of that.

THE ERROR WAS THE SCORING FUNCTION, not its constants. Using a full
training run to choose a feature's period is a wrapper method paying
wrapper prices for a decision that does not need one. The reference book
does not do this: ch. 3.3 selects inputs by measuring each candidate
indicator's CORRELATION with the target and dropping the ones with none,
with no network involved.

So: rank candidates by the MUTUAL INFORMATION between the resulting
feature vector and the triple-barrier label. MI rather than correlation
because the label is 3-class categorical and the features are not
monotonically related to it. Equal-FREQUENCY binning (rank-based),
because these features are ATR-normalised and heavy-tailed - fixed-width
bins put nearly everything in one bucket and report ~0 information for a
genuinely useful feature.

Scoring is arithmetic over the feature cache, so it costs seconds and its
cost is independent of topology: LSTM now tunes as fast as the MLP.
Coordinate sweep, not product sweep - cost is the SUM of per-parameter
candidate counts, so enabling every indicator stays affordable - with a
second pass that breaks early once nothing moves.

Sampling is IS-ONLY. Letting the OOS window influence which indicator
settings ship would mean the holdout had been used for selection and had
stopped being a holdout.

HONEST LIMIT, recorded because it is the price: MI is marginal, so a
parameter that only pays off in combination with another can be missed
(Guyon & Elisseeff 2003, filter vs wrapper). Given the wrapper it
replaces was ranking pure noise at 48 h a run, this is strictly better.

Deleted with it: GaRungEras/GaExtract/GaStore/GaMutate/GaRandomCandidate/
GaBlockCrossover/GaSortAliveByScoreDesc/GaBreedNextGeneration, 14 m_ga*
members, the GA_*/TUNE_POP_* constants, and ComputeTuneTrialBudget.

AND m_evalMode/m_evalEraBudget, because nothing set them any more - 28
read sites all permanently inert. That is not a tidy-up: the `if
(!m_evalMode)` guard on UpdateClassPriors is exactly what silently
disabled the imbalance correction for entire runs two commits ago. Dead
machinery that still reads like live machinery is this codebase's most
expensive recurring bug, and leaving 28 more instances of it would have
been indefensible.

The panel's tuning-progress state goes too - tuning no longer takes long
enough to need one.

Both builds compile 0 errors / 0 warnings.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 10:29:31 -04:00
AnimateDread
f48bc93f9b refactor(inputs): 96 -> 70 inputs; remove two untested/unusable filter modules
Every removal below is FINGERPRINT-NEUTRAL by construction: each retired
input is pinned to the exact value it already shipped with, so running
models keep their filenames and resume rather than restarting at era 0.
Verified field by field against BuildConfigFingerprint.

Removed as inputs, kept as pinned constants (the value was never a
preference the user had a basis to change):

- OutputNeuronsCount. The regression head predicts a continuous quantity
  the triple-barrier label does not contain; the target is an EVENT, so
  the right output is its probability. The regression code paths stay
  implemented and dormant - they cost nothing and removing them would
  touch every scoring path at once.
- MinRecall. A safety floor, not a preference, and the only direction a
  user can move it is the harmful one: raising it past what the config
  reaches yields NO model, not a better one (observed repeatedly at 60).
- SwingConfirmationBars. Stopped gating the labels with the relabel, but
  is STILL load-bearing for the swing-context input features - it is the
  ZigZag repainting embargo, and without it those 9 features read a leg
  the live bar could not have had yet. Pinned, not deleted.
- MaxErasPerRun (runaway backstop, never reached in a healthy run),
  FreezePriorCalibration (unanswerable by a user; near-balanced labels
  make the priors stable anyway), VerboseMode (developer view, joins
  DebuggingMode), MACD/Ichimoku periods x6 (both indicators ship
  disabled, and as optimizer dimensions they are pure overfitting
  surface - the AI auto-tuner is the supported way to move them).
- SignalClusterWindow -> 3, no longer an input. Barrier labels make
  consecutive setups real, which argued for 0; it is not 0 because on D1+
  a 6-bar window spans over a week and two arrows a day apart on a
  weekly-scale move are one event. 3 splits it correctly by timeframe.
- EnableOnlineLearning -> ON. Adapting to a changing market is what keeps
  a months-attached model from going stale, and the rolling-accuracy
  freeze is what makes it safe. See the caveat noted in the handoff: it
  had not been forward-tested on a live feed when this became default.

Removed entirely:

- Intraday Time Filter (5 inputs + Signals/SignalITF.mqh). Two of its
  five inputs were raw BITMASKS, which is an implementation detail
  exposed as a control. The job is covered three times over by things
  that are declarative or that learn: the session filter, the
  time-of-day/day-of-week input features (the network discovers which
  hours are good rather than being told), and the journal's time buckets.
- Market Depth Filter (5 inputs + Signals/SignalMarketDepth.mqh, plus
  its OnInit probe and OnDeinit release). It needs real level-2 data
  that this broker - and most retail MT5 brokers - do not provide, so
  the module has never once executed against real data. Shipping four
  tuning dropdowns for an untested path is worse than shipping nothing:
  the only users who could enable it would be its first-ever testers,
  live. If DOM returns it should be a FEATURE fed to the network, not a
  rule-based veto with hand-tuned thresholds - imbalance is data.
- IndicatorTuneTrials, replaced by ComputeTuneTrialBudget(). The useful
  budget depends on how many parameters are actually being searched,
  which depends on which features are enabled - so one number meant
  wildly different things run to run. The shipped 32 was ~10 candidates
  per dimension against one enabled indicator (wasteful: each costs
  GA_SEEDS full training runs) and under one per dimension against all
  nine (blind). Now population ~ 4 x active dimensions, clamped [8,64],
  with CADIndicatorTuner::ActiveDimensions() defined immediately above
  PerturbRandom() so the two cannot drift apart.
- Six orphaned enums (TUNE_TRIALS_PRESET, DOM_*, ENTRY_HOUR_OF_DAY,
  TIME_FILTER_DAY_OF_WEEK), 81 lines.

Other UX:

- SL_ATR_x1 / TP_ATR_x3 now carry the "(classic)" default marker every
  other preset enum in the file already used. Nothing in the SL/TP
  dropdowns previously told a user which pair was the shipped default -
  which matters far more since the relabel, because those two define the
  labels and changing either forces a retrain.
- Neural Network section moved directly ABOVE AI Input Features: choose
  the architecture, then choose what it sees. NN Optimizer / Performance
  stays last - the Adam/Sgd inputs are declared in AI/Network.mqh and
  render immediately after that divider.
- News feature + window moved to the end of the AI feature list, below
  Wyckoff Bar Inversion.
- Dropped "(0-100)" from Min vote to open - it is an enum, not a number.

Both builds compile 0 errors / 0 warnings. No retrain forced.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 21:22:02 -04:00
AnimateDread
2de93539d4 refactor: split CExpertSignalAIBase implementation by responsibility
ExpertSignalAIBase.mqh was 8216 lines: the class declaration followed by 87
method bodies covering training, labelling, feature extraction, persistence,
chart drawing, online learning, the GA auto-tuner and inference, all in one
file. Train() alone is 1492 lines; a change to arrow drawing meant scrolling
past the era loop.

Moved the bodies into Expert\AIBase\, included at the bottom of the original
after the class declaration:

  Training.mqh        1607  era loop, plateau ladder, checkpoint select, deploy
  Features.mqh        1093  indicator creation + per-bar input feature vector
  ChartUI.mqh          634  arrows, arrow persistence, status panel, cleanup
  Persistence.mqh      492  .stats/.cfg sidecars, CPU-inference validation, copy
  OnlineLearning.mqh   461  live continual learning, EMA shadow, OOS simulator
  Labels.mqh           309  ZigZag pivot labels, async label-cache prebuild
  AutoTune.mqh         275  genetic tuner (population, crossover, halving)
  Inference.mqh        235  softmax, prior calibration, class priors

  ExpertSignalAIBase.mqh  8216 -> 3131 (declaration + topology build only)

This is a pure relocation - verified mechanically, not by eye: HEAD's file
reconstructed from the eight partials plus the surviving remainder is
byte-identical to HEAD, span for span (scratchpad verify_split.py). No
declaration moved, no signature changed, no code rewritten, so behaviour is
unchanged by construction.

Compiles 0 errors, 0 warnings.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 00:42:45 -04:00