Commit graph Warrior_EA/Warrior_EA.mq5
Author SHA1 Message Date
AnimateDread
d70efc0f74 fix(gate): the book had to beat zero, not beat what it costs to collect
WarriorRungBookProfitable tested (hLong+hShort)/n > 0.0. The round-turn spread was
computed a few lines away in the same function, printed in the payoff line, and
carried an explicit comment saying it gates NOTHING - "spreadR is a live snapshot
to be judged by hand".

That objection was correct about the QUANTITY and wrong about the CONCLUSION.
SymbolInfoInteger(SYMBOL_SPREAD) is whatever the book looks like at the instant an
era happens to end, which is not what the historical trades in that window would
have paid - so it should not gate. But this project has already measured that "4 of
5 die to spread and the survivor dies on commission"
(project_cost_boundary_equilibrium), and EXPECTANCY IS THE BAR. The answer is to
measure cost properly, not to leave it out of the gate.

MEASURED, NOT SNAPSHOTTED. g_ensVoteCostR carries the round-turn spread at each
vote row, taken from the broker's own per-bar history (CopySpread, already
maintained as m_spreadSeries for the feature block) and divided by that bar's ATR,
so it lands in the SAME units as the ride book it has to beat. Filled once per ROW
in the block that already measures the ride and the excursions, for the reason
stated there: cost is a property of the CHART at that bar, identical for every
member.

PER RUNG, over the bars that rung actually fires on - not a chart-wide average.
Signals cluster, and a cluster can sit in a wider-spread regime than the window
mean, so a window average would understate the cost of exactly the bars being
traded. sweepCost/sweepCostN accumulate alongside sweepFired.

AN UNKNOWN COST DOES NOT WAIVE THE TEST. g_ensVoteCostR is -1 where the spread
series or the ATR was unavailable and those rows are DROPPED from the mean rather
than read as zero - a cost of zero is the one answer that can never be right. If a
rung has no measurable cost at all the gate falls back to the pre-existing `> 0`
bar, which is "cannot price this", not a free pass.

NOT a multiple of the cost. The "expected payoff should be at least double the
spread" rule is a common heuristic (and is what prompted this - MQL5 article 8410,
Ilin, sent by the operator) but it is not measured here, so it is not imposed.
Sampling error on the book is already carried by the exact binomial bar the same
verdict applies to precision.

WHAT IT CHANGES, on the fleet as it stands. Median book per call against the
census's measured spread/ATR:

    SP500   +0.02  vs ~0.065   -> now correctly REFUSED (was certified)
    GBPUSD  +0.05  vs ~0.020   -> marginal
    DAX40   +0.08  vs ~0.029   -> marginal
    BTCUSD  +0.23  vs ~0.006   -> clears
    XTIUSD  +0.57  vs ~0.083   -> clears
    NAS100  +0.75  vs ~0.025   -> clears

Both charts currently DEPLOYABLE clear it comfortably, so this closes a hole rather
than reversing a live decision - but SP500's book was being certified while sitting
below its own spread.

The sweep line now prints cost beside the book at every rung, so BOOK-FAIL names a
visible reason instead of an invisible one, and its legend says what the number is
and what "n/a" falls back to.

Three call sites updated, not one - the rung walk, the derived-rung verdict and the
sweep report all ask this question (feedback_rename_leaves_readers_behind).

Not retrain-forcing: no fingerprint member moves, and g_ensVoteCostR is an in-memory
per-era buffer.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-03 13:25:48 -04:00
AnimateDread
e13e317e54 feat(signal): one operating point per side - a single global cut had silenced the sell book
User report: "some charts are not showing sell signals at all, only buys". It is
real. By-side ride book and n on the fleet's most recent era per chart - the
drift-free test the era log already prints:

    NAS100  160 long vs      5 short   -    32 : 1
    DAX40    30 long vs      0 short   -   no short book at all

(Earlier eras from the same charts read 744:1, 260:1 and 41:1, but those lines
predate today's restart and are quoted nowhere as current - the two above are the
post-restart measurement.)

THE LABEL IS NOT THE CAUSE. P(ride pays | up-leg) vs P(| down-leg) is 0.99x to
1.55x across the fleet and up-legs are ~50% of bars. A 1.2x rate asymmetry was
being amplified into 744:1.

THE MECHANISM is the operating point's depth. FitDirConfThreshold places one cut
where coverage matches the pooled directional label rate, and at 3-25% coverage
that cut sits deep in the tail. There, a small shift in one side's margin
distribution moves nearly ALL of that side across it: the threshold hits its
target POOLED and misses it per side by two orders of magnitude. The model was not
failing to see sells - BTCUSD's pre-threshold sell recall is 84% and its raw argmax
is B45/S40 - the single cut was discarding them.

Each side now clears the margin ITS OWN measured label rate asks for.
m_labelPrebuildBuyCount/SellCount already existed, so the two targets are
measurements, not choices.

THE PROPERTY THAT MAKES THIS SAFE, AND IT IS EXACT: the per-side targets are
100*buy/tot and 100*sell/tot against the same denominator the pooled rule uses, so
they SUM TO THE POOLED TARGET. Total coverage is preserved; only its split between
the books changes. This redistributes exposure rather than increasing it.

THE TRAP, NAMED: 5a7d96c (project_gate_book_not_sides) retired a per-side test
because it selected on noise. This is not that, and the difference is the whole
argument. That was a best-of-2 VERDICT on an outcome metric - selection. This fits
a RATE to a separately-MEASURED rate and maximises nothing, which is
FitDirConfThreshold's own stated rule ("call a direction as often as a direction
actually occurs") applied honestly per side instead of only in aggregate. The
precision at each fitted point is logged as REPORTED, not optimised, for the same
reason the pooled line already says so. And the deploy gate still judges the
POOLED book, so if the newly admitted calls do not pay, the gate closes -
restoring the sell side's COVERAGE is not a claim that it is profitable.

EACH SIDE MUST EARN ITS OWN FIT: DIR_CONF_MIN_FIT_CALLS per side, the same
evidence the pooled point needs, because a side's operating point deserves as much
as the pooled one did. A side that cannot reach it keeps DIR_CONF_SIDE_UNFITTED and
falls back to the pooled threshold - byte-for-byte today's behaviour - so this can
only change a decision where there is enough evidence to change it, and a model
whose RAW argmax really is one-sided has no sells manufactured for it. The
sentinel is negative rather than 0.0 because 0.0 would mean "call every bar on
this side", the trapdoor DIR_CONF_MIN_FIT_CALLS exists to keep the pooled point
away from.

ONE SWEEP, THREE FITS. FitDirConfBinFor() is shared by the pooled fit and both
per-side fits so the rule cannot differ between them - two copies of a sweep is
how one of them ends up with a different tie-break or a different denominator. The
coverage denominator stays m_dirConfPrimaryBars in all three: using a side's own
call count would make every side's target 100% by construction.

THE OPERATING POINT BELONGS TO THE WEIGHTS IT WAS FITTED FOR, so the per-side pair
travels through every one of the four checkpoint sites the pooled value already
uses - the ensemble joint commit (via the m_eraStatThresholdSide stash, because no
member "won" the era, the vote did), the solo capture at a new best, the mid-run
rollback, and the deploy-the-checkpoint restore. Leaving one behind would pair
checkpointed weights with cuts fitted for a rejected era.

Persisted in the .cfg, appended AFTER the variable-length pins so no existing
reader's offsets move, and length-guarded like every field above it: a .cfg from
before today ends early, both guards yield the UNFITTED sentinel, and the model
runs exactly the policy it was written under. This matters because a DEPLOYED
model runs no further eras - without it, every restart would drop both sides back
to the pooled point and re-silence the sell book on the charts this exists for.

WHAT THIS WILL AND WILL NOT FIX, measured before deploying rather than claimed
after. Fresh numbers taken at 11:45-11:47, after the restart, say the member layer
is NOT where the amplification happens on the charts that are currently firing:

    XAUUSD members, post-operating-point:  2161 Buy / 1692 Sell   = 1.28 : 1
    NAS100 combined VOTE direction:        2654 buy / 2345 sell   = 1.13 : 1
    NAS100 after the ENSEMBLE RUNG:         160 long /    5 short = 32   : 1
    DAX40  after the ENSEMBLE RUNG:          30 long /    0 short

So both the member operating point and the vote's DIRECTION are already balanced,
and the 32:1 appears at the single ensemble rung applied to |vote|. That rung sits
at the very top of the vote distribution - NAS100's strongest vote is 5.5% against
a 5.0% threshold, drawing 43 arrows on 4999 bars - and the top of that
distribution is where the members AGREE, which the era log's own note says is
where they agree WITH THE DRIFT. Cutting there selects drift-aligned calls, and on
this fleet the drift is long.

This commit therefore fixes the member instance of the defect, which is real and
which XAUUSD's 1.28:1 shows is already being held in check there, and it is NOT
expected to move the fleet's 32:1 on its own. The ensemble rung is the same defect
one layer up and is the next change; it is deliberately separate so that each
effect stays attributable, and because a per-side rung must be fitted to a
MEASURED RATE rather than selected on payoff - selecting a rung per side on an
outcome metric is exactly what 5a7d96c retired.

Not retrain-forcing: nothing here is a fingerprint member. Deployed ALONE rather
than bundled with mask v5 as project_next_three_tasks proposed - mask v5 is
retrain-forcing, and measuring the sell book's return needs the SAME mature models
this baseline was taken on (the method project_honest_floor_silenced_four_charts
used). Bundling would have made the effect unattributable.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-03 11:49:57 -04:00
AnimateDread
3c0e2facb2 fix(chart): a signal mark's persisted time was the left edge of its line, not its bar
User report: "some labels are completly off (arrow not at the low/high, and entry
price way higher than candle body and much much away from what the spread would
add)". Not the training label, and not a fill - the only order attempt on the day
was refused by the client. Both halves are one display defect.

A mark is TWO objects since 2026-08-20: an OBJ_TREND segment from t-half to
t+half so it is wide enough to see, plus an OBJ_ARROW glyph. half is
PeriodSeconds * WARRIOR_SIG_LEVEL_HALF_SPAN = 4680 s on H1, so OBJPROP_TIME index
0 of the line is its LEFT EDGE. FIVE call sites read index 0 and every one of
them treated it as the bar time, because "which bar does this mark belong to" had
no single implementation and each site re-derived it from rendering geometry.

MEASURED on the live files, and it COMPOUNDS: Snapshot() wrote t-half, the restore
redrew a line centred on the stored value, and the next save read THAT line's left
edge. Predicted residues are the 10 multiples of 360 s, not the 60 possible
minutes - observed 704 of 704 persisted vote arrows across fifteen files on the
k*1080s ladder, ZERO off it, with k=1 where a mark had been saved once and k up to
6 where it had been carried through sessions. XAUUSD and EURUSD sat at k=6: 7.8 H1
bars adrift, which is why a mark's stored trigger price was nowhere near the candle
it was drawn over - the price was right for a bar 7.8 bars away. The arrow half
lost its low/high anchor for the same reason: a shifted time is not a bar open, so
iBarShift(exact) returned -1 and the glyph fell back to the trigger price, landing
inside the candle body.

WarriorSignalMarkBarTime() is now the one accessor, and it takes the MIDPOINT of
the two ends rather than t0+half: the midpoint recovers t exactly by construction
AND stays correct if the half-span is ever retuned, including for marks already
drawn under the old value. Its t1 <= t0 branch answers correctly for the
single-anchor arrow half too, which the rescan sweep needs. Routed through it:
CVoteArrowStore::Snapshot, CChartUI::SaveChartSignals,
WarriorReconcileVoteCooldown, WarriorLatestVoteArrowTime, and the rescan's
typed-blind scope sweep.

Two more defects the same read was hiding:

  * the cooldown reconciliation deleted by a RECONSTRUCTED name built from the
    shifted time it had just read, so it could not remove a freshly drawn arrow at
    all and its log over-reported kills. The name now travels with the time
    through an insertion sort over both arrays.
  * the rescan's scope sweep gave the two halves of ONE mark two different times,
    so at the window edge it deleted an arrow and left its line - exactly the
    split WarriorDeleteSignalMark exists to prevent.
  * WarriorLatestVoteArrowTime seeded the live cooldown clock 1.3 bars EARLY on
    every restart, so the first signal after a restart could fire inside the
    window the arrow on the chart was enforcing.

WarriorSignalMarkOnBarGrid() stops the drift surviving a restart. Arithmetic
rather than a history search, and the residue is READ off iTime(sym,period,1)
instead of assuming UTC alignment, because where the bar grid sits in epoch
seconds is the broker's day start. It KEEPS on an unknown - no history yet, or a
weekly/monthly frame that is not a modular grid - since deleting on an unknown is
the failure mode that cost this chart 272 of 273 arrows in 2ca32e9. Applied in
both sidecar loaders and in the reconciliation, which is the only layer that
reaches what MT5 restores from profiles\Charts\*.chr.

An off-grid mark is DROPPED, not snapped to the nearest bar. Its price names a bar
that cannot be recovered, so placing it anywhere would assert that a signal fired
on a bar where it did not. Verified in situ on the deployed fleet: nine charts
screened their files at init and dropped 236 of 236 records, every one off-grid,
matching the offline measurement exactly. The persisted arrow history is therefore
gone - it was already misinformation.

Also killed an IMMORTAL zero-price record measured in the USDCAD file.
WarriorPlotSignalLevel rejects price <= 0 so it was never drawn, but the save path
copied the undrawn restore queue straight back to disk every session, so one bad
record survived forever.

Not retrain-forcing: nothing here is a fingerprint member. Deployed as
mark-bartime-1.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-03 11:25:31 -04:00
AnimateDread
8280a7cd40 feat(fleet): open the charts the census says are worth having
The census answers which instruments carry deep history at a spread the book can
cover; this is the half that acts on it, so adding an instrument is a reviewable
list in source control rather than six manual chart operations nobody can
reconstruct later.

BTCUSD, GBPUSD, NAS100, DAX40 - every one with more than twelve years of server
H1 and a sampled spread under 3% of ATR, against a measured ride book of +0.05 to
+0.63 ATR per call. BTCUSD is the cheapest deep instrument the broker offers
(0.63%) and the only one in a different asset class, which is worth the most to a
pooled certificate that declines the diversification credit. DAX40 trades a
different session, so its bars are not the same hours as everything else.

IDEMPOTENT: a chart is opened only when none exists on that symbol and period, so
a restart re-opens nothing and the list can be edited freely. That is what makes
it safe to leave armed.

THE TEMPLATE CARRIES THE EA. ChartSaveTemplate on the running chart captures this
Expert Advisor and its inputs, so a new member comes up configured exactly like
the one that spawned it - which is the point: a fleet whose members differ by
attach order is how four charts ended up on a 10-bar cooldown and two on 30.

Leased like the census and the alt-data fetch, because six instances would each
try to open the same four charts - and the charts this opens initialise an EA that
reaches this same code.

COST, STATED IN THE FILE: four added charts are sixteen more models on a six-core
box already at ~72% with six charts. Expect eras to slow across the whole fleet.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-03 08:26:39 -04:00
AnimateDread
7a5dddc61a fix(census): one chart, not six - and sample the spread instead of snapshotting it
TWO DEFECTS IN MY OWN CENSUS, both found by reading its output rather than by the
compiler.

IT RAN ON ALL SIX CHARTS. The guard was a plain global, and MQL5 globals are per
PROGRAM INSTANCE - six charts are six instances, so every one of them walked all
64 symbols and overwrote the same file. Replaced with the atomic
GlobalVariableSetOnCondition lease that System\AltDataFetch.mqh already uses for
exactly this problem: six charts racing one shared file. Losing the lease is the
correct outcome, not an error.

THE SPREAD WAS ONE SNAPSHOT. Cost is the number this whole decision turns on - the
measured ride book runs +0.05 to +0.63 ATR per call, so an instrument costing a
tenth of an ATR a round turn has already spent most of what it could earn - and a
single reading is one moment of one session. Spreads widen at rollover and around
news. Now ten readings thirty seconds apart, reporting mean AND max, because the
max is what says whether an instrument is quietly untradeable at the wrong hour.

A reading only counts behind a real two-sided quote; a symbol with none reads
NO-QUOTE rather than 0, since 0 would rank it as the cheapest instrument on offer.
That is the same error the previous commit fixed one layer up, and the guard now
sits at the sample rather than only at the report.

WHAT THE FIRST GOOD RUN ALREADY SETTLED: 64 symbols openable both ways, and all 64
report a SERVER first-H1 date with server == local. So depth is real, not a sync
artifact - the broker genuinely offers deep H1 on twelve instruments and added the
other fifty-two in July/August 2026 with no history at all. The expansion universe
is six symbols, not fifty-eight.

Also promotes the derived-cooldown line out of PrintVerbose. It reports a change
to the TRADING POLICY, and this codebase's rule is that a line reporting a state
change never sits behind the verbosity gate - the tier-ladder restore learned that
the expensive way.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-03 07:15:09 -04:00
AnimateDread
81efd64a7f feat(signal): derive the cooldown from the measured label lifespan, and census the broker's symbols
TWO CHANGES, ONE CAUSE: a per-chart input could not reach the fleet, and the fleet
had no data on which symbols were worth adding.

THE COOLDOWN IS NOW DERIVED, NOT SET. An input was the wrong shape for it twice:

  * It is a measurable property of the LABEL, not a preference. The leg-ride label
    resolves when the ZigZag leg flips, so the mean label lifespan IS the average
    leg duration in bars. Two calls closer together than that concern the SAME
    leg - the second pays a second spread for a move the first already owns. One
    lifespan apart is where consecutive trades concern DISTINCT legs, which is
    also what makes the deploy gate's independence assumption exact rather than
    approximate: EffectiveSampleSizeDeclustered's divisor becomes 1.
  * An input could not reach an attached EA. MT5 stores inputs per chart, so
    "raising Signal_CooldownBars from 10 to 30 changed nothing on six live
    charts" - and yesterday the fleet was found running 10 on four charts and 30
    on two, two different trading policies inherited from attach order rather
    than chosen. A derived value cannot drift that way.

The fraction is 1.0 because the argument picks it: less re-admits same-leg
duplicates, more declines distinct legs for no stated reason. At the measured
33-34.5 bar lifespan it lands within a few bars of the 30 the default intended.
Recomputed wherever the lifespan is measured - the label-cache build - so the
measurement and its consumer cannot drift apart. SignalCooldownOverrideBars still
wins, and switching declustering off entirely is still possible.

THE SYMBOL CENSUS answers "which symbols are worth adding" with data. The binding
constraint on this system is independent observations, and instruments are the
only lever that multiplies them, so it writes the three numbers that decide it:
history depth, spread against ATR, and whether both sides can be opened at all (a
close-only symbol can never satisfy the gate's two-sidedness test).

AND IT RUNS ON THE TIMER, NOT AT INIT, which the first version got wrong. At init
the terminal has just reconnected and nothing has a quote, so SYMBOL_SPREAD reads
0 everywhere - the first run duly ranked thirty untradeable-on-cost symbols as the
cheapest the broker offers. Caught by noticing GBPUSD reported a zero spread
beside EURUSD's 2 points. Now deferred three minutes, and a symbol with no tick is
written NO-QUOTE rather than a number, so the column cannot be sorted on by
mistake. It selects nothing: sixty symbols added to Market Watch is a change to
the operator's terminal, and this is a report.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-03 07:05:49 -04:00
AnimateDread
8fd2756aa5 fix(gate): the member gate certified a population the EA never trades
The deploy gate judged m_oos.DirCalls() - every bar the prior-corrected posterior
fired - and deflated it by the full 34.5-bar label lifespan. The EA does not trade
that set. It trades what survives declustering: live, a call the cooldown rejects
has its signal zeroed before the vote is published, so it produces no arrow, no
vote and no position.

The honest counts were already being computed, by the same CSignalDeclusterPolicy
object the live signal applies, on the same adjusted decision it feeds - printed
every era as "TRADED (declustered)". The gate simply never read them.

BOTH HALVES OF THE ERROR POINTED THE SAME WAY, which is what made it expensive:

  * the traded set is measurably CLEANER. Live USDJPY eras: judged 24% where
    traded was 26%, judged 26% where traded was 27%. The gate was reading a
    precision the EA would never have realised.
  * the traded set is far more INDEPENDENT. The cooldown is 30 bars against a
    34.5-bar label, so consecutive traded labels overlap by at most 4.5 bars - the
    stream is ~87% independent, not ~3%. Deflating it by the full lifespan applies
    a correction the cooldown has already made.

On a live USDJPY era that is the whole verdict: judged 24% against a 25.4% bar
FAILS; traded 27% against a 24.7% bar PASSES.

AND THE SELECTION IS UNBIASED, which is what makes the traded precision usable at
all: the declustering keeps the chronologically FIRST bar of each run, never the
highest-confidence one, so this is not cherry-picking winners.

COVERAGE STAYS ON THE SIGNAL POPULATION, and that split is now explicit in the
signature rather than implied. Coverage asks "did this model call often enough to
be a strategy", which is a question about the signal; precision asks "were the
calls it took right", which is a question about the account. Coverage cannot move
to the traded set for a structural reason: a 30-bar cooldown caps traded coverage
at 1/30 = 3.3% while the floor is a quarter of a ~20% base rate, so every chart
would fail it forever for reasons unrelated to the model.

THE ENSEMBLE GATE STILL HAS THIS DEFECT and is passed through unchanged, on
purpose. Its OOS rows carry each member's raw adjusted vote - the decluster replay
is per-member state that never reaches the shared vote buffer - so there is no
traded population there to read yet. Fixing it needs the per-member decluster
decision carried on the vote row. One gate at a time, so the effect of this one
stays attributable.

Not retrain-forcing: gate arithmetic only, no fingerprint field moves.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-03 06:14:28 -04:00
AnimateDread
6d266d4e62 feat(features): every price column here is memoryless - fractional differentiation
AFML ch.5. Every price-derived feature in this set is a FULL difference:
(close-open)/atr, (close-MA)/atr, the 20-bar return, the leg extension.
Differencing is what makes a series learnable - a model fitted on 2019 EURUSD
levels cannot read 2026 ones - but a first difference is memoryless BY
CONSTRUCTION. It keeps the last step and throws away the series. That is the
trade this feature set has been making silently at every column.

Fractional differentiation is the observation that the exponent need not be an
integer. (1-B)^d for 0<d<1 interpolates between the raw level (all memory, not
stationary, useless to a learner) and the return (stationary, no memory). Three
orders are emitted - d = 0.3/0.5/0.7 - in ATR-relative units.

A LADDER, NOT A CHOSEN d. AFML picks the minimum d passing an ADF test; there is
no ADF test here, and adding one to tune a feature's parameter per instrument per
era would fit the feature to the data before the model saw it. Three fixed orders
go in and the FLEET KEEP-SCREEN votes on them, exactly as it has already deleted
five ZigZag geometry columns and demoted the alt-data block. Kept by construction
in mask v4 for one outing only: the screen reports only on columns that are
EMITTED, so a block masked off at birth can never be measured.

THE BUG THAT WOULD HAVE SHIPPED THREE DEAD COLUMNS, caught before deploy by
checking the arithmetic at realistic price levels rather than on a series
starting at zero:

The weights of (1-B)^d sum to zero in the LIMIT - that is what makes it a
differencing operator. TRUNCATED at 64 bars they do not: the residual is 0.222 at
d=0.3. That residual multiplies the LOG PRICE LEVEL, so the raw sum carries
0.222 x log(2400) = 1.73 on XAUUSD against 0.222 x log(1.08) = 0.017 on EURUSD -
an instrument-identity constant hundreds of times larger than the signal. Divided
by the relative ATR it pins every bar to the clamp: measured 100% of XAUUSD and
SP500 bars clipped, 68% of EURUSD. Three constant columns that FEATURE HEALTH
would have reported only after a full retrain had been spent on them.

Anchoring every term at this bar's log price removes exactly the level component
and leaves a fracdiff-weighted combination of the multi-horizon RETURNS ending
here - stationary, scale-free, same meaning on every instrument, and the long
memory intact. Same series after: mean ~0, sd ~1.1-1.3, nothing clipped.

The weight recursion is checked against two known values: at d=1 it terminates to
[1,-1,0,...], the plain first difference, and at d=0.5 it reproduces the standard
expansion of (1-B)^0.5 to six terms.

REJECTS RATHER THAN DEGRADES at the oldest edge, unlike the swing and volume
windows beside it. A 30-bar Donchian range is still a Donchian range; a
fractional difference over a short window is a DIFFERENT OPERATOR - a different
effective d - reported in the same column, which is a silently wrong number
rather than a degraded one. Costs ~64 bars of 130,000.

RETRAIN-FORCING twice over: the input width is field 4 of the fingerprint and
FEATURE_MASK_VERSION rides it too.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-02 23:12:02 -04:00
AnimateDread
f063cfd6ba fix(calibration): the confidence rescale was blended with a per-SAMPLE time constant, once per era
m_confidenceCalScale has been frozen at its 1.0 constructor default for the life
of the mechanism, and the units error that froze it is one line:

    m_confidenceCalScale += (eraScale - m_confidenceCalScale)
                            / Net.recentAverageSmoothingFactor;

recentAverageSmoothingFactor is 10000. It is a PER-TRAINING-SAMPLE constant - it
exists to average the network error over ten thousand samples inside backprop.
Applied ONCE PER ERA it moves the value by one hundredth of one percent, so after
a hundred eras the scale has travelled 1% of the way to its target and after a
thousand it is still not halfway.

That reframes what was already known. ConfidenceBridge.mqh records that the
confidence is miscalibrated and that five trade-management modes were deleted
because of it. The mechanism meant to fix it was not merely shape-blind, as the
isotonic curve's commit message argued - it was NUMERICALLY INERT, and never had
the chance to correct anything at all. A quantity blended per era needs a per-era
time constant; borrowing one from a per-sample loop reads as a working mechanism
in every review, because the line is shaped exactly like a working EMA.

The new curve inherited the same divisor when it was written yesterday, which
would have frozen it at its first fit - visible in the log as a "carried" column
identical to "refit" to four decimals on every chart. Both now use CAL_BLEND_ERAS
(5 eras): slow enough that one noisy band cannot swing the live number, fast
enough to track a model whose output distribution moves every era.

VERIFY, DON'T ASSERT: the curve's log line now prints the scalar's current value,
so the first line after this deploy reports the number restored from .stats - the
value it reached over that model's entire history. 1.000 is the claim above,
measured rather than argued from the arithmetic.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-02 22:57:43 -04:00
AnimateDread
5ec8902020 fix(gate): the zero-skill floor was still halved one level down - and the measurement sizing actually needs
THREE FINDINGS, ALL FROM READING THE FLEET'S OWN LOG RATHER THAN THE COMPILER.

1. THE MEMBER DEPLOY GATE STILL USED THE 3-CLASS ZERO-SKILL FLOOR.

The ensemble gate's chance rate was corrected on 2026-09-03 when the head stopped
choosing sides. SOosTally::ChancePrecPct - one level down - was not, and it feeds
more than the ensemble copy does: the MEMBER deploy gate, the certified pair
HasDemonstratedEdge() reads to decide whether a member may vote at all, and the
cross-instrument pooled certificate. On a live H1 window it read 12.5% where the
honest always-ride floor is 21.1%, so every member showed "+13pp edge, PASSES"
and was admitted to the vote. A model with no skill whatsoever cleared it.

The same halved floor was in BaselineComparator's Alglib comparison, which scored
the forest and the linear baseline against a bar half the height of the one the
net is judged by.

ChancePrecPct now TAKES THE POLICY AS A PARAMETER WITH NO DEFAULT, so the next
head change is a compile error at every reader instead of a silent wrong answer
at some of them. That is the whole lesson of finding this one a week late: the
first fix was applied where the bug was noticed, not everywhere the assumption
lived.

The era line said "the gate ranks on the LARGER of the two" for that entire week,
on six charts, every era. It now states the policy actually in force.

2. THE ERA LINE PRINTED +/-1.79e308 IN THE MIDDLE OF EVERY RECORD. The binary
head has no third output neuron, so slot 2's min/max kept their DBL_MAX ctor
values and were formatted anyway. Width-aware now, and the slots are labelled
PAYS/DOESNT rather than B/S/N, which is what they hold.

3. THE SIZING QUESTION NEEDED A DIFFERENT MEASUREMENT THAN THE ONE I BUILT.

The reliability curve says the confidence is now HONEST - carried, out-of-sample,
ECE 38pp -> 1.3-2.4pp. It says nothing about whether it RANKS, and ranking is
what a bet size needs. The existing conviction curve cannot answer it either:
coverage collapses above the lowest rung, so every fired call sits in one bucket
and there is no curve to read.

So the ensemble now carries the calibrated confidence per OOS row - the mean over
members that actually called, which is exactly what LiveSignedConfidence()
publishes - and the era verdict splits the certified rung's fired calls at their
median confidence and compares what the two halves earned, in ATR per call.

A MEDIAN SPLIT, NOT A DECILE CURVE, and the reason is power: per-call SD is ~3
ATR and the labels overlap ~34 bars, so a decile of a few hundred raw calls holds
under ten INDEPENDENT ones and its error bar is wider than the whole book. Ten
noisy points would invite the best-of-N reading this project has already crowned
four times. Two halves is the most the data can be asked for, and it is reported
with its standard error on INDEPENDENT counts plus the size of difference this
window could ever resolve - so "not measurable" is distinguishable from "no
effect", which is a statement about the data rather than a verdict on the idea.

Nothing here sizes anything. This is the evidence ConfidenceBridge.mqh's standing
rule demands before anything may.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-02 22:50:19 -04:00
AnimateDread
18ada65d85 feat(calibration): the refit column is not evidence - carry the curve and measure the live range
Two corrections to what the first commit reported, both found by reading its own
first live output rather than by the compiler.

THE "AFTER" NUMBERS WERE IN-SAMPLE. ECE 44.7pp -> 0.5pp was the curve scored on
the very band it had just been fitted on, and isotonic regression is expressive
enough to drive that near zero whatever the data says. It is the attainable
floor, not a result. Quoting it would be the calibration-slice error - a
threshold fitted on memorised bars - one layer along.

The fix is free, because the previous era's curve is already sitting there when
the new one is fitted: score it too. That curve was fitted on EARLIER bands and
has never seen these bars, so it is a genuine out-of-sample calibration
measurement. The log now reads raw -> carried -> refit, and says in the line
itself that the middle column is the one to read.

THE SPREAD, NOT THE ERROR, IS THE SIZING QUESTION. The first fit showed claimed
0.75 and claimed 0.99 mapping to the SAME calibrated 0.273 - the curve is flat
at the top, meaning higher confidence there does not mean higher accuracy. A
well-calibrated constant is still a constant: no amount of calibration makes a
flat map rank anything, and a bet scaled by it would be a bet scaled by noise -
the precise error that killed the five confidence-scaled modes on 2026-08-25.

So the report now restricts to the calls the operating point actually ADMITTED -
the only ones that can become a trade - and states the calibrated probability at
the lowest and highest admitted bin, their spread, and the map's mean against the
realised rate on that same set. Carried as its own counts rather than derived by
cutting the histogram at a magnitude, because the operating point is a MARGIN and
this histogram is keyed on a MAGNITUDE; those are monotone in each other on the
binary head and not on the 3-class one, and a reparameterisation guess is exactly
the kind of thing that reads as a measurement.

This is the number the sizing decision will be made on, and it is deliberately
being gathered BEFORE anything is built on it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-02 22:36:08 -04:00
AnimateDread
92ccf20385 feat(calibration): a scalar has no shape - the isotonic reliability curve
The model's confidence was corrected by ONE number: accuracy divided by mean
claimed confidence, EMA-blended per era. That can move the reliability curve up
or down and can do nothing else. A model that is honest at 0.55 and wildly
over-confident at 0.95 has a SHAPE problem, and no scalar has a shape.

The scaffolding for the fix already existed and was better than expected.
RunCalibrationPass already walks a PURGED, held-out band with batch norm frozen,
visiting each bar exactly once, and fills a 50-bin (margin, hit-rate) histogram.
That is a reliability diagram on out-of-fold predictions. Isotonic regression is
a fit over data already being collected, on a band already purged - no new walk,
no new holdout, no new cost.

WHAT THIS ADDS

  * A SECOND histogram in the same walk, keyed on |dPrevSignal| - the magnitude
    the consumer actually holds - not on the winner-vs-rival margin. Deliberately
    not a re-key of the existing one: on the 3-class head the two statistics are
    not the same quantity, and FitDirConfThreshold's own note records what
    happened the last time one curve served two fits.
  * FitCalibrationCurve: pool-adjacent-violators over the occupied bins, weighted
    by call count. The least-squares monotone fit (Ayer et al. 1955); the
    monotonicity constraint IS the regularisation, so there is no smoothing
    parameter to tune and it cannot fit a shape the data does not show.
  * EMA-blended across eras with the same smoothing the scalar used. A convex
    combination of monotone sequences is monotone, so blending costs nothing the
    fit exists to impose.
  * Brier, ECE and MCE reported before and after the map, on the band the map was
    fitted on. REPORTED, NOT OPTIMISED - nothing selects on them. A model that was
    already calibrated shows all three pairs unchanged, which is the outcome that
    says this map is not needed.
  * Persisted (.stats WSTE), because the curve is produced only by a completed
    calibration band and a DEPLOYED model runs no more eras - the same failure the
    tier ladder had. The bin count leads the block so a future
    DIR_CONF_THRESHOLD_BINS change is a mismatch the reader detects rather than
    fifty doubles landing in the wrong slots.

THE SCALAR STAYS as the unfitted-model answer, and the two are never applied
together: they answer the identical question, the curve per magnitude and the
scalar on average, so stacking them would correct the same error twice. Same
discipline as the logit adjustment being backward-pass only.

NOT RETRAIN-FORCING. The fingerprint is untouched, the .nnw is untouched, and a
WSTD file still loads - it simply comes back with no curve and keeps the scalar.
The fleet resumes exactly where it was.

STILL TELEMETRY. Variables\ConfidenceBridge.mqh carries a standing rule that
nothing there may steer a trade, and CalibratedConfidenceMagnitude has exactly
one consumer: the trade journal's aiConfidence bucket. Five confidence-scaled
trade-management modes were deleted on 2026-08-25 for one stated reason - the
confidence was known to be miscalibrated, so they scaled money by a quantity
whose units were never established. This is the measurement that establishes
them. Sizing is deliberately NOT in this commit: shipped alongside its own
calibration it would be untestable, because if the book moves nothing says which
half did it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-02 22:30:51 -04:00
AnimateDread
84789ea992 fix(head): the binary logit branch printed the 3-class banner
ApplyLogitAdjustment's meta-head branch sat BELOW the 3-class tau computation
and its log block. The offsets installed were correct - the branch returns
before the 3-class ones are written - but the 3-class banner printed first and
latched m_logitAdjustLogged, so every era log claimed "APPLIED across all
three" while the head had two logits, and the binary line was unreachable.
A log line that misdescribes a live gate-adjacent path is the same class of
defect as a stale comment.

Moved above the 3-class work. Also corrects this file's own claim about the
imbalance against the measurement it now has: the H1 fleet's shares are
9.4/9.0/81.6, so the binary split is ~4.4:1 rather than the ~3:1 estimated
before the run - and tau goes from 0.57 capped on three classes to ~0.86 on
two, which is the point: most of the correction gets through where most of it
was being clipped.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-02 21:50:05 -04:00
AnimateDread
8d6cff7fcb feat(head): the side was never a prediction - the binary meta-label head
The leg-ride target's SIDE is the direction of the ZigZag leg in progress as
of that bar, readable from bars t and older with no lookahead. The 3-class
head made the network re-derive it anyway: that is Lopez de Prado's PRIMARY
MODEL being learned instead of used, and it cost on four axes at once -
chance at 33% instead of 50%, an 8:1 imbalance instead of 3:1 against a tau
already pinned at its cap, every Buy row spent as evidence about "is this a
long leg" rather than about payoff, and a directional error scored the same
as a payoff error though only one of them was a question we asked.

The head is now binary: slot 0 "riding this leg to the flip pays at least
LEG_LABEL_MIN_RIDE_ATR", slot 1 "it does not". The side comes from
LegDirAsOf() at read time.

TWO OUTPUT NEURONS, NOT ONE, and that is what made this small. Both backward
passes in AI\Impl\NetForward.mqh already carry a `total == 2` softmax+CE arm,
left in deliberately when the old meta head was removed on 2026-08-25 - and a
2-class softmax IS a logistic/BCE head, the logit difference being the
log-odds. So this reuses the exact gradient path the 3-class head uses instead
of growing a second one.

NOTHING DOWNSTREAM CHANGED. ApplyClassificationSoftmax() expands (pGo, side)
back into the [pBuy, pSell, pNeutral] triple every reader already consumes, so
Argmax3, the operating-point histogram, per-class recall, confidence tiers,
the vote currency and the chart arrows are untouched. Neutral stops being a
class the net competes for and becomes what it always meant: pGo below the
operating point. The margin the threshold is expressed in becomes 2*pGo-1,
monotone in pGo, so the calibration walk fits the same statistic.

THE HEAD STAYS BOUNDED, deliberately, and the 2026-07-27 unbinding retry is
NOT bundled here. Reading the gradient showed why it need not be: the softmax
arm OVERWRITES the output neuron's gradient with (target - softmax), so the
sigmoid derivative and MIN_ACTIVATION_DERIVATIVE are already bypassed at the
head. What SIGMOID x CLASS_LOGIT_SCALE actually costs is p in [0.0025,0.9975]
- three orders of magnitude wider than the band where the live question is
0.35 versus 0.60. Unbinding has its own failure history and deserves its own
measurement; batch norm before the head, its precondition, already ships.

THE ZERO-SKILL FLOOR HAD TO MOVE WITH IT. The rate gate's chance was the
better of always-Buy and always-Sell. This head cannot choose a side, so its
no-skill policy is ALWAYS-RIDE - every bar taken in its own leg's direction -
which is right on EVERY directional-label bar, not the better half. Left
alone it would have handed a model with no skill whatsoever a ~+11pp edge.
The book gate already measured against always-ride; the rate gate now agrees
with it about what zero skill means. Same correction in the module-weight
shrinkage prior (Lifecycle.mqh's 50.0 side coin-flip).

RETRAIN-FORCING twice over: |MHEAD:1 joins the fingerprint and the output
count is field 6 of the .nnw filename, so no existing model or pool row can
be adopted. Verified live - all six charts rejected every peer file by
fingerprint and restarted from era 0.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-02 21:46:50 -04:00
AnimateDread
8d7cbdbf9d feat(training): sample the queue by uniqueness instead of weighting it - the update itself was the cost
Weighting by uniqueness (8351c39) redistributed gradient mass but still spent an
update on every redundant label. Lopez de Prado's actual prescription for
redundancy is the other one: bag size = average uniqueness (AFML ch.4). A label
now enters an era's training queue with probability proportional to its own
uniqueness, and carries no uniqueness weight once it is in.

UNIQUE_WEIGHT_VERSION becomes a MODE so the two can never stack - one integer
cannot be two values, which is the same class of mistake as the focal-loss and
logit-adjustment stacking this codebase already removed. 0 off, 1 weight, 2
sample. Mode 2 is the default.

- The draw is STRATIFIED BY CLASS. Uniqueness is correlated with the label here,
  since a label's span IS its ride and whether a ride clears 1 ATR is what
  separates Buy/Sell from Neutral. An unstratified draw would shift the training
  set's class balance away from the priors the logit-adjusted loss corrects
  against - the exact failure the queue site's own comment warns about. Dividing
  by the class mean equalises the expected acceptance rate across classes.
- It also decorrelates the ensemble: each member draws its own subset each era,
  so the four models no longer see one identical stream. That is the diversity
  lever the era report has asked for every era (4 models worth 1.3-1.9).
- Measured: eras now queue 7.4-9.0% of eligible bars, matching each chart's mean
  uniqueness, and era time fell from ~107s to ~22-27s on the index and oil.

One bug shipped and was caught in the first run. The probability was
rate * u / classMean, whose mean inside a class is 1.0 by construction, so at
rate 1.0 it kept 73% of bars and the reduction never happened. The class mean is
the right stratifier but the wrong scale; the scale is the global mean. Fixed
before this commit, and the rate now rides in the fingerprint (UWGT:2:R100)
because a model trained on a twelfth of the bars is not interchangeable with one
trained on all of them.

RETRAIN-FORCING. Compiled 0/0. Deployed 20:15 as uniq-sample-2.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-02 20:18:40 -04:00
AnimateDread
8351c3933e feat(training): weight every sample by its label's uniqueness - the loss was counting 34x redundancy
Every statistic in this program deflates overlapping labels. The training loop
did not. EffectiveSampleSize() is called in exactly two places, the feature
keep-screen and the deploy gate, and both are reporting paths; the gradient step
weighted each sample by |ride| alone, with no term for how many other labels are
made of the same bars.

Measured on the H1 fleet: mean label overlap 31.6-37.4 bars. USDJPY carries
125,214 labels that this same binary reports as ~3,344 independent ones. The
optimiser was told it had thirty-four times more evidence than it has, which is
the textbook cause of a strong in-sample fit with a thin out-of-sample book -
the pattern every era log has shown since this label shipped.

The correction is Lopez de Prado, Advances in Financial Machine Learning ch.4:
average uniqueness is 1/concurrency averaged over the bars a label spans, and
the prescribed weight is uniqueness x attributed return, i.e. the product of the
new Variables\UniqueWeight.mqh and the existing MoneyWeight.mqh.

- ComputeLabelUniqueness() runs once when the label cache pre-build completes,
  O(bars) via a difference array for concurrency plus a prefix sum over 1/c, so
  a 125k-label cache costs two sweeps rather than millions of span walks.
- On the leg-ride target the two weights are anti-correlated: a label's span IS
  its ride, so a long leg earns a big money weight and sits where concurrency is
  highest. Money weight alone concentrated gradient on exactly the least
  independent evidence in the set. The product is re-clamped to [0.25, 4.00]
  because two clamped factors multiply to a 16x tail.
- Measured even when the knob is off, so the report can show the spread this
  would apply without training on it. First run: mean uniqueness 0.076-0.085.

This REDISTRIBUTES gradient mass; it does not reduce the update count. Training
on fewer, more independent samples is the sequential bootstrap (ch.4 s.4.5) and
gets its own commit and its own measurement.

UNIQUE_WEIGHT_VERSION 0 restores the prior behaviour and the prior fingerprint
exactly. RETRAIN-FORCING (|UWGT:1). Compiled 0/0. Deployed 19:40 as
uniq-weight-1; all six H1 charts from era 0.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-02 19:42:43 -04:00
AnimateDread
3aa15c8b4b feat(altdata): publication stamps by series, alt block back in, one pool for all six charts
Operator asked for the alt data to be properly mapped on H1. Three findings:

1. The as-of join was already whole-day: any H1 bar of day D reads row D, and the
   window layout (ALTW:2) carries that reading once per window at the anchor.
   What was wrong was the ROW DATE. Every FRED series was stamped "knowable next
   day", which is right for a market close and five to six weeks early for a
   monthly print: July CPI (dated 07-01) entered the export on 07-02 and was
   released 08-12. Unemployment the same; the H.10 dollar index (weekly, posted
   the following Monday) a week early; the effective funds rate a day early. On
   H1 that is ~1,000 bars of a value nobody had, in three of the twelve fleet
   columns. FredPublishLagDays() stamps by series (CPI +48d, UNRATE +40d,
   DTWEXBGS +8d, DFF +2d, closes +1d), cached rows are re-stamped on load, and
   ALTFETCH_EXPORT_VERSION (a .ver sidecar beside each export) forces one rebuild
   at the next init so every chart reads the corrected export immediately.
   Verified on XAUUSD_D1.csv: mac_cpi now changes on 08-18, mac_unemp on 08-10.

2. The alt block was not reaching the model at all. Keep mask v2 dropped all
   twelve alt columns on a screen measured under the pivot label on H4, and the
   screen only reports on emitted columns. v3 emits them again (28 of 47 columns,
   input width 168); the H1 keep-screen will say which of them clear.
   The window dedupe now uses the EMITTED alt width, not the panel's: under v2
   it placed a 12-wide block over the last twelve of sixteen emitted columns.

3. The training pool ran as two groups because the cross-asset block encoded
   index mode (base == quote: SP500, and this broker's XAUUSD/XTIUSD) with a
   different meaning per slot than FX mode, and the fingerprint tagged it
   ":IDX2". U3 gives both modes one layout (proxy fast/slow in 0/2, denomination
   fast/slow in 1/3, own move minus the proxy-vs-denomination cross in 4), so all
   six charts print the same fingerprint and pool together.

RETRAIN-FORCING (XA:6:U3, ALTV:2, FMASK:3). Deployed 15:00 as alt-stamps-1; all
six H1 charts from era 0, one shared fingerprint, exports rebuilt at v2.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-02 15:02:27 -04:00
AnimateDread
8c37869c63 fix(init): a warm re-init kept the H4 run's globals; the training pool adopted other timeframes
Operator report 2026-09-02: the six charts were switched from H4 to H1 and the
status label kept showing the H4 run. MetaTrader does not unload the program on a
timeframe, symbol or input change - it runs OnDeinit and OnInit inside the same
instance and every file-scope global survives the pair. The members' converged
flags, the live vote line, the panel rows, the shared best-era/plateau state and
the once-per-chart report flags all belonged to the models just torn down.

- Warrior_EA.mq5: WarriorResetWarmReinitState() runs first in OnInit and puts
  every such global back to its cold-start default. Kept on purpose: the chart's
  book magic (positions opened before the switch stay owned), the alt-data fetch
  throttle (rate-limited APIs), the tester profile, the RNG, the OpenCL flags.
- TrainingPool.mqh: peers must be the caller's own timeframe. The fingerprint
  does not carry the period, so the first H1 census adopted 39,754 H4 rows from
  three peer files and credited them to the capacity budget. Files are named
  SYMBOL_PERIOD.bin; the suffix decides, and the reader names the rejection.

Build tag warm-reinit-1. Compiled 0 errors / 0 warnings. RETRAIN-NEUTRAL.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-02 14:25:20 -04:00
AnimateDread
980f10b60c feat(label): ride the leg, exit on the flip - the leg-ride target replaces the pivot-event target
The pivot label paid +1.6 ATR on a hit and -1.7 on a miss at 50-65% precision:
a zero book by arithmetic, because a "bottom" that is not one is a move that
kept going, and the model called it because a big move had just happened.
This target never asks for a turn. Direction is the leg in progress, known on
the bar; the label is whether riding it from here until the leg flips pays at
least LEG_LABEL_MIN_RIDE_ATR (1.0). The flip is the exit the EA now places.

- Labeling/LegState.mqh: a line-for-line replica of ZigZag.mq5 (12/5/3) run
  over bars <= t only, so "which leg am I in" is what the chart showed on t,
  never the final buffer. Verified against the real indicator at every init
  (LEG STATE REPLICA line) and by Tests/Test_LegState.mq5 on hand-built bars.
- LegRideLabel (Labels.mqh): ride = legDir * (close[flip] - close[t]) / ATR,
  the flip found by asking the same as-of function of each newer bar in turn.
  Unresolved until the leg has flipped inside loaded history. Lifespan = the
  leg, so EffectiveSampleSize deflates honestly (~3x fewer than the window
  constant claimed); the topology's overlap is the median leg again.
- Money weight = |ride|. Online step's exit bar = the flip.
- Three as-of leg features in the swing block (direction, extension in ATR,
  age); FMASK:2 keeps them. Fingerprint TGT:LEG1:10 - RETRAIN-FORCING.
- Era verdict: the book WarriorRungBookProfitable gates is the RIDE (entry at
  the call, exit at the flip), printed per rung and at the certified rung
  beside ALWAYS-RIDE (zero skill) and the ORACLE RIDE (the ceiling). Fixed-
  horizon payoff stays as the signal diagnostic. The by-distance profile,
  its slot arithmetic and the reversal-exit pass are deleted (measured:
  the vote's reversals land 33-137 bars late; dead).
- Live: Exit_On_Leg_Flip (default on, not in the fingerprint) closes when
  the as-of leg flips against the position and places no take-profit; the
  measured stop stays. Build tag leg-ride-1.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-02 11:02:47 -04:00
AnimateDread
39e4bb587a refactor(signal): delete the barrier sweep, the automatic direction policy and the pivot-geometry features
Three layers out, ~1,000 lines, RETRAIN-NEUTRAL - the fingerprint and the emitted feature vector are
byte-identical, and a WSTC .stats still restores under the new WSTD reader.

BARRIER SWEEP. The 64-cell (stop, target) grid, its 999-draw Westfall-Young permutation null, the
first-touch offsets recorded per OOS row, the swept-pair persistence (WSTA) and the tier-1 branch of
ResolveBarrierMultiplier are gone. It never cleared its own null on any chart (p 0.19 to 1.00), so it
gated nothing and cost a grid plus 999 rescans per era. Live stops and targets come from the mean
excursions at the hold horizon, as they did in practice.

DIRECTION IS THE INPUT. WarriorEffectiveDirection() returns tradingdirection. The bar-body drift
screen (Signal\DriftScreen.mqh) and the 2-SE by-side expectancy test that outranked it are removed
with their unfiltered accumulator and persistence (WSTB/WSTC). On a fleet where no side clears 2 SE
on payoff the test could never speak - SP500 and XAUUSD never once printed its verdict across two
days of logs - so the screen decided by default. The era verdict now certifies BOTH sides whatever
the input says: the checkpoint no longer depends on a per-chart input, so one trained model serves
LONG_ONLY, SHORT_ONLY and BOTH and the input can be optimized in the tester without a retrain. The
blocked side's vote still closes a position under Exit_On_Reversal_Vote.

PIVOT GEOMETRY. The five confirmed-pivot swing features (leg direction, distance since pivot,
prior-leg magnitude, retracement ratio, bars since pivot) scored 4/2/0/2/1 of 24 on the keep-screen
and the mask had already dropped them from every emitted vector. Deleted from the builder; the kept
trend-position columns are swing[0..3] now and the vector is unchanged.

SWING LEG SIZE, measured. MeasureSwingGeometry now prints the leg-size distribution in ATR at the
leg's start pivot. Offline on the archived H4 feeds (stock ZigZag 12/5/3, ATR 20): median leg 4.5-5.0
ATR, p90 9-10, max 35-76; median length 13-14 bars, a third of legs outlast the 18-bar hold. The era
report's ORACLE (+1.6 ATR) is a fixed-horizon close-to-close capture from the call bar, not the leg -
a perfect caller entering at the pivot and holding 18 bars earns ~1.9-2.0 ATR on the same feeds.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-02 09:55:44 -04:00
AnimateDread
d82e7b8ed5 feat(training): money-weighted loss - cross-entropy could not see magnitude
The objective was never aligned with the book. Cross-entropy against the
pivot label scores a 0.2 ATR pivot and a 3 ATR pivot identically, so the
network optimises how OFTEN it is right and nothing tells it that being
wrong on the big ones is what kills the P&L. That is exactly the measured
pathology: precision 50.4-65.2% against a 13.4-14.5% base rate (23-52
sigma, unambiguous) while the book earns +0.002 to +0.045 ATR per call and
loses to always-long on five of six charts.

The label is NOT the problem. The oracle - a perfect caller of this exact
label - earns +1.167 ATR @5 and +1.627 @19 over 3,986 calls. The target is
rich and we capture ~2% of it.

Per-sample weight = |forward return at the hold horizon| / mean(|r|),
clamped to [0.25, 4.00]. Normalised by the mean, so an era trains on the
same TOTAL gradient mass as before - redistributed from cheap bars to
expensive ones - and the effective learning rate does not move.

- Variables/MoneyWeight.mqh: the knob (MONEY_WEIGHT_VERSION 0 restores
  unweighted training and the pre-change fingerprint exactly), the
  transform, the fingerprint tag. Runtime predicate, not #if - MQL5 has no
  #if <expression>.
- The floor is load-bearing: sampleWeight = 0 IS NOT A SKIP in this
  optimiser (Adam momentum and weight decay still apply, t and
  m_batchCount still advance), so weights must never decay toward zero.
- Cached as the RAW |r| under the label's own validity flag, normalised at
  use: the mean keeps moving as bars resolve, so a cached weight would be
  stale for every bar but the last.
- Peer pool rows stay at 1.0 - the pool record carries no forward return
  and deliberately does not name its source bar. Normalisation is what
  keeps the local:peer gradient ratio unchanged.
- Not lookahead: measured from future bars exactly like the label, reaches
  the loss only, and has no route into the feature vector.
- RETRAIN-FORCING by design. A weighted and an unweighted model have
  identical topology and identical weight-file shape and differ only in
  what they were taught to value; the fingerprint is the only thing that
  can tell them apart.
- Reports the realised weight distribution once per cache, including the
  share pinned at each clamp - a saturated scheme is invisible in every
  downstream number.

Compiled 0 errors, 0 warnings.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-09-01 22:32:12 -04:00
AnimateDread
724f8dce83 feat(features): fleet keep mask - 13 of 49 columns, the false-pivot lever
RETRAIN-FORCING. Re-keys every fingerprint (column count 49 -> 13, plus
an explicit |FMASK: token).

Decomposing the deployed fleet's payoff by distance-to-pivot says the
whole edge is one number. A hit pays +0.94 to +1.57 ATR and is FLAT
across d=-3..0; a miss costs -0.76 to -1.62; the miss rate is 34.5% to
49.4%. Remove the misses and every chart earns +0.97 to +1.22 instead of
+0.05 to +0.32. One percentage point off that rate is worth +0.011 to
+0.018 ATR per call.

Two obvious levers are already spent, and the logs say so. Raising the
vote threshold buys precision and hands the expectancy back - SP500's
precision climbs 38.9% -> 57.3% across the six rungs while its book
expectancy FALLS +0.39 -> +0.21 - and the rung search already takes the
best rung available. Exits do not rescue it either: the 64-cell sweep
fails Sidak on all six, p 0.19 to 1.00. Narrowing the label to drop d=+1
was checked and rejected: it is negative on SP500/USDCAD/XTIUSD, neutral
on XAUUSD/USDJPY and strongly POSITIVE on EURUSD (+0.769, n=116).

What is left is discrimination at fixed coverage, and 294 inputs against
~4,700 independent observations is 16 per first-layer weight - the
overfit regime, where an out-of-sample overfit IS a false pivot.

The keep-screen already measured which columns carry information, and
the open question blocking a prune was whether the masks agree. They do.
Intersected across all 24 members - six charts x four architectures -
swing[5..8] and ma[0..1] are UNANIMOUS, ma[2..4] carry 23/23/20 votes,
and 16 of 49 columns are kept by nobody. This mask is every column kept
by at least half the fleet: swing[5..8], ma[0..4], crossasset[3,4],
volume[1,3]. Width 294 -> 78, first-layer budget 16 -> ~60.

Two of the votes are worth reading twice. swing[0..4] - the five
CONFIRMED-PIVOT features - scored 4, 2, 0, 2 and 1 of 24: the label is a
ZigZag pivot event and the ZigZag geometry carries almost nothing, while
trend position carries everything. And alt[0..11] scored 4, 1, 1, 1 with
atr at zero, so a whole rate-limited external pipeline buys nothing
under this label. Alt stays ENABLED and merely unmasked, so the screen
keeps reporting and the finding stays falsifiable.

Implementation keeps one authority for block order and width:
CFeatureBuilder::FeatureBlockTable, which FeatureSlotName, the kept
count and the emit-time compaction all read. The mask is stated
BLOCK-RELATIVE so toggling a block cannot silently re-point it, and
applied once per bar at the existing sanitize seam rather than inside
nine emitting blocks. m_neuronsCount is set from the counted kept
columns, never from a hand-written subtraction, and a one-shot width
check fires if the two ever disagree.

FEATURE_MASK_VERSION 0 restores the full set exactly, including the
pre-mask fingerprint. This is a measurement, not a conclusion: the
screen tests each column's MARGINAL information and cannot see a column
that is useless alone and useful in combination, so judge the retrain
against the per-chart false-pivot rate and book expectancy above.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-01 11:12:39 -04:00
AnimateDread
2c2cdfed6a feat(exits): sweep the barrier grid on held-out rows; remove SL_Mode/TP_Mode
PHASE 1 of doing the trade-management search inside the EA instead of in the
MT5 optimizer, for the one pair that cannot wait: the stop and the target.

WHY NOT ALGLIB. The space is discrete - 8 stop widths x 8 target widths = 64
cells. At that size you do not search, you ENUMERATE. Exhaustive has no seed,
no convergence question and no tuning of its own, and it answers the thing a
search cannot: whether the good region is a broad plateau or one lucky cell.

WHERE IT RUNS. Inside the era report's existing OOS threshold sweep - same
held-out rows, same purge, same certified rung the deploy gate uses. No tester,
no agents, no .set.

ORDERING IS THE WHOLE POINT. MFE and MAE cannot say which barrier a trade hit
FIRST, and "both were touched" is the common case, so a grid evaluated from the
extremes alone would be guesswork dressed as measurement. MeasureBarPayoff()
already walks the forward window bar by bar; it now records the first-touch bar
offset for each grid level in each direction. One compare per level per bar on
a loop that already runs. The sweep then costs 64 integer compares per fired
row and reads no prices at all.

Same-bar ties resolve to the STOP. Bar data cannot order two touches inside one
bar, and assuming the target would be the optimistic half of an unknowable coin
flip, on exactly the half that flatters the result.

JUDGED AGAINST THE NULL OF THE MAXIMUM, NOT ZERO. Picking the best of 64 cells
and reporting its own z is the best-of-N error this project has already made
four times, including on the deploy decision. SidakFamilyP over the cells
actually scored is the same correction the rung sweep and the baseline
comparator use. A cell needs BARRIER_MIN_CALLS before it may win at all - a
thin cell tops the grid on noise alone. The pair is STORED either way: a reader
must be able to tell "swept and rejected" from "never swept", so the p travels
with the pair (WSTA in .stats) and gates its use at read time, not its record.

SL_Mode and TP_Mode inputs are REMOVED, and STOP_LOSS_MODE/TAKE_PROFIT_MODE
with them - deleted rather than left dangling, per the RISK_REWARD_RATIO rule:
a live enum with no input behind it is the shape of the 2026-07 incident where
a saved .set kept feeding a deleted ordinal back in. With the inputs gone there
is no ordinal left to feed, and ValidateTradeManagementInputs() loses two
members.

Three tiers at read time, most trusted first: the swept pair when it cleared
its own family-wise test; else the mean excursions (cruder, but nothing was
SELECTED to produce them, so they need no such test); else a fixed fallback
that announces itself and is unreachable on a deployed chart, since a chart
with no completed era cannot pass the deploy gate.

Entry offset, expiration, trailing and the exit-vote flag stay inputs for the
MT5 GA - they need entry-fill and path-stepping simulation this phase does not
have.

Compiled 0 errors, 0 warnings.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-31 14:28:34 -04:00
AnimateDread
f942cb2ef2 feat(exits): size the stop and target from measured excursions; gate indicators on a new bar
Two changes, both driven by the same single-backtest profile (6,607,688 ticks,
165 s, SP500 H4 2019-2026).

INDICATORS, 39% OF THE PASS. m_indicators.Refresh() cost 61.4 s at 9.30 us per
tick, running on every quote regardless of Expert_EveryTick. Every indicator
this EA registers is computed from CLOSED bars - ATR, the MA, the ZigZag, the
feature block - so their value cannot change between two ticks of the same bar
and the refresh was recomputing a constant. It is now gated on a new bar, with
its own CNewBar watermark: IsNewBar() consumes the transition and Refresh()
runs before Processing(), so sharing m_newBar would have silently disabled the
SetDirection gate.

The exit invariant holds. What changes intrabar is price and position, and
neither comes from an indicator buffer: m_symbol.RefreshRates() still runs on
every tick and is what CheckClose/CheckTrailingStop/pending maintenance read. A
trailing stop still moves on any tick; it now compares the live price against
an ATR from the last closed bar, which is the ATR it should have been using.
ProtectOpenPosition() keeps its ungated copy - it runs only on ticks Refresh()
declined, and never appeared in the profile.

SL_MEASURED / TP_MEASURED (both value 100), now the shipped defaults. The fixed
pair was SL x2 / TP x6 against a measured SP500 MFE of 2.59 ATR and MAE of
2.71: the stop sat INSIDE the average adverse move and the target BEYOND the
average favourable one, so the average trade was stopped out before reaching a
target it does not reach. That converts winners into losers mechanically at any
precision, and no model work can fix a barrier pair pointing the wrong way.

These are NOT the removed SL_INTELLIGENT/TP_INTELLIGENT. Those scaled the
barriers by the model's own CONFIDENCE - an over-confident model gave itself a
tighter stop and a wider target, which is why they were deleted. These read a
MEASUREMENT of what the market did on the bars this ensemble fired on: the mean
adverse and favourable excursions in ATR at the hold horizon, already computed
by the era report that certifies the deploy and previously printed and thrown
away. Only the WIDTH is measured; the reward:risk that falls out is reported,
never targeted - the ratio is policy, the width is what pays.

A FRESH ENUM VALUE, never the vacated -1 the removed members held: MetaTrader
does not validate enum inputs, so a .set saved by that build still feeds -1 in,
and reusing it would silently give a stale file a new meaning.
ValidateTradeManagementInputs() keeps rejecting -1 and now accepts 100.

Plumbing follows the derived-threshold route exactly, because
ExpertSignalCustom.mqh is the PARENT of the filter that owns g_ensBest* and
cannot read them: stashed inside the isBetter block (so the widths describe the
bars the CHECKPOINT fired on, never a later era's), persisted as WST9 in
.stats (a deployed ensemble runs no further eras - the tier-ladder failure one
layer along), and published per tick via PublishMeasuredBarriers(). The -1/-1
"not measured" state is published too, so a reset-weights cannot leave a stop
sized off a dead ensemble.

With no measurement it falls back to the shipped fixed presets and says so once
per run. It deliberately does not substitute a plausible number: an invented
width would be indistinguishable from a measured one in every log afterwards.

Compiled 0 errors, 0 warnings.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-31 13:31:30 -04:00
AnimateDread
04cc345661 perf(tester): split the per-tick profile so the slow bucket names itself
Three turns of reasoning about where a 6.3-minute pass goes have produced
three hypotheses and no measurement. The profiler that would answer it
already existed and printed nothing: across 129 optimization passes on 12
agents the "tester pass profile" line appeared ZERO times, while OnInit's
output appeared on every one.

Two fixes, both aimed at ending the guessing rather than at being right.

1. The line is now built by WarriorTesterProfileLine() and printed from
   OnTester() as well as OnDeinit(). OnTester runs on the agent at the end of
   the pass, BEFORE OnDeinit. If the line appears there and not in OnDeinit,
   an optimization agent is discarding OnDeinit's Print; if it appears in
   neither, g_tpTicks is genuinely 0 and the instrumentation never ran. Those
   need different fixes. The zero-tick case now prints its own explicit line
   instead of staying silent, because a profiler that says nothing when it
   fails is indistinguishable from a fast pass.

2. "Expert.OnTick" was one bucket containing both halves of the question.
   CExpertCustom::Refresh() runs on EVERY tick whatever Expert_EveryTick says
   - correctly, since an open position must be manageable on any quote - and
   under a 1-minute-OHLC model that body executes millions of times per pass.
   Its three steps are now timed separately: TCHasEnoughHistory(),
   RefreshRates(), and m_indicators.Refresh(). They are reported as a SUBSET
   of Expert.OnTick, not as siblings, because double-counting a bucket is how
   a profile lies. ProtectOpenPosition()'s copy of the same refresh counts
   into the same bucket rather than hiding on the declined-tick path.

The accumulators move to System\TesterProfile.mqh: ExpertCustom.mqh has to
see them and is included long before Warrior_EA.mq5's own globals. Every
bracket is guarded on g_tpActive, which OnInit sets only for
MQL_TESTER/OPTIMIZATION/FORWARD - a live chart never reads the clock for it.

No behavioural change to any path. Compiled 0 errors, 0 warnings.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-31 12:07:09 -04:00
AnimateDread
a9c21fdf10 feat(tester): gate input RANGES at OnTesterInit, not just input values
ValidateTradeManagementInputs() hard-gates input VALUES at OnInit because
MetaTrader replays a saved .set without validating it. Nothing gated input
RANGES, so an optimization would happily sweep a parameter that cannot affect
a trade - or one that destroys the run - and score every pass as if it meant
something.

Measured on the 2026-08-31 session: Signal_ThresholdOpen was swept over
10..80, which MetaTrader expands to twelve PERCENTAGE_PRESETS members. Its own
label reads "SEED (derived after era 1)": every pass restores the pinned rung
from .stats into g_ensDerivedThreshold, and OnTick publishes it over
m_threshold_open before the first bar closes. Twelve identical strategies, a
12x multiplier on a fifteen-hour estimate, and no warning anywhere. A wasted
dimension does not only multiply runtime - it fills the GA's fitness landscape
with plateaus, so crossover produces clones. 115 of 129 passes scored 0.

WarriorGuardOptimizationRanges() runs once in the controlling terminal, the
only place ParameterGetRange/ParameterSetRange are legal and the last moment a
bad sweep is free to stop. Two classes:

  REFUSE - keys BuildModelFingerprint(), or the capacity budget upstream of the
  neuron count inside it (Use_Training_Pool). These select a different MODEL,
  not a different strategy: the agent resolves a .nnw filename that does not
  exist and either votes silently or trains inside the backtest, and the pass
  still lands in the results table looking real. Session stopped with
  INIT_PARAMETERS_INCORRECT.

  PIN - a seed, or inert on an inference-only pass: the training-only inputs,
  the DB ranking pair, presentation flags, and the news filter (calendar error
  4806 over historical dates, so it fails open on every bar - the message says
  so, because a config optimised here trades through news live). The sweep is
  switched off and the operator's own value kept.

tradingdirection is deliberately NOT in the table - it is a legitimate
override, and WarriorEffectiveDirection() is explicit that the screen never
overrules an operator who chose. A degenerate range spanning BOTH and the
screened side is reported instead, from the controlling terminal where the
drift screen has full history.

An unresolved name warns rather than refuses: the table addresses inputs by
string, so a rename unbinds it, but a false refusal would make the EA
impossible to optimize at all. The summary line prints on every optimization,
clean ones included, so a table that has come unstuck shows up on the first
run - same doctrine as the drift screen reporting on charts it does not
restrict.

Compiled 0 errors, 0 warnings. Live paths untouched: nothing outside
OnTesterInit() is reachable from this file.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-31 11:52:42 -04:00
AnimateDread
f7adb0bf16 feat(panel): Pause works on a deployed chart - it holds ONLINE LEARNING, and owes the bars back
Operator asked for a queue: hold online-training signals while an optimization
runs, then push them "as if the stop never happened". The queue already exists
and is better than a queue - it is a WATERMARK. m_onlineLearnedUpToTime
advances one LEARNED bar at a time inside the catch-up walk and is never
bulk-set at the end, so a step that returns early leaves it exactly where it
was; the next walk starts at the frontier, runs back to the watermark and
learns oldest -> newest in strict chronological order. It rides in .stats, so
it survives a restart. Nothing to serialize, nothing to drain.

What was missing was the switch. OnlineLearnStep() already gates on
m_trainingPaused - but HandleCpTogglePause() REFUSED on a deployed chart ("there
is no training run to pause"), which was true of the training loop and wrong
about the chart. Online continual learning runs only on deployed charts, so the
one state where pausing matters was the one state the button would not enter.

WHY IT MATTERS FOR THE GA: online learning re-saves the deployed .nnw, and every
tester agent re-seeds its optcache when the production file is newer than its
copy. A model that keeps learning mid-run makes later passes evaluate different
weights than earlier ones and the GA reads that as a parameter effect - and the
config you optimised was tuned against a model that no longer exists by the time
you deploy it. Pausing freezes the weights for the run and gives the bars back
afterwards.

Bounded by ONLINE_LEARN_MAX_CATCHUP (64 confirmed bars, ~10 days on H4) - the
walk stops there and older bars stay behind the watermark. Covers an
optimization run; not a way to park a chart for a month. Stated in the alert.

Trading is unaffected: nothing in the live signal path reads m_trainingPaused;
the only other deployed-state reader is the one-shot ladder/snapshot rebuild,
which defers. The alert names which thing was paused, because "training paused"
on a deployed chart reads as pausing something that was not running.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 11:55:15 -04:00
AnimateDread
eed262c416 fix(chart): the vote-arrow file stamped the threshold from INIT, not the one in force
An arrow is a claim about a threshold, so the store discards its cache when
that threshold changes between sessions. Correct rule, stale input: the stamp
was captured once in Configure() during OnInit and never refreshed, while the
derived rung goes on MOVING for the rest of the session (it follows each era's
rung until a checkpoint pins it).

So the two files written in the same OnDeinit disagreed by construction:
.stats saved the CURRENT threshold, .votearrows saved the init-time one. Next
start compared them and threw away a perfectly good arrow set.

Measured 2026-08-27 22:26 - four of six charts lost everything:
  XAUUSD stored 25 / now 20   XTIUSD stored 25 / now 5
  SP500  stored 25 / now 5    USDJPY stored 15 / now 20
The two that survived, EURUSD (173 arrows) and USDCAD (13), were the two
whose rung happened not to move.

WarriorVoteArrowThreshold() is now the one expression, read live at the load
comparison AND at every save; the cached member is gone, so it cannot go stale
again. Only the close threshold is still passed to Configure(), because
VOTE_EXIT_DISABLED_THRESHOLD genuinely cannot move.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 22:37:49 -04:00
AnimateDread
5304cfb876 refactor: delete the machinery I added that was not earning its place
Self-audit against YAGNI. Three things from this session's commits were
configuration nobody asked for, and one was a measurement report pasted into
source.

DELETED - SignalCooldownScopeOverride + WarriorSignalCooldownScope(). I added a
source override for the cooldown scope, then concluded in the very next commit
that source-overrides-operator is the wrong pattern and stood it down to -1.
What was left was a mechanism whose only state is "disabled" - the definition of
speculative generality. The scope is read straight from its input now. The thing
that actually solves "which value is running" is the init line that PRINTS the
resolved value, not a second place to set it.

DELETED - REQUIRE_BOTH_SIDES_PAY. A #define that was always true, added so the
new gate could be "reverted". Nobody asked for that switch and a gate condition
that is optional is not a gate condition. The rule is either right or it is not;
if it turns out wrong, git is the revert mechanism.

TRIMMED - 38 lines of comment in the ensemble verdict listing all six charts'
by-side payoff numbers. Those are a MEASUREMENT: they belong in the commit that
made the change and in the session record, not pinned in source where they go
stale on the next run and start actively misinforming. Two lines left saying
what the code does and when it can be false.

TRIMMED - the cooldown override comment from 11 lines to 4. Same reasoning.

KEPT, with the case for each:
  CSignalDeclusterPolicy   replaced FOUR copies of one rule that had silently
                           diverged into two different units and two different
                           windows. Net negative lines.
  PivotLabelFillTarget     replaced FIVE hand-rolled target vectors and is the
                           only place the sum-to-1.0 gradient invariant lives.
  PivotLabel* window fns   the -1 sentinel collision they fix was a live defect.
  CRunningMean             two meanings were sharing one accumulator, which is
                           how the purge came to be sized off the wrong one.
  LABEL_RESOLUTION_CAP_BARS  a bound on a MEASURED quantity feeding the training
                           split. The floor is load-bearing; the cap stops a
                           degenerate ZigZag eating the training set.
  Test_PivotLabelWindow    161 of the 195 net added lines. Two of the three
                           things it pins were live bugs this session.

Net across the whole session, production code only, excluding tests:
+479 / -390 = +89 lines, and in exchange: decluster rule 4 copies -> 1,
target builders 5 -> 1, chart reconciliation 2 -> 1, twelve loose NMS members
-> two policy objects.

Compiles clean, 0 errors 0 warnings.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 21:02:04 -04:00
AnimateDread
eac8912dd0 fix(gate,cooldown): both sides must PAY, not just fire; stand down the source overrides
Three fixes off the 2026-08-27 session logs (six H4 charts, fresh from era 0).

1. THE DEPLOY GATE NEVER READ ITS OWN DRIFT-FREE TEST.

   twoSided was `firedLong > 0 && firedShort > 0` - an anti-degeneracy check.
   The by-side PAYOFF prints three lines below it under the heading "THE TEST
   THAT IS DRIFT-FREE", and nothing consumed it. Measured at each chart's own
   derived rung, hold horizon, ATR per call:

     EURUSD  long +0.073  short +0.115   <- the only one where both pay
     USDJPY  long +0.358  short -0.136
     USDCAD  long +0.177  short -0.070
     XAUUSD  long +0.076  short -0.379
     XTIUSD  long +0.222  short -0.065
     SP500   long +0.657  short -0.273

   Five of six were stamped DEPLOYABLE on 30-37% precision (chance 13-14%)
   while their short book lost money on every call. XAUUSD's whole vote earned
   -0.157 against its own always-long book at +0.436 and still passed.

   Precision is measured against the LABEL: it says a turn was called, not that
   the leg after it paid. On a trending instrument the down-legs are shorter, so
   a symmetric caller with real skill still bleeds on one side - and the higher
   rungs make it worse, because more member agreement means more agreement WITH
   the drift (SP500 at the 15% rung fires short n=0).

   tradeableOK now ANDs in bothSidesPay. A floor at zero, not a significance
   test - no per-side variance is tracked, so "positive" is the honest claim.
   A side that never fired is not judged; an era where neither side has a
   forward window is not judged either. REQUIRE_BOTH_SIDES_PAY restores the old
   behaviour. The verdict line now NAMES this when it withholds DEPLOYABLE,
   because "not good enough yet" and "one side of the book loses money" are
   different problems with different fixes.

   Timing is deliberate: ENSEMBLE_CHECKPOINT_MIN_ERA is 20 and the fastest chart
   is at era 17, so nothing has deployed yet and this costs nothing to land now.

2. THE SOURCE OVERRIDE WAS DISCARDING THE OPERATOR'S COOLDOWN.

   SignalCooldownOverrideBars was 30. The charts are configured for 10. The
   session log says "signal cooldown (30 bars)" 43 times across all six. That is
   the same failure the override exists to fix, pointed the other way - source
   defeating the operator instead of a stale profile defeating source - and it
   is WORSE, because a stale profile is at least visible in the chart's own
   inputs dialog and this is not. Stood down to 0, which is the exit condition
   its own comment always described.

   The scope override (added earlier this session) is stood down to -1 for
   exactly the same reason rather than kept out of convenience.

   The real fix for "which value is running" is not a second place to set it:
   ConfigureAISignal now prints the RESOLVED window and scope once per chart
   alongside the chart input and both overrides, and says so explicitly when
   they disagree. One line kills the whole class.

3. The overlay sweep's hard-coded OVERLAY_NMS_WINDOW of 6 is gone with the
   window/decluster work replayed onto main - it now uses the resolved cooldown
   like every other layer.

MEASURED, and it corrects a number I reported earlier in the session: against
the 30 bars the override was forcing, 89.8% of consecutive vote-arrow pairs sat
inside the window. Against the 10 bars actually configured, the figure is 11.7%
(1038 pairs across six .votearrows sidecars, gaps counted in real H4 bars with
weekends excluded). The earlier figure was true of the window that was running
and NOT of the one the operator set; the 10-bar number is the one to judge the
fix against. USDJPY carries most of what remains (18.4%, and all 21 fleet pairs
below 7 bars) - it is also the least-trained chart, at era 3.

NOT CHANGED, and flagged rather than fixed: the source default is still SCB_30
and its comment argues 20 is the smallest defensible value, because a trade on
this label is held 5 + the median leg = 18-19 bars. A 10-bar cooldown re-announces
inside that hold. That is the operator's call, not a bug.

Compiles clean: EA and both test suites, 0 errors 0 warnings.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 20:02:47 -04:00
AnimateDread
184de331fc fix(label,signal): move the pivot window into the leg; one declustering rule, measured in bars
Two independent defects, both about a rule written down more than once.

THE LABEL WINDOW SAT ENTIRELY ON THE APPROACH. It labelled d=1..5 bars BEFORE
the turn. Measured payoff by distance, fleet-pooled over 102 era rows and 3
charts at the hold horizon: d=1 +2.095, d=2 +1.743, d=3 +1.300, d=4 +0.969,
d=5 +0.769 ATR/call. Slope -0.343 per bar, monotone on all three individually.
Pure hold truncation predicts d=5 earns 14/18 = 0.78x of d=1 (~1.63); it earns
0.769, which is 0.37x - so truncation explains less than half and the rest is
the ADVERSE APPROACH, the call sitting through bars of price moving against it.
Every bar of lead the window granted was paying that.

The window now straddles the turn: d=+1..-3 (LEAD 1, LAG 3). It keeps the
best-paying approach bar and spends the rest of its width on the far side,
where direction is confirmed. WIDTH IS UNCHANGED AT 5, deliberately - it sets
the class balance, and ApplyLogitAdjustment's tau is capped at 1.2 logits with
the cap binding at EVERY imbalance, so only the uncorrected residual moves
(1.9x at width 5). Simulated over three leg-length regimes the balance shifts
by <=0.2pp, so nothing downstream of it had to move. That is also why the
target stays FLAT across the window rather than Gaussian-tapered: a taper
shrinks directional mass ~40% and tau has no budget left to absorb it.

Finality is re-derived, not weakened. One committed pivot strictly newer than
the window's newest slot freezes the whole window - the labelling pivot's
finality and the negative verdict together - because ZigZag can only ever move
its newest pivot, and only toward newer bars. Verified: the witness is the next
pivot newer than the labelling pivot in 100% of simulated cases, so minMove
still measures the leg AFTER the turn.

TGT:PVT1 -> TGT:PVT2:lead:lag. Every .nnw re-keys; the fleet retrains from era 0.

-1 IS NO LONGER A VALID SENTINEL - it is one bar into the leg. Added
PIVOT_LABEL_NO_PIVOT and PivotLabelDistanceSlot(); the by-distance payoff report
bucketed the no-pivot population and the leg bars into the same slot otherwise,
and its "directional labels only" test (`d >= 0`) would have dropped every leg
bar it now exists to measure.

THE DECLUSTERING RULE WAS WRITTEN OUT FOUR TIMES and the copies disagreed.
Rules 0-3 lived in NmsLiveAccept, PruneDirectionalClusters, pass 3's OOS tally
and the overlay sweep, each carrying a comment insisting it must match the
others. It did not:

  * the two sweeps compared bar INDICES; the live gates compared wall-clock
    seconds. Elapsed time is always >= bars*period because weekends and session
    breaks add time without adding bars, so the live rule was strictly the most
    permissive and could only ever UNDER-suppress. On H1 a 30-bar window is 30
    hours against a ~50 hour FX weekend: a Friday signal never blocked a Monday
    one, once a week, per chart. On session-break instruments, daily.
  * the overlay sweep had rules 1 and 2 only - no cooldown, no alternation -
    against a hard-coded OVERLAY_NMS_WINDOW of 6 while the live gate required
    30, and its comment claimed parity with m_signalClusterWindow. It drew
    reconstructed arrows 7 bars apart onto a chart whose gate requires 30.

Now one CSignalDeclusterPolicy, measured in bars via WarriorBarsBetween(), with
an unresolvable frame SUPPRESSING rather than passing. The vote layer configures
it as a pure cooldown because that is what CheckOpenPosition actually applies.

Three more that let clusters through:
  * THE CLOCK WAS NEVER SEEDED. MT5 restores arrows from the .chr profile but
    not the state that spaced them, so every restart, recompile or timeframe
    change let the next bar fire regardless. Seeded from the newest arrow on the
    chart once the progressive restore completes.
  * THE COOLDOWN WAS CONSUMED BEFORE THE TRADE EXISTED. VoteCooldownAccept both
    tested and committed, ahead of order-parameter validation - and the failure
    branch restores the vote for retry but could not un-burn 30 bars. Split into
    a pure test plus VoteCooldownCommit() at the draw.
  * SCOPE HAD NO SOURCE OVERRIDE. Bars did; scope did not, so a chart whose
    profile held PER_DIRECTION silently dropped rule 0 with no way to correct it.

Also DRY: five hand-rolled 3-class target vectors -> PivotLabelFillTarget(), the
one place the slot order, the smoothing constants and the sum-to-1.0 invariant
live (the gradient is target_i - softmax_i, so a vector that does not sum to 1
adds a constant drift to all three logits). Two copies of the chart-wide
reconciliation -> WarriorReconcileVoteCooldown(). Twelve loose NMS members ->
two policy objects.

Documented at CNet::backProp: sampleWeight = 0 IS NOT A SKIP. Adam still applies
decayed momentum and decoupled weight decay, t still advances, and m_batchCount++
is unconditional - so masking a class-skewed subset that way shrinks every weight
rather than ignoring the example. Drop the bar before queueing instead.

Compiles clean (0 errors, 0 warnings). Window arithmetic verified by simulation;
the label's in-situ behaviour against real ZigZag output is NOT verified here.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 19:57:50 -04:00
AnimateDread
213b3aac15 fix(chart): the reconciliation never ran - it was hooked to a sweep deployed charts do not do
cooldown-recon put the chart-wide cooldown at the end of the overlay sweep. It
executed ZERO times. This store's own header already said why: the sweep
're-arms only when an era ends. A DEPLOYED ensemble runs no further eras'. Five
of six charts were deployed, so there were zero 'Filtered view: swept' lines in
the entire session while the saved files still held 148 same-side pairs under 30
bars on XTIUSD.

Moved to the completion of the progressive vote-arrow restore, which runs on
every chart including deployed ones.

The restore thinning alone was never going to be enough either: MT5 persists
chart objects in profiles\Charts\*\chart*.chr, so arrows drawn under an older
window are ALREADY on the chart when the process starts, and a freshly-thinned
restore just adds to them. Two correctly-thinned sets still union into clusters.
The chart is the only authority.

Same construction as before: OBJ_TREND only (the line is the canonical half of a
mark, matching Snapshot()), sorted by time first because object order is not time
order, and the gap>0 guard so a mis-ordered set fails visibly by keeping rather
than silently by deleting.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 16:16:40 -04:00
AnimateDread
4e2cdd3aff fix(chart): reconcile the cooldown over the CHART, not over one producer's record
The record-based prune was individually correct and still left clusters. It is
not the only producer of a vote arrow: the persisted-arrow RESTORE thins its own
list from its own state, the overlay sweep thins its own list from its own state,
and the live gate marks the current bar from a third. Each spaces ITS OWN
survivors 30 bars apart; interleaved on one chart the union sits 1 bar apart.

Two independent thinning passes over one namespace produce a union, not an
intersection.

MEASURED from the saved .votearrows files, which is what the chart actually
holds:

  XTIUSD  284 arrows, 180 gaps under 30 bars, 148 of them SAME-SIDE, min gap 0
  XAUUSD  263 arrows, 177 gaps under 30 bars, 144 same-side, min gap 1
  EURUSD  309 arrows, 176 gaps under 30 bars, 110 same-side, min gap 0

while every producer's own log reported it had thinned correctly. The sweep's
'drew' counter says what ONE producer drew; the chart is the union. Verifying on
that counter is what let this stand through four builds.

The authority is now the chart itself: after the sweep, walk every
SIG_VOTE_PREFIX OBJ_TREND object, sort by time, enforce one window. Whatever drew
an arrow, this runs last.

OBJ_TREND only - a mark is a line AND an arrow and the line is canonical, the
same test Snapshot() uses. Sorted first, because object order is not time order
and an unsorted forward walk yields negative gaps, which is how a prune once
deleted 272 of 273 arrows.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 16:12:49 -04:00
AnimateDread
844aac653a fix(train): the OOS final pass ran a full epoch at an undecayed rate
Capping the pass at m_etaCeiling was not enough. Measured on the first two live
runs: USDCAD 0.00085 over 13,335 bars, EURUSD 0.00242 over 15,041 - a 3x spread
across charts, because a chart whose plateau ladder reset recently still carries
a high eta and the cap never bound.

The slice turns out to be roughly HALF the data, not a tail, so one pass over it
at the model's own rate is a full training epoch on a model that has already been
selected and certified. That is materially more than the 'just a bit finer
weights' this was asked for.

OOS_FINAL_PASS_ETA_SCALE (0.25) now scales the rate. Scaling rather than
shortening the pass keeps the whole slice in play - seeing the held-out bars at
all is the point - while making the step proportionate to an already-selected
model.

USDCAD and EURUSD have already taken the unscaled pass; that is not reversible
without a retrain. USDJPY has not converged yet and will get the corrected one.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 13:46:48 -04:00
AnimateDread
c5b9a1ad10 fix(signal): a changed input default cannot reach an already-attached EA
Raising Signal_CooldownBars from 10 to 30 changed nothing. All six live charts
kept reporting a 10-bar window, because MT5 stores an input PER CHART in
profiles\Charts\*\chart*.chr and an already-attached EA ignores a changed
default entirely. This codebase already documents that trap, in the derived-
threshold comment in Training.mqh - and converting SignalClusterWindow from a
const to an input reintroduced the exact problem the const existed to avoid.

SignalCooldownOverrideBars (const, 30) now wins over the input; 0 hands control
back to the panel. Tunability per chart is kept, source-correctability is back.

Not applied when the input says OFF: an operator who switched the cooldown off
meant it, and silently re-enabling it from source would be the same surprise
pointed the other way.

ALSO gates the per-model arrow restore on DrawUnfilteredSignals. DrawObject()
returns early when the raw view is off, but AdvanceChartSignalRestore called
WarriorPlotSignalLevel DIRECTLY and never checked - so every restart repainted up
to MAX_PERSISTED_ARROWS per-model opinions per member, four members per chart, on
top of the combined-vote arrows. Same shape as the vote-arrow restore bug in
322c052: a restore path that does not obey the rule its own draw path does.
Stale arrows already on the chart are purged too, since MT5 persists objects in
the profile and nothing else would ever remove them.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 12:04:03 -04:00
AnimateDread
09af7d5ee9 feat(signal): default the cooldown to 30 bars - 10 thinned almost nothing
Measured on the live fleet at 10 bars: the restore thinned 100 of 364 and the
overlay prune 23-66 per chart, leaving 182-250 arrows over ~5000 bars. Every
layer was working; the window was simply below the ~20-bar natural spacing
between vote arrows, so it could only catch the tightest pairs.

The floor is principled, not cosmetic. A trade on this label is held for
5 + the median ZigZag leg = 18-19 bars, so any second signal inside that window
is the same trade being re-announced. 20 is the smallest defensible value and 30
is one step above it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 11:59:17 -04:00
AnimateDread
3218db4a38 feat(train): ONE pass over the held-out slice at deploy, on the restored checkpoint
The OOS slice is the newest history and the model never trains on it, while
online learning adapts to every bar resolving AFTER deployment. That leaves a gap
exactly at the handover, over the most regime-relevant data there is. This closes
it: select on validation, then refit on everything, which is standard practice.

Placed AFTER Net.RestoreWeights() and ResetOptimizerState() and BEFORE
PersistDeployedModel(), so it refines the weights that were actually SELECTED
rather than whatever the run happened to end on, and what it produces is what
gets written down.

THE COST IS REAL AND IS NOW STATED IN THE LOG. The deploy line promises "every
model reverts to the weights it held at the era whose combined vote scored best,
so the ensemble that trades is exactly the one that was measured". After this
pass that is no longer literally true, so the pass prints that the certified
numbers belong to the PRE-PASS weights and must be quoted that way. Set
EnableOosFinalPass=false to keep certified == traded exactly.

Guards:

* ONE-SHOT PER RUN, and the flag is set BEFORE the loop so no early return inside
  it can leave the pass eligible to fire twice over bars it already trained on.
  Reset at m_trainRunActive=true, because a retrain is a fresh selection and
  earns a fresh pass.
* THE CONVERGED RATE, never a plateau-boosted one: m_modelEta can still carry
  PLATEAU_RESTART_BOOST from an escape attempt, and this is a refinement of a
  selected model, not another warm restart. g_eta is what backProp reads, so that
  is what is capped and restored.
* OLDEST -> NEWEST. Series indices count backwards, so decreasing i moves forward
  in time - the order the bars happened in.
* A failed feedForward is never followed by backProp; the output layer would
  still hold the previous sample's activations and the update would be this bar's
  label against another bar's prediction.
* m_oosFinalPassCutoff records the newest bar consumed and is deliberately NOT
  cleared on a new run, so a later run can say plainly that its out-of-sample
  window reaches back into bars this model has already seen.

Expect the gain to come from CURRENCY rather than finer weights: OOS precision
was measured flat from era 20 while in-sample error kept falling, so the data
this model can already see is exhausted.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 11:05:43 -04:00
AnimateDread
322c052a65 fix(chart): the persisted vote-arrow restore put back every arrow the cooldown removed
Clutter remained after cooldown-v3 because there is a FOURTH producer of
SIG_VOTE_PREFIX arrows: CVoteArrowStore, which replays a .votearrows file and
draws into the SAME object names as the overlay. Replaying a file written before
the cooldown existed therefore resurrects exactly the arrows the prune deleted.

On a DEPLOYED chart that is the entire arrow set. The store's own header says
why: the overlay re-sweep "re-arms only when an era ends. A DEPLOYED ensemble
runs no further eras" - which is the reason this store exists at all, and also
the reason nothing would ever have removed those arrows again.

The restore now thins to the cooldown at Load(), before the progressive draw is
armed, so it is idempotent: a file already written from a cooled chart passes
through untouched, an older one is corrected once.

IT SORTS BY TIME FIRST, AND THAT IS NOT OPTIONAL. Snapshot() walks
ObjectsTotal(), so the record is in OBJECT order - its own comment says so, and
the existing MAX_KEPT trim already sorts a copy for exactly this reason. Applying
a spacing rule to an unsorted record yields negative gaps, and a negative gap is
inside any window: that is the bug that wiped 272 of 273 arrows in 2ca32e9, which
would have been reproduced here verbatim.

Insertion sort on the four parallel arrays - n is capped at VOTE_ARROWS_MAX_KEPT
(1000) and this runs once per chart per session on a path that has just done file
I/O.

Same `gap > 0` guard as the overlay prune, so a future ordering change fails
visibly by KEEPING rather than silently by deleting, and it logs what it thinned.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 10:58:54 -04:00
AnimateDread
2ca32e933f fix(signal): the overlay cooldown prune ran backwards and left ONE arrow per chart
cooldown-v2 suppressed 272 of 273 on SP500, 320 of 321 on EURUSD, 329 of 330 on
USDCAD - one surviving arrow on every chart in the fleet.

The record is OLDEST-FIRST. The prune walked it backwards, so every gap came out
NEGATIVE, and a negative gap is always <= the window: everything after the first
arrow was suppressed.

The direction was taken from the member comment on m_overlayIndex, which reads
"walking newest -> oldest" and is WRONG. The sweep DECREMENTS a SERIES index
(0 = newest) from MathMin(span, barsAvail-150) down to m_overlayStopIndex, so it
walks OLDEST -> NEWEST. The pre-existing overlay NMS at the draw site agrees -
it tests (m_overlayNmsKeptIdx - idx) and expects that to be positive for later
bars. A stale comment counts as a guess, and this one cost a build.

Comment corrected at the declaration so the next reader is not misled the same
way.

Guard added: the gap must be > 0 as well as <= the window. A non-positive gap
means the record is not in the order this loop assumes, and suppressing the whole
chart is precisely what that looks like from the outside - so it now fails
visibly by KEEPING rather than silently by deleting.

Found only because the verification was the drawn arrow count rather than an
assertion that the code was correct. Compiling clean said nothing about it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 10:34:46 -04:00
AnimateDread
110080dc86 fix(signal): the cooldown belonged at the VOTE layer, as a filter - not per member
cooldown-v1 extended NmsLiveAccept, which declusters each MEMBER's own signal.
That is not what the charts show and not what trades. The combined vote in
CExpertSignalCustom had NO spacing rule at all - grep found not one reference to
the cluster window in that file - so four individually-declustered members were
averaged into a vote that could fire on consecutive bars. Measured live: 2,970
voting bars becoming 299-328 vote arrows.

Proof of the diagnosis, from the deployed fleet under cooldown-v1: SP500 273 and
XAUUSD 212 arrows, unchanged from before the change. The member-level rule could
not touch them.

Gated where the vote becomes a trade - CheckOpenPosition, beside the
open-prohibition and open-market-closed checks, tracing as "open-cooldown". That
is the filter chain the request asked for from the start and it is where this
should have gone first.

Suppression there means no order AND no live arrow, honouring the same "no arrow,
no vote, no position" contract the member rule already had.

THE DRAWN HISTORY NEEDED A SECOND PASS, NOT AN INLINE TEST. The overlay sweep
walks NEWEST->OLDEST and is chunked across ticks, so an inline cooldown would
keep the NEWEST bar of a cluster while the live gate keeps the FIRST, and the
drawn set would contradict the traded set - the exact defect the renderer's own
comments warn about. The sweep now records what it drew and prunes it backwards
over that record, which is forward in time.

Direction() is a TRANSACTION that can run more than once on a bar, so the live
accept is cached per bar time. Without that a second call flips the bar's verdict
after it has already journaled one.

One resolver, WarriorSignalCooldownBars(), now serves both layers so they can
never disagree about the window.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 10:29:12 -04:00
AnimateDread
6308a19f27 feat(signal): make the signal cooldown tunable, and add a hard any-direction gate
The declustering the charts needed already existed - NmsLiveAccept, per-direction
run-collapse plus cross-direction resolution plus strict alternation - and it was
already set to 10 bars. It could not be TUNED: SignalClusterWindow was a compile-
time const, so finding the right value needed a rebuild. That is the actual gap.

Now three inputs, as enum dropdowns:
  Signal_CooldownScope    per-direction, or a hard any-direction gate on top
  Signal_CooldownBars     SCB_OFF..SCB_50, default 10
  Signal_CooldownMinutes  SCM_OFF..SCM_1440, overrides bars when set

Minutes resolve against the CHART period and round UP, so a cooldown asked for in
wall-clock is never silently shorter than requested and survives a timeframe
change.

SCB_/SCM_ prefixes are deliberately unique. M15/M30/M60 are ALREADY members of
NF_LOOKBACK_PRESETS, and MQL5 binds a duplicated enum member to the first-declared
enum silently - the obvious names would have compiled straight into the news
filter's values.

THE ANY-DIRECTION GATE IS ADDITIVE, NOT A REPLACEMENT, and the first cut of this
had it backwards. Measured on the live log: the current rules draw 222 arrows over
4999 bars, while a BARE 10-bar cooldown permits up to 454 - because ALTERNATION is
what declutters today, not the window. Swapping the rules out would have roughly
doubled the clutter it was asked to remove. Layered, it can only ever suppress
more. Suppressed bars still advance the per-direction last-SEEN cursors, so a run
straddling the boundary does not restart as if it were fresh.

Applied at all THREE sites that must agree - live inference, OOS pass-3 scoring
and the chart renderer. Their own comments say why: an arrow set that does not
obey the same rule as the traded set shows calls the EA would never take.

Also corrects a stale comment that called this window "display only". It is not:
when it suppresses, the live path zeroes the signal outright - no arrow, no vote,
no position. Training never sees it, so these cost no retrain and are correctly
absent from the fingerprint.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 09:48:28 -04:00
AnimateDread
8083a31754 diag(gate): move the conviction curve to the horizon that has value, and add mean-d per rung
The 5-bar conviction curve cannot answer the question it was built for. The
oracle measures ~0 at 5 bars across three charts (+0.012, -0.054, +0.064), so
PERFECT foresight earns nothing there and no rung can show payoff either. Every
reading it produced was null by construction. It was placed at 5 bars for
statistical power, before the oracle showed what that horizon is worth. Kept as
a control; the hold-horizon curve is the one to read.

Also adds MEAN DISTANCE-TO-PIVOT PER RUNG, which is the high-power form of the
same question. Payoff falls ~0.34 ATR for every bar of distance to the pivot
(fleet-pooled: d=1 +2.095, d=2 +1.743, d=3 +1.300, d=4 +0.969, d=5 +0.769,
wrong calls -0.668). So a rung that selects NEARER pivots is worth more per call
even at unchanged precision - and mean-d is a far tighter statistic than
mean-payoff, because d spans five bars where payoff spans several ATR.

That matters because it can REOPEN a lever I closed. Precision does not rise
with the rung - every 15-vs-10 comparison across six charts sits below 0.71
sigma - so the threshold looked exhausted. But precision is not the only thing a
threshold can select for. If conviction correlates with proximity to the pivot,
raising it buys payoff without buying precision.

Directional labels only: an incorrect call has no pivot and therefore no
distance, and folding those in as zero would read as "this rung picks pivots
that are imminent" when it means the opposite.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 09:12:28 -04:00
AnimateDread
0d9320cc87 diag(gate): the ORACLE - what a perfect caller of this label would earn
The ceiling on the target, and the measurement that decides where the work goes.
Same payoff arithmetic, signed by the LABEL's direction instead of the vote's,
over every directionally-labelled shared bar.

If a model that got EVERY pivot right still earns nothing over the holding
horizon then the target carries no money and no amount of model improvement
reaches any - the label, not the network, is what has to change. If the oracle
earns well the target is sound and the shortfall is the model's. Those are
completely different programmes and nothing so far distinguishes them.

It uses no forecast, so it is not a leak: it is the value of perfect foresight
OF THIS LABEL, reported as a benchmark. Nothing may trade on it.

Accumulated above the voter and direction-policy filters, like the zero-skill
book, because it is a property of the bars and their labels rather than of what
the vote did with them. A bar with no directional label offers a perfect caller
nothing to take and is skipped rather than counted as zero - the benchmark is
"every call it COULD make".

Motivated by the first skill-by-distance row, which already reframes the day:
correct calls earn +0.75 to +1.90 ATR against a spread of 0.005-0.042, and
incorrect ones cost -0.66. That puts break-even precision near 32% against a
measured 33-37% - thin, but on the right side, and utterly unlike the "no
payoff" reading the confounded 5-bar window suggested.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 08:21:59 -04:00
AnimateDread
1a9b56e3b0 diag(label): expose bars-to-pivot - the confound the payoff test was missing
CORRECTION to what the payoff instrument was measuring. The 5-bar horizon looked
like the powered test and it is confounded.

SwingPivotDirectionLabel returns Buy when a swing LOW lands up to
PIVOT_LABEL_TOLERANCE_BARS bars AHEAD, and says the quiet part itself: gating on
where the pivot sits relative to entry "would drop exactly the bars where the
turn has not finished coming to us", and how much adverse move remains before
the turn "is a trade-management question".

So on a CORRECT Buy call price is often still falling for d more bars. A window
shorter than d measures the APPROACH, not the leg, and its negative contribution
is expected on the calls that are RIGHT. The tight null at 5 bars
(-0.012 +/- 0.074) is therefore not evidence of no payoff. Neither horizon is
both clean and powered: 5 bars is powered and confounded, 18-19 is clean and has
an SE of 0.277.

(idx - P1) was computed in the label and thrown away. Now cached beside
m_labelResolveAge under the same validity flag, and bucketed in the era verdict.

DELIBERATELY NOT USED AS A PER-CALL HORIZON, which is the trap sitting right
next to this: d exists only on bars the label found a pivot for, so a horizon
that varied with d would hand correct and incorrect calls different windows and
bias the comparison outright. The horizon stays fixed; d only buckets.

The bucket for "the label called no pivot here" is reported by name rather than
folded in, because it is the control the others are read against. Buckets 1..N
condition on the label, so they describe the MECHANISM, not what a book earns.

Reads: rising with d means the edge is in EARLY calls and the tolerance window
is spending it - fixable by reweighting the loss, not by a new label. Flat means
that hypothesis dies.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 08:16:47 -04:00
AnimateDread
3372b82dfa diag(gate): the conviction curve - does payoff rise with vote magnitude?
The practical question behind "can I just trade the strongest signals" is
whether payoff rises with vote magnitude. The threshold sweep already visits
every rung, so the whole curve costs four arrays and no extra pass.

Reported as the DRIFT-FREE statistic per rung - long plus short, both sign
corrected - with the two halves alongside. The halves alone invite reading a
drift-fed long side as skill, which is exactly the error the zero-skill book
caught at the certified rung: an always-long book earns MORE than the vote on
two of three charts.

Taken at the SHORT horizon, which is the one with the power. Pooled across the
three training charts the certified rung reads -0.012 +/- 0.074 ATR - a tight
null, 95% interval [-0.16, +0.13], with the long/short pattern (+0.030 against
-0.041) being the drift signature exactly. The hold horizon agrees and is 3.7x
noisier, so the answer is not a horizon artifact.

Precision is already known not to rise significantly with the rung (every
15-vs-10 comparison across six charts sits below 0.71 sigma). If payoff rises
anyway that is a surprise worth having; if it does not, the two agree and the
threshold lever is closed on both counts.

Still gates nothing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 08:10:47 -04:00
AnimateDread
feaadd80a2 diag(gate): split the payoff by side at the horizon that can actually resolve it
The by-side test is the one that separates directional skill from drift, but at
the HOLD horizon it cannot answer: payoff overlap is the horizon itself, so an
18-bar window leaves ~65 independent observations per chart and a standard error
of 0.25-0.45 ATR against an effect that would matter at 0.1.

The 5-bar window carries ~3.8x the independent observations and roughly half the
standard error. It buys that power by risking a window that ends before the
pivot has committed - which is exactly why the horizon was widened in 98f485b.

So neither horizon alone is trustworthy and both are now reported. Agreement
between them is the evidence; disagreement localises the problem to the horizon
rather than to the signal.

Measured so far, and the reason this was worth adding: the hold-horizon split
puts every chart inside one standard error - undecided, on all three - while the
zero-skill always-long book earns MORE than the vote on two of three. The raw
positive mean was drift, which is what that book was built to catch.

The drift check itself passes: base@hold / base@short lands at 3.39 and 3.41
against an expected 3.60 and 3.80, so the always-long book scales with time the
way real drift does and the payoff arithmetic is sound.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 07:48:53 -04:00
AnimateDread
ce4f74fe2c diag(ensemble): measure how much the four members actually disagree
The ensemble beats its best single member by +2.2 to +6.8pp on all six charts -
sign-stable across six instruments, so the ensemble is doing real work rather
than diluting. How much MORE is available depends entirely on how decorrelated
the members are: the variance of an m-member average scales as (1+(m-1)r)/m, so
at r=0.8 four models are worth about 1.2 independent ones and at r=0.3 nearly 3.

Nothing measured that, so the obvious next lever - different feature subsets per
member, or a fifth architecture - could not be costed. Both force a full retrain
of 24 models, which is not a price to pay on a guess.

Measured on the SIGNED VOTE, which is what actually gets averaged: not accuracy,
not raw confidence. Two members can agree on direction almost always and still
contribute independently through magnitude.

Accumulated over every SHARED row rather than fired ones - restricting to fired
rows would measure agreement only where the members already agreed enough to
fire, which is the sample most biased toward agreement.

A member whose signed vote never varies (all abstentions, a dead tier) is
SKIPPED rather than counted as r=0, which would drag the mean toward
"decorrelated" using a member carrying no information at all.

Reported as an effective member count, which is the honest way to say what four
models are worth.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 07:41:42 -04:00
AnimateDread
7500e08e17 feat(gate): the payoff number needed a zero-skill book and a by-side split
payoff-v1 reported what a call was worth and nothing to compare it against. A
positive mean R is not a finding on its own: if the instrument drifts, an
ALWAYS-LONG book earns a positive mean too, and drift is the one anomaly family
this project has found that survives cost - so the vote would be reporting the
market's own move as if it were its own.

Two comparisons, and the second is the one that decides it:

ZERO-SKILL BOOK - the same forward move accumulated with a fixed long sign over
every SHARED row, not only fired ones. Accumulated above the voter and
direction-policy filters deliberately: restricting it to bars the vote fired on
would compare the vote against a baseline the vote itself selected. Always-short
is exactly its negative, so one pass covers both.

BY SIDE - the vote's own payoff split by the direction it took, still sign
corrected, at the rung the live signal is actually trading:

  both sides positive          -> directional skill, it pays going either way
  one positive, one negative
  and roughly cancelling       -> it found the drift, and the pooled mean is
                                  saying nothing about skill

This is drift-free BY CONSTRUCTION - drift enters both sides with opposite sign
after the correction, so it cannot manufacture a two-sided positive. That is
precisely what a pooled mean cannot tell you and what no baseline subtraction
fully recovers.

The split is taken at the CHECKPOINTED rung, not this era's derived one: the
derived rung is not known until after the row loop that accumulates the split,
and the checkpointed rung is the operating point the question is actually about.

Still gates nothing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 07:36:25 -04:00
AnimateDread
43c1b27654 feat(gate): measure what a call was WORTH, not only how often it was right
The ensemble deploy gate certifies PRECISION against a chance rate and has
never known whether a correct call pays for its own spread. Every verdict this
project has recorded - 33% precision against a 14% chance rate, an edge that
clears its exact-binomial bar comfortably - is silent on the one question that
decides whether any of it is tradeable, and the cost boundary is exactly where
several earlier edges died with their precision already believed.

Adds a per-row payoff measurement, taken once per ROW (a chart property, not a
member one) at the same time the label is written:

  * forward close move over K = round(SwingLifespanEstimate()) bars,
  * the up and down extreme excursions over the same window,

each divided by the bar's own ATR. K is deliberately the label lifespan the
effective-sample-size deflation already uses, so precision and payoff describe
the same window and can be read in one sentence.

POLICY-FREE: no stop, no target, no trailing rule. It measures the SIGNAL, not
a trade-management choice layered on top - exit shaping moves payoff around
without creating any, so mixing the two would hide which was responsible.

Stored unsigned by direction; the sign comes from the vote at verdict time, and
a short's excursions SWAP rather than negate - negating them would report a
short's worst case as a negative best case.

The newest K bars of the OOS slice have no forward window and are dropped from
the tally with their own denominator, never counted as a zero move: that is the
leading-edge trap that made the lag profile's first run a false positive.

The era verdict now prints mean R, MFE and MAE at the certified rung against
the spread in the same ATR units. It GATES NOTHING - wiring a policy to an
unvalidated payoff number is how a measurement becomes a decision before anyone
has checked it.

Build tag payoff-v1.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 07:14:15 -04:00
AnimateDread
0e1e952b96 feat(topology): cap the input window at 6 bars for capacity - 588 inputs -> 294
Three charts (SP500, XAUUSD, XTIUSD) sat on the FIRST_LAYER_MIN_WIDTH floor
even after pooling took SP500 from 4.1 to 1.8 weights per independent
observation. ComputeFirstLayerWidth needs width <= ~331 to clear it; 49 columns
x 12 bars = 588.

TWO QUESTIONS, AND THE WINDOW IS NOW THE SMALLER ANSWER. The ZigZag ladder
answers "how far back is a swing worth looking" and says 12. The capacity
budget answers "how far back can this much data support" and says 6. Taking the
min stops the first writing a cheque the second cannot cover.

WHY THE LAG AXIS AND NOT THE COLUMN AXIS - the choice was between this and a
per-column mask (designed, parked on feature/column-mask):

  - On the LAG axis there is a measured null. The corrected lag profile finds no
    linear structure at any lag within +/-50, on all six charts, family-wise
    p=1.0000, argmax scattered across different columns and lags per chart.
  - On the COLUMN axis the two measures that would justify a mask - marginal MI
    retention and variance share - are explicitly blind to joint and temporal
    structure, and the columns they would delete include the entire price core,
    which is the one place such structure would plausibly live.

Cutting where there is a measured null beats cutting where the instrument
cannot see. Corroborating: PAI/CONV/LSTM/HYBRID score within ~1pp of each
other, so the temporal machinery is not visibly earning the deeper lags.

THE CAP IS A FLEET CONSTANT, NOT A PER-CHART DERIVATION. Pool rows are keyed on
`bars x columns`, so a capacity cap computed from a chart's own observation
count would differ across the fleet by construction and hand every chart its
own layout, its own fingerprint and its own pool of one - exactly what orphaned
SP500. Set from the most starved chart; every chart shares it.

Conv survives: CONV_RECEPTIVE_FIELD_BARS is 3, so a 6-bar window still leaves 4
sliding positions. LSTM sequence length becomes 6.

RETRAIN-FORCING and POOL-INVALIDATING: width changes, so old .nnw and old
TrainPool rows are both incompatible. Wipe both - which puts the fleet back in
the cold-start condition 6c2959d was written for, and will exercise it.

Build tag -> window6-v1.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 06:16:36 -04:00