forked from mnbvc188199/Warrior_EA
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
d70efc0f74 |
fix(gate): the book had to beat zero, not beat what it costs to collect
WarriorRungBookProfitable tested (hLong+hShort)/n > 0.0. The round-turn spread was
computed a few lines away in the same function, printed in the payoff line, and
carried an explicit comment saying it gates NOTHING - "spreadR is a live snapshot
to be judged by hand".
That objection was correct about the QUANTITY and wrong about the CONCLUSION.
SymbolInfoInteger(SYMBOL_SPREAD) is whatever the book looks like at the instant an
era happens to end, which is not what the historical trades in that window would
have paid - so it should not gate. But this project has already measured that "4 of
5 die to spread and the survivor dies on commission"
(project_cost_boundary_equilibrium), and EXPECTANCY IS THE BAR. The answer is to
measure cost properly, not to leave it out of the gate.
MEASURED, NOT SNAPSHOTTED. g_ensVoteCostR carries the round-turn spread at each
vote row, taken from the broker's own per-bar history (CopySpread, already
maintained as m_spreadSeries for the feature block) and divided by that bar's ATR,
so it lands in the SAME units as the ride book it has to beat. Filled once per ROW
in the block that already measures the ride and the excursions, for the reason
stated there: cost is a property of the CHART at that bar, identical for every
member.
PER RUNG, over the bars that rung actually fires on - not a chart-wide average.
Signals cluster, and a cluster can sit in a wider-spread regime than the window
mean, so a window average would understate the cost of exactly the bars being
traded. sweepCost/sweepCostN accumulate alongside sweepFired.
AN UNKNOWN COST DOES NOT WAIVE THE TEST. g_ensVoteCostR is -1 where the spread
series or the ATR was unavailable and those rows are DROPPED from the mean rather
than read as zero - a cost of zero is the one answer that can never be right. If a
rung has no measurable cost at all the gate falls back to the pre-existing `> 0`
bar, which is "cannot price this", not a free pass.
NOT a multiple of the cost. The "expected payoff should be at least double the
spread" rule is a common heuristic (and is what prompted this - MQL5 article 8410,
Ilin, sent by the operator) but it is not measured here, so it is not imposed.
Sampling error on the book is already carried by the exact binomial bar the same
verdict applies to precision.
WHAT IT CHANGES, on the fleet as it stands. Median book per call against the
census's measured spread/ATR:
SP500 +0.02 vs ~0.065 -> now correctly REFUSED (was certified)
GBPUSD +0.05 vs ~0.020 -> marginal
DAX40 +0.08 vs ~0.029 -> marginal
BTCUSD +0.23 vs ~0.006 -> clears
XTIUSD +0.57 vs ~0.083 -> clears
NAS100 +0.75 vs ~0.025 -> clears
Both charts currently DEPLOYABLE clear it comfortably, so this closes a hole rather
than reversing a live decision - but SP500's book was being certified while sitting
below its own spread.
The sweep line now prints cost beside the book at every rung, so BOOK-FAIL names a
visible reason instead of an invisible one, and its legend says what the number is
and what "n/a" falls back to.
Three call sites updated, not one - the rung walk, the derived-rung verdict and the
sweep report all ask this question (feedback_rename_leaves_readers_behind).
Not retrain-forcing: no fingerprint member moves, and g_ensVoteCostR is an in-memory
per-era buffer.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
e13e317e54 |
feat(signal): one operating point per side - a single global cut had silenced the sell book
User report: "some charts are not showing sell signals at all, only buys". It is
real. By-side ride book and n on the fleet's most recent era per chart - the
drift-free test the era log already prints:
NAS100 160 long vs 5 short - 32 : 1
DAX40 30 long vs 0 short - no short book at all
(Earlier eras from the same charts read 744:1, 260:1 and 41:1, but those lines
predate today's restart and are quoted nowhere as current - the two above are the
post-restart measurement.)
THE LABEL IS NOT THE CAUSE. P(ride pays | up-leg) vs P(| down-leg) is 0.99x to
1.55x across the fleet and up-legs are ~50% of bars. A 1.2x rate asymmetry was
being amplified into 744:1.
THE MECHANISM is the operating point's depth. FitDirConfThreshold places one cut
where coverage matches the pooled directional label rate, and at 3-25% coverage
that cut sits deep in the tail. There, a small shift in one side's margin
distribution moves nearly ALL of that side across it: the threshold hits its
target POOLED and misses it per side by two orders of magnitude. The model was not
failing to see sells - BTCUSD's pre-threshold sell recall is 84% and its raw argmax
is B45/S40 - the single cut was discarding them.
Each side now clears the margin ITS OWN measured label rate asks for.
m_labelPrebuildBuyCount/SellCount already existed, so the two targets are
measurements, not choices.
THE PROPERTY THAT MAKES THIS SAFE, AND IT IS EXACT: the per-side targets are
100*buy/tot and 100*sell/tot against the same denominator the pooled rule uses, so
they SUM TO THE POOLED TARGET. Total coverage is preserved; only its split between
the books changes. This redistributes exposure rather than increasing it.
THE TRAP, NAMED:
|
||
|
|
3c0e2facb2 |
fix(chart): a signal mark's persisted time was the left edge of its line, not its bar
User report: "some labels are completly off (arrow not at the low/high, and entry
price way higher than candle body and much much away from what the spread would
add)". Not the training label, and not a fill - the only order attempt on the day
was refused by the client. Both halves are one display defect.
A mark is TWO objects since 2026-08-20: an OBJ_TREND segment from t-half to
t+half so it is wide enough to see, plus an OBJ_ARROW glyph. half is
PeriodSeconds * WARRIOR_SIG_LEVEL_HALF_SPAN = 4680 s on H1, so OBJPROP_TIME index
0 of the line is its LEFT EDGE. FIVE call sites read index 0 and every one of
them treated it as the bar time, because "which bar does this mark belong to" had
no single implementation and each site re-derived it from rendering geometry.
MEASURED on the live files, and it COMPOUNDS: Snapshot() wrote t-half, the restore
redrew a line centred on the stored value, and the next save read THAT line's left
edge. Predicted residues are the 10 multiples of 360 s, not the 60 possible
minutes - observed 704 of 704 persisted vote arrows across fifteen files on the
k*1080s ladder, ZERO off it, with k=1 where a mark had been saved once and k up to
6 where it had been carried through sessions. XAUUSD and EURUSD sat at k=6: 7.8 H1
bars adrift, which is why a mark's stored trigger price was nowhere near the candle
it was drawn over - the price was right for a bar 7.8 bars away. The arrow half
lost its low/high anchor for the same reason: a shifted time is not a bar open, so
iBarShift(exact) returned -1 and the glyph fell back to the trigger price, landing
inside the candle body.
WarriorSignalMarkBarTime() is now the one accessor, and it takes the MIDPOINT of
the two ends rather than t0+half: the midpoint recovers t exactly by construction
AND stays correct if the half-span is ever retuned, including for marks already
drawn under the old value. Its t1 <= t0 branch answers correctly for the
single-anchor arrow half too, which the rescan sweep needs. Routed through it:
CVoteArrowStore::Snapshot, CChartUI::SaveChartSignals,
WarriorReconcileVoteCooldown, WarriorLatestVoteArrowTime, and the rescan's
typed-blind scope sweep.
Two more defects the same read was hiding:
* the cooldown reconciliation deleted by a RECONSTRUCTED name built from the
shifted time it had just read, so it could not remove a freshly drawn arrow at
all and its log over-reported kills. The name now travels with the time
through an insertion sort over both arrays.
* the rescan's scope sweep gave the two halves of ONE mark two different times,
so at the window edge it deleted an arrow and left its line - exactly the
split WarriorDeleteSignalMark exists to prevent.
* WarriorLatestVoteArrowTime seeded the live cooldown clock 1.3 bars EARLY on
every restart, so the first signal after a restart could fire inside the
window the arrow on the chart was enforcing.
WarriorSignalMarkOnBarGrid() stops the drift surviving a restart. Arithmetic
rather than a history search, and the residue is READ off iTime(sym,period,1)
instead of assuming UTC alignment, because where the bar grid sits in epoch
seconds is the broker's day start. It KEEPS on an unknown - no history yet, or a
weekly/monthly frame that is not a modular grid - since deleting on an unknown is
the failure mode that cost this chart 272 of 273 arrows in
|
||
|
|
8280a7cd40 |
feat(fleet): open the charts the census says are worth having
The census answers which instruments carry deep history at a spread the book can cover; this is the half that acts on it, so adding an instrument is a reviewable list in source control rather than six manual chart operations nobody can reconstruct later. BTCUSD, GBPUSD, NAS100, DAX40 - every one with more than twelve years of server H1 and a sampled spread under 3% of ATR, against a measured ride book of +0.05 to +0.63 ATR per call. BTCUSD is the cheapest deep instrument the broker offers (0.63%) and the only one in a different asset class, which is worth the most to a pooled certificate that declines the diversification credit. DAX40 trades a different session, so its bars are not the same hours as everything else. IDEMPOTENT: a chart is opened only when none exists on that symbol and period, so a restart re-opens nothing and the list can be edited freely. That is what makes it safe to leave armed. THE TEMPLATE CARRIES THE EA. ChartSaveTemplate on the running chart captures this Expert Advisor and its inputs, so a new member comes up configured exactly like the one that spawned it - which is the point: a fleet whose members differ by attach order is how four charts ended up on a 10-bar cooldown and two on 30. Leased like the census and the alt-data fetch, because six instances would each try to open the same four charts - and the charts this opens initialise an EA that reaches this same code. COST, STATED IN THE FILE: four added charts are sixteen more models on a six-core box already at ~72% with six charts. Expect eras to slow across the whole fleet. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
7a5dddc61a |
fix(census): one chart, not six - and sample the spread instead of snapshotting it
TWO DEFECTS IN MY OWN CENSUS, both found by reading its output rather than by the compiler. IT RAN ON ALL SIX CHARTS. The guard was a plain global, and MQL5 globals are per PROGRAM INSTANCE - six charts are six instances, so every one of them walked all 64 symbols and overwrote the same file. Replaced with the atomic GlobalVariableSetOnCondition lease that System\AltDataFetch.mqh already uses for exactly this problem: six charts racing one shared file. Losing the lease is the correct outcome, not an error. THE SPREAD WAS ONE SNAPSHOT. Cost is the number this whole decision turns on - the measured ride book runs +0.05 to +0.63 ATR per call, so an instrument costing a tenth of an ATR a round turn has already spent most of what it could earn - and a single reading is one moment of one session. Spreads widen at rollover and around news. Now ten readings thirty seconds apart, reporting mean AND max, because the max is what says whether an instrument is quietly untradeable at the wrong hour. A reading only counts behind a real two-sided quote; a symbol with none reads NO-QUOTE rather than 0, since 0 would rank it as the cheapest instrument on offer. That is the same error the previous commit fixed one layer up, and the guard now sits at the sample rather than only at the report. WHAT THE FIRST GOOD RUN ALREADY SETTLED: 64 symbols openable both ways, and all 64 report a SERVER first-H1 date with server == local. So depth is real, not a sync artifact - the broker genuinely offers deep H1 on twelve instruments and added the other fifty-two in July/August 2026 with no history at all. The expansion universe is six symbols, not fifty-eight. Also promotes the derived-cooldown line out of PrintVerbose. It reports a change to the TRADING POLICY, and this codebase's rule is that a line reporting a state change never sits behind the verbosity gate - the tier-ladder restore learned that the expensive way. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
81efd64a7f |
feat(signal): derive the cooldown from the measured label lifespan, and census the broker's symbols
TWO CHANGES, ONE CAUSE: a per-chart input could not reach the fleet, and the fleet
had no data on which symbols were worth adding.
THE COOLDOWN IS NOW DERIVED, NOT SET. An input was the wrong shape for it twice:
* It is a measurable property of the LABEL, not a preference. The leg-ride label
resolves when the ZigZag leg flips, so the mean label lifespan IS the average
leg duration in bars. Two calls closer together than that concern the SAME
leg - the second pays a second spread for a move the first already owns. One
lifespan apart is where consecutive trades concern DISTINCT legs, which is
also what makes the deploy gate's independence assumption exact rather than
approximate: EffectiveSampleSizeDeclustered's divisor becomes 1.
* An input could not reach an attached EA. MT5 stores inputs per chart, so
"raising Signal_CooldownBars from 10 to 30 changed nothing on six live
charts" - and yesterday the fleet was found running 10 on four charts and 30
on two, two different trading policies inherited from attach order rather
than chosen. A derived value cannot drift that way.
The fraction is 1.0 because the argument picks it: less re-admits same-leg
duplicates, more declines distinct legs for no stated reason. At the measured
33-34.5 bar lifespan it lands within a few bars of the 30 the default intended.
Recomputed wherever the lifespan is measured - the label-cache build - so the
measurement and its consumer cannot drift apart. SignalCooldownOverrideBars still
wins, and switching declustering off entirely is still possible.
THE SYMBOL CENSUS answers "which symbols are worth adding" with data. The binding
constraint on this system is independent observations, and instruments are the
only lever that multiplies them, so it writes the three numbers that decide it:
history depth, spread against ATR, and whether both sides can be opened at all (a
close-only symbol can never satisfy the gate's two-sidedness test).
AND IT RUNS ON THE TIMER, NOT AT INIT, which the first version got wrong. At init
the terminal has just reconnected and nothing has a quote, so SYMBOL_SPREAD reads
0 everywhere - the first run duly ranked thirty untradeable-on-cost symbols as the
cheapest the broker offers. Caught by noticing GBPUSD reported a zero spread
beside EURUSD's 2 points. Now deferred three minutes, and a symbol with no tick is
written NO-QUOTE rather than a number, so the column cannot be sorted on by
mistake. It selects nothing: sixty symbols added to Market Watch is a change to
the operator's terminal, and this is a report.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
8fd2756aa5 |
fix(gate): the member gate certified a population the EA never trades
The deploy gate judged m_oos.DirCalls() - every bar the prior-corrected posterior
fired - and deflated it by the full 34.5-bar label lifespan. The EA does not trade
that set. It trades what survives declustering: live, a call the cooldown rejects
has its signal zeroed before the vote is published, so it produces no arrow, no
vote and no position.
The honest counts were already being computed, by the same CSignalDeclusterPolicy
object the live signal applies, on the same adjusted decision it feeds - printed
every era as "TRADED (declustered)". The gate simply never read them.
BOTH HALVES OF THE ERROR POINTED THE SAME WAY, which is what made it expensive:
* the traded set is measurably CLEANER. Live USDJPY eras: judged 24% where
traded was 26%, judged 26% where traded was 27%. The gate was reading a
precision the EA would never have realised.
* the traded set is far more INDEPENDENT. The cooldown is 30 bars against a
34.5-bar label, so consecutive traded labels overlap by at most 4.5 bars - the
stream is ~87% independent, not ~3%. Deflating it by the full lifespan applies
a correction the cooldown has already made.
On a live USDJPY era that is the whole verdict: judged 24% against a 25.4% bar
FAILS; traded 27% against a 24.7% bar PASSES.
AND THE SELECTION IS UNBIASED, which is what makes the traded precision usable at
all: the declustering keeps the chronologically FIRST bar of each run, never the
highest-confidence one, so this is not cherry-picking winners.
COVERAGE STAYS ON THE SIGNAL POPULATION, and that split is now explicit in the
signature rather than implied. Coverage asks "did this model call often enough to
be a strategy", which is a question about the signal; precision asks "were the
calls it took right", which is a question about the account. Coverage cannot move
to the traded set for a structural reason: a 30-bar cooldown caps traded coverage
at 1/30 = 3.3% while the floor is a quarter of a ~20% base rate, so every chart
would fail it forever for reasons unrelated to the model.
THE ENSEMBLE GATE STILL HAS THIS DEFECT and is passed through unchanged, on
purpose. Its OOS rows carry each member's raw adjusted vote - the decluster replay
is per-member state that never reaches the shared vote buffer - so there is no
traded population there to read yet. Fixing it needs the per-member decluster
decision carried on the vote row. One gate at a time, so the effect of this one
stays attributable.
Not retrain-forcing: gate arithmetic only, no fingerprint field moves.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
6d266d4e62 |
feat(features): every price column here is memoryless - fractional differentiation
AFML ch.5. Every price-derived feature in this set is a FULL difference: (close-open)/atr, (close-MA)/atr, the 20-bar return, the leg extension. Differencing is what makes a series learnable - a model fitted on 2019 EURUSD levels cannot read 2026 ones - but a first difference is memoryless BY CONSTRUCTION. It keeps the last step and throws away the series. That is the trade this feature set has been making silently at every column. Fractional differentiation is the observation that the exponent need not be an integer. (1-B)^d for 0<d<1 interpolates between the raw level (all memory, not stationary, useless to a learner) and the return (stationary, no memory). Three orders are emitted - d = 0.3/0.5/0.7 - in ATR-relative units. A LADDER, NOT A CHOSEN d. AFML picks the minimum d passing an ADF test; there is no ADF test here, and adding one to tune a feature's parameter per instrument per era would fit the feature to the data before the model saw it. Three fixed orders go in and the FLEET KEEP-SCREEN votes on them, exactly as it has already deleted five ZigZag geometry columns and demoted the alt-data block. Kept by construction in mask v4 for one outing only: the screen reports only on columns that are EMITTED, so a block masked off at birth can never be measured. THE BUG THAT WOULD HAVE SHIPPED THREE DEAD COLUMNS, caught before deploy by checking the arithmetic at realistic price levels rather than on a series starting at zero: The weights of (1-B)^d sum to zero in the LIMIT - that is what makes it a differencing operator. TRUNCATED at 64 bars they do not: the residual is 0.222 at d=0.3. That residual multiplies the LOG PRICE LEVEL, so the raw sum carries 0.222 x log(2400) = 1.73 on XAUUSD against 0.222 x log(1.08) = 0.017 on EURUSD - an instrument-identity constant hundreds of times larger than the signal. Divided by the relative ATR it pins every bar to the clamp: measured 100% of XAUUSD and SP500 bars clipped, 68% of EURUSD. Three constant columns that FEATURE HEALTH would have reported only after a full retrain had been spent on them. Anchoring every term at this bar's log price removes exactly the level component and leaves a fracdiff-weighted combination of the multi-horizon RETURNS ending here - stationary, scale-free, same meaning on every instrument, and the long memory intact. Same series after: mean ~0, sd ~1.1-1.3, nothing clipped. The weight recursion is checked against two known values: at d=1 it terminates to [1,-1,0,...], the plain first difference, and at d=0.5 it reproduces the standard expansion of (1-B)^0.5 to six terms. REJECTS RATHER THAN DEGRADES at the oldest edge, unlike the swing and volume windows beside it. A 30-bar Donchian range is still a Donchian range; a fractional difference over a short window is a DIFFERENT OPERATOR - a different effective d - reported in the same column, which is a silently wrong number rather than a degraded one. Costs ~64 bars of 130,000. RETRAIN-FORCING twice over: the input width is field 4 of the fingerprint and FEATURE_MASK_VERSION rides it too. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
f063cfd6ba |
fix(calibration): the confidence rescale was blended with a per-SAMPLE time constant, once per era
m_confidenceCalScale has been frozen at its 1.0 constructor default for the life
of the mechanism, and the units error that froze it is one line:
m_confidenceCalScale += (eraScale - m_confidenceCalScale)
/ Net.recentAverageSmoothingFactor;
recentAverageSmoothingFactor is 10000. It is a PER-TRAINING-SAMPLE constant - it
exists to average the network error over ten thousand samples inside backprop.
Applied ONCE PER ERA it moves the value by one hundredth of one percent, so after
a hundred eras the scale has travelled 1% of the way to its target and after a
thousand it is still not halfway.
That reframes what was already known. ConfidenceBridge.mqh records that the
confidence is miscalibrated and that five trade-management modes were deleted
because of it. The mechanism meant to fix it was not merely shape-blind, as the
isotonic curve's commit message argued - it was NUMERICALLY INERT, and never had
the chance to correct anything at all. A quantity blended per era needs a per-era
time constant; borrowing one from a per-sample loop reads as a working mechanism
in every review, because the line is shaped exactly like a working EMA.
The new curve inherited the same divisor when it was written yesterday, which
would have frozen it at its first fit - visible in the log as a "carried" column
identical to "refit" to four decimals on every chart. Both now use CAL_BLEND_ERAS
(5 eras): slow enough that one noisy band cannot swing the live number, fast
enough to track a model whose output distribution moves every era.
VERIFY, DON'T ASSERT: the curve's log line now prints the scalar's current value,
so the first line after this deploy reports the number restored from .stats - the
value it reached over that model's entire history. 1.000 is the claim above,
measured rather than argued from the arithmetic.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
5ec8902020 |
fix(gate): the zero-skill floor was still halved one level down - and the measurement sizing actually needs
THREE FINDINGS, ALL FROM READING THE FLEET'S OWN LOG RATHER THAN THE COMPILER. 1. THE MEMBER DEPLOY GATE STILL USED THE 3-CLASS ZERO-SKILL FLOOR. The ensemble gate's chance rate was corrected on 2026-09-03 when the head stopped choosing sides. SOosTally::ChancePrecPct - one level down - was not, and it feeds more than the ensemble copy does: the MEMBER deploy gate, the certified pair HasDemonstratedEdge() reads to decide whether a member may vote at all, and the cross-instrument pooled certificate. On a live H1 window it read 12.5% where the honest always-ride floor is 21.1%, so every member showed "+13pp edge, PASSES" and was admitted to the vote. A model with no skill whatsoever cleared it. The same halved floor was in BaselineComparator's Alglib comparison, which scored the forest and the linear baseline against a bar half the height of the one the net is judged by. ChancePrecPct now TAKES THE POLICY AS A PARAMETER WITH NO DEFAULT, so the next head change is a compile error at every reader instead of a silent wrong answer at some of them. That is the whole lesson of finding this one a week late: the first fix was applied where the bug was noticed, not everywhere the assumption lived. The era line said "the gate ranks on the LARGER of the two" for that entire week, on six charts, every era. It now states the policy actually in force. 2. THE ERA LINE PRINTED +/-1.79e308 IN THE MIDDLE OF EVERY RECORD. The binary head has no third output neuron, so slot 2's min/max kept their DBL_MAX ctor values and were formatted anyway. Width-aware now, and the slots are labelled PAYS/DOESNT rather than B/S/N, which is what they hold. 3. THE SIZING QUESTION NEEDED A DIFFERENT MEASUREMENT THAN THE ONE I BUILT. The reliability curve says the confidence is now HONEST - carried, out-of-sample, ECE 38pp -> 1.3-2.4pp. It says nothing about whether it RANKS, and ranking is what a bet size needs. The existing conviction curve cannot answer it either: coverage collapses above the lowest rung, so every fired call sits in one bucket and there is no curve to read. So the ensemble now carries the calibrated confidence per OOS row - the mean over members that actually called, which is exactly what LiveSignedConfidence() publishes - and the era verdict splits the certified rung's fired calls at their median confidence and compares what the two halves earned, in ATR per call. A MEDIAN SPLIT, NOT A DECILE CURVE, and the reason is power: per-call SD is ~3 ATR and the labels overlap ~34 bars, so a decile of a few hundred raw calls holds under ten INDEPENDENT ones and its error bar is wider than the whole book. Ten noisy points would invite the best-of-N reading this project has already crowned four times. Two halves is the most the data can be asked for, and it is reported with its standard error on INDEPENDENT counts plus the size of difference this window could ever resolve - so "not measurable" is distinguishable from "no effect", which is a statement about the data rather than a verdict on the idea. Nothing here sizes anything. This is the evidence ConfidenceBridge.mqh's standing rule demands before anything may. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
18ada65d85 |
feat(calibration): the refit column is not evidence - carry the curve and measure the live range
Two corrections to what the first commit reported, both found by reading its own first live output rather than by the compiler. THE "AFTER" NUMBERS WERE IN-SAMPLE. ECE 44.7pp -> 0.5pp was the curve scored on the very band it had just been fitted on, and isotonic regression is expressive enough to drive that near zero whatever the data says. It is the attainable floor, not a result. Quoting it would be the calibration-slice error - a threshold fitted on memorised bars - one layer along. The fix is free, because the previous era's curve is already sitting there when the new one is fitted: score it too. That curve was fitted on EARLIER bands and has never seen these bars, so it is a genuine out-of-sample calibration measurement. The log now reads raw -> carried -> refit, and says in the line itself that the middle column is the one to read. THE SPREAD, NOT THE ERROR, IS THE SIZING QUESTION. The first fit showed claimed 0.75 and claimed 0.99 mapping to the SAME calibrated 0.273 - the curve is flat at the top, meaning higher confidence there does not mean higher accuracy. A well-calibrated constant is still a constant: no amount of calibration makes a flat map rank anything, and a bet scaled by it would be a bet scaled by noise - the precise error that killed the five confidence-scaled modes on 2026-08-25. So the report now restricts to the calls the operating point actually ADMITTED - the only ones that can become a trade - and states the calibrated probability at the lowest and highest admitted bin, their spread, and the map's mean against the realised rate on that same set. Carried as its own counts rather than derived by cutting the histogram at a magnitude, because the operating point is a MARGIN and this histogram is keyed on a MAGNITUDE; those are monotone in each other on the binary head and not on the 3-class one, and a reparameterisation guess is exactly the kind of thing that reads as a measurement. This is the number the sizing decision will be made on, and it is deliberately being gathered BEFORE anything is built on it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
92ccf20385 |
feat(calibration): a scalar has no shape - the isotonic reliability curve
The model's confidence was corrected by ONE number: accuracy divided by mean
claimed confidence, EMA-blended per era. That can move the reliability curve up
or down and can do nothing else. A model that is honest at 0.55 and wildly
over-confident at 0.95 has a SHAPE problem, and no scalar has a shape.
The scaffolding for the fix already existed and was better than expected.
RunCalibrationPass already walks a PURGED, held-out band with batch norm frozen,
visiting each bar exactly once, and fills a 50-bin (margin, hit-rate) histogram.
That is a reliability diagram on out-of-fold predictions. Isotonic regression is
a fit over data already being collected, on a band already purged - no new walk,
no new holdout, no new cost.
WHAT THIS ADDS
* A SECOND histogram in the same walk, keyed on |dPrevSignal| - the magnitude
the consumer actually holds - not on the winner-vs-rival margin. Deliberately
not a re-key of the existing one: on the 3-class head the two statistics are
not the same quantity, and FitDirConfThreshold's own note records what
happened the last time one curve served two fits.
* FitCalibrationCurve: pool-adjacent-violators over the occupied bins, weighted
by call count. The least-squares monotone fit (Ayer et al. 1955); the
monotonicity constraint IS the regularisation, so there is no smoothing
parameter to tune and it cannot fit a shape the data does not show.
* EMA-blended across eras with the same smoothing the scalar used. A convex
combination of monotone sequences is monotone, so blending costs nothing the
fit exists to impose.
* Brier, ECE and MCE reported before and after the map, on the band the map was
fitted on. REPORTED, NOT OPTIMISED - nothing selects on them. A model that was
already calibrated shows all three pairs unchanged, which is the outcome that
says this map is not needed.
* Persisted (.stats WSTE), because the curve is produced only by a completed
calibration band and a DEPLOYED model runs no more eras - the same failure the
tier ladder had. The bin count leads the block so a future
DIR_CONF_THRESHOLD_BINS change is a mismatch the reader detects rather than
fifty doubles landing in the wrong slots.
THE SCALAR STAYS as the unfitted-model answer, and the two are never applied
together: they answer the identical question, the curve per magnitude and the
scalar on average, so stacking them would correct the same error twice. Same
discipline as the logit adjustment being backward-pass only.
NOT RETRAIN-FORCING. The fingerprint is untouched, the .nnw is untouched, and a
WSTD file still loads - it simply comes back with no curve and keeps the scalar.
The fleet resumes exactly where it was.
STILL TELEMETRY. Variables\ConfidenceBridge.mqh carries a standing rule that
nothing there may steer a trade, and CalibratedConfidenceMagnitude has exactly
one consumer: the trade journal's aiConfidence bucket. Five confidence-scaled
trade-management modes were deleted on 2026-08-25 for one stated reason - the
confidence was known to be miscalibrated, so they scaled money by a quantity
whose units were never established. This is the measurement that establishes
them. Sizing is deliberately NOT in this commit: shipped alongside its own
calibration it would be untestable, because if the book moves nothing says which
half did it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
84789ea992 |
fix(head): the binary logit branch printed the 3-class banner
ApplyLogitAdjustment's meta-head branch sat BELOW the 3-class tau computation and its log block. The offsets installed were correct - the branch returns before the 3-class ones are written - but the 3-class banner printed first and latched m_logitAdjustLogged, so every era log claimed "APPLIED across all three" while the head had two logits, and the binary line was unreachable. A log line that misdescribes a live gate-adjacent path is the same class of defect as a stale comment. Moved above the 3-class work. Also corrects this file's own claim about the imbalance against the measurement it now has: the H1 fleet's shares are 9.4/9.0/81.6, so the binary split is ~4.4:1 rather than the ~3:1 estimated before the run - and tau goes from 0.57 capped on three classes to ~0.86 on two, which is the point: most of the correction gets through where most of it was being clipped. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
8d6cff7fcb |
feat(head): the side was never a prediction - the binary meta-label head
The leg-ride target's SIDE is the direction of the ZigZag leg in progress as of that bar, readable from bars t and older with no lookahead. The 3-class head made the network re-derive it anyway: that is Lopez de Prado's PRIMARY MODEL being learned instead of used, and it cost on four axes at once - chance at 33% instead of 50%, an 8:1 imbalance instead of 3:1 against a tau already pinned at its cap, every Buy row spent as evidence about "is this a long leg" rather than about payoff, and a directional error scored the same as a payoff error though only one of them was a question we asked. The head is now binary: slot 0 "riding this leg to the flip pays at least LEG_LABEL_MIN_RIDE_ATR", slot 1 "it does not". The side comes from LegDirAsOf() at read time. TWO OUTPUT NEURONS, NOT ONE, and that is what made this small. Both backward passes in AI\Impl\NetForward.mqh already carry a `total == 2` softmax+CE arm, left in deliberately when the old meta head was removed on 2026-08-25 - and a 2-class softmax IS a logistic/BCE head, the logit difference being the log-odds. So this reuses the exact gradient path the 3-class head uses instead of growing a second one. NOTHING DOWNSTREAM CHANGED. ApplyClassificationSoftmax() expands (pGo, side) back into the [pBuy, pSell, pNeutral] triple every reader already consumes, so Argmax3, the operating-point histogram, per-class recall, confidence tiers, the vote currency and the chart arrows are untouched. Neutral stops being a class the net competes for and becomes what it always meant: pGo below the operating point. The margin the threshold is expressed in becomes 2*pGo-1, monotone in pGo, so the calibration walk fits the same statistic. THE HEAD STAYS BOUNDED, deliberately, and the 2026-07-27 unbinding retry is NOT bundled here. Reading the gradient showed why it need not be: the softmax arm OVERWRITES the output neuron's gradient with (target - softmax), so the sigmoid derivative and MIN_ACTIVATION_DERIVATIVE are already bypassed at the head. What SIGMOID x CLASS_LOGIT_SCALE actually costs is p in [0.0025,0.9975] - three orders of magnitude wider than the band where the live question is 0.35 versus 0.60. Unbinding has its own failure history and deserves its own measurement; batch norm before the head, its precondition, already ships. THE ZERO-SKILL FLOOR HAD TO MOVE WITH IT. The rate gate's chance was the better of always-Buy and always-Sell. This head cannot choose a side, so its no-skill policy is ALWAYS-RIDE - every bar taken in its own leg's direction - which is right on EVERY directional-label bar, not the better half. Left alone it would have handed a model with no skill whatsoever a ~+11pp edge. The book gate already measured against always-ride; the rate gate now agrees with it about what zero skill means. Same correction in the module-weight shrinkage prior (Lifecycle.mqh's 50.0 side coin-flip). RETRAIN-FORCING twice over: |MHEAD:1 joins the fingerprint and the output count is field 6 of the .nnw filename, so no existing model or pool row can be adopted. Verified live - all six charts rejected every peer file by fingerprint and restarted from era 0. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
8d7cbdbf9d |
feat(training): sample the queue by uniqueness instead of weighting it - the update itself was the cost
Weighting by uniqueness (
|
||
|
|
8351c3933e |
feat(training): weight every sample by its label's uniqueness - the loss was counting 34x redundancy
Every statistic in this program deflates overlapping labels. The training loop did not. EffectiveSampleSize() is called in exactly two places, the feature keep-screen and the deploy gate, and both are reporting paths; the gradient step weighted each sample by |ride| alone, with no term for how many other labels are made of the same bars. Measured on the H1 fleet: mean label overlap 31.6-37.4 bars. USDJPY carries 125,214 labels that this same binary reports as ~3,344 independent ones. The optimiser was told it had thirty-four times more evidence than it has, which is the textbook cause of a strong in-sample fit with a thin out-of-sample book - the pattern every era log has shown since this label shipped. The correction is Lopez de Prado, Advances in Financial Machine Learning ch.4: average uniqueness is 1/concurrency averaged over the bars a label spans, and the prescribed weight is uniqueness x attributed return, i.e. the product of the new Variables\UniqueWeight.mqh and the existing MoneyWeight.mqh. - ComputeLabelUniqueness() runs once when the label cache pre-build completes, O(bars) via a difference array for concurrency plus a prefix sum over 1/c, so a 125k-label cache costs two sweeps rather than millions of span walks. - On the leg-ride target the two weights are anti-correlated: a label's span IS its ride, so a long leg earns a big money weight and sits where concurrency is highest. Money weight alone concentrated gradient on exactly the least independent evidence in the set. The product is re-clamped to [0.25, 4.00] because two clamped factors multiply to a 16x tail. - Measured even when the knob is off, so the report can show the spread this would apply without training on it. First run: mean uniqueness 0.076-0.085. This REDISTRIBUTES gradient mass; it does not reduce the update count. Training on fewer, more independent samples is the sequential bootstrap (ch.4 s.4.5) and gets its own commit and its own measurement. UNIQUE_WEIGHT_VERSION 0 restores the prior behaviour and the prior fingerprint exactly. RETRAIN-FORCING (|UWGT:1). Compiled 0/0. Deployed 19:40 as uniq-weight-1; all six H1 charts from era 0. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
3aa15c8b4b |
feat(altdata): publication stamps by series, alt block back in, one pool for all six charts
Operator asked for the alt data to be properly mapped on H1. Three findings: 1. The as-of join was already whole-day: any H1 bar of day D reads row D, and the window layout (ALTW:2) carries that reading once per window at the anchor. What was wrong was the ROW DATE. Every FRED series was stamped "knowable next day", which is right for a market close and five to six weeks early for a monthly print: July CPI (dated 07-01) entered the export on 07-02 and was released 08-12. Unemployment the same; the H.10 dollar index (weekly, posted the following Monday) a week early; the effective funds rate a day early. On H1 that is ~1,000 bars of a value nobody had, in three of the twelve fleet columns. FredPublishLagDays() stamps by series (CPI +48d, UNRATE +40d, DTWEXBGS +8d, DFF +2d, closes +1d), cached rows are re-stamped on load, and ALTFETCH_EXPORT_VERSION (a .ver sidecar beside each export) forces one rebuild at the next init so every chart reads the corrected export immediately. Verified on XAUUSD_D1.csv: mac_cpi now changes on 08-18, mac_unemp on 08-10. 2. The alt block was not reaching the model at all. Keep mask v2 dropped all twelve alt columns on a screen measured under the pivot label on H4, and the screen only reports on emitted columns. v3 emits them again (28 of 47 columns, input width 168); the H1 keep-screen will say which of them clear. The window dedupe now uses the EMITTED alt width, not the panel's: under v2 it placed a 12-wide block over the last twelve of sixteen emitted columns. 3. The training pool ran as two groups because the cross-asset block encoded index mode (base == quote: SP500, and this broker's XAUUSD/XTIUSD) with a different meaning per slot than FX mode, and the fingerprint tagged it ":IDX2". U3 gives both modes one layout (proxy fast/slow in 0/2, denomination fast/slow in 1/3, own move minus the proxy-vs-denomination cross in 4), so all six charts print the same fingerprint and pool together. RETRAIN-FORCING (XA:6:U3, ALTV:2, FMASK:3). Deployed 15:00 as alt-stamps-1; all six H1 charts from era 0, one shared fingerprint, exports rebuilt at v2. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
8c37869c63 |
fix(init): a warm re-init kept the H4 run's globals; the training pool adopted other timeframes
Operator report 2026-09-02: the six charts were switched from H4 to H1 and the status label kept showing the H4 run. MetaTrader does not unload the program on a timeframe, symbol or input change - it runs OnDeinit and OnInit inside the same instance and every file-scope global survives the pair. The members' converged flags, the live vote line, the panel rows, the shared best-era/plateau state and the once-per-chart report flags all belonged to the models just torn down. - Warrior_EA.mq5: WarriorResetWarmReinitState() runs first in OnInit and puts every such global back to its cold-start default. Kept on purpose: the chart's book magic (positions opened before the switch stay owned), the alt-data fetch throttle (rate-limited APIs), the tester profile, the RNG, the OpenCL flags. - TrainingPool.mqh: peers must be the caller's own timeframe. The fingerprint does not carry the period, so the first H1 census adopted 39,754 H4 rows from three peer files and credited them to the capacity budget. Files are named SYMBOL_PERIOD.bin; the suffix decides, and the reader names the rejection. Build tag warm-reinit-1. Compiled 0 errors / 0 warnings. RETRAIN-NEUTRAL. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
980f10b60c |
feat(label): ride the leg, exit on the flip - the leg-ride target replaces the pivot-event target
The pivot label paid +1.6 ATR on a hit and -1.7 on a miss at 50-65% precision: a zero book by arithmetic, because a "bottom" that is not one is a move that kept going, and the model called it because a big move had just happened. This target never asks for a turn. Direction is the leg in progress, known on the bar; the label is whether riding it from here until the leg flips pays at least LEG_LABEL_MIN_RIDE_ATR (1.0). The flip is the exit the EA now places. - Labeling/LegState.mqh: a line-for-line replica of ZigZag.mq5 (12/5/3) run over bars <= t only, so "which leg am I in" is what the chart showed on t, never the final buffer. Verified against the real indicator at every init (LEG STATE REPLICA line) and by Tests/Test_LegState.mq5 on hand-built bars. - LegRideLabel (Labels.mqh): ride = legDir * (close[flip] - close[t]) / ATR, the flip found by asking the same as-of function of each newer bar in turn. Unresolved until the leg has flipped inside loaded history. Lifespan = the leg, so EffectiveSampleSize deflates honestly (~3x fewer than the window constant claimed); the topology's overlap is the median leg again. - Money weight = |ride|. Online step's exit bar = the flip. - Three as-of leg features in the swing block (direction, extension in ATR, age); FMASK:2 keeps them. Fingerprint TGT:LEG1:10 - RETRAIN-FORCING. - Era verdict: the book WarriorRungBookProfitable gates is the RIDE (entry at the call, exit at the flip), printed per rung and at the certified rung beside ALWAYS-RIDE (zero skill) and the ORACLE RIDE (the ceiling). Fixed- horizon payoff stays as the signal diagnostic. The by-distance profile, its slot arithmetic and the reversal-exit pass are deleted (measured: the vote's reversals land 33-137 bars late; dead). - Live: Exit_On_Leg_Flip (default on, not in the fingerprint) closes when the as-of leg flips against the position and places no take-profit; the measured stop stays. Build tag leg-ride-1. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
39e4bb587a |
refactor(signal): delete the barrier sweep, the automatic direction policy and the pivot-geometry features
Three layers out, ~1,000 lines, RETRAIN-NEUTRAL - the fingerprint and the emitted feature vector are byte-identical, and a WSTC .stats still restores under the new WSTD reader. BARRIER SWEEP. The 64-cell (stop, target) grid, its 999-draw Westfall-Young permutation null, the first-touch offsets recorded per OOS row, the swept-pair persistence (WSTA) and the tier-1 branch of ResolveBarrierMultiplier are gone. It never cleared its own null on any chart (p 0.19 to 1.00), so it gated nothing and cost a grid plus 999 rescans per era. Live stops and targets come from the mean excursions at the hold horizon, as they did in practice. DIRECTION IS THE INPUT. WarriorEffectiveDirection() returns tradingdirection. The bar-body drift screen (Signal\DriftScreen.mqh) and the 2-SE by-side expectancy test that outranked it are removed with their unfiltered accumulator and persistence (WSTB/WSTC). On a fleet where no side clears 2 SE on payoff the test could never speak - SP500 and XAUUSD never once printed its verdict across two days of logs - so the screen decided by default. The era verdict now certifies BOTH sides whatever the input says: the checkpoint no longer depends on a per-chart input, so one trained model serves LONG_ONLY, SHORT_ONLY and BOTH and the input can be optimized in the tester without a retrain. The blocked side's vote still closes a position under Exit_On_Reversal_Vote. PIVOT GEOMETRY. The five confirmed-pivot swing features (leg direction, distance since pivot, prior-leg magnitude, retracement ratio, bars since pivot) scored 4/2/0/2/1 of 24 on the keep-screen and the mask had already dropped them from every emitted vector. Deleted from the builder; the kept trend-position columns are swing[0..3] now and the vector is unchanged. SWING LEG SIZE, measured. MeasureSwingGeometry now prints the leg-size distribution in ATR at the leg's start pivot. Offline on the archived H4 feeds (stock ZigZag 12/5/3, ATR 20): median leg 4.5-5.0 ATR, p90 9-10, max 35-76; median length 13-14 bars, a third of legs outlast the 18-bar hold. The era report's ORACLE (+1.6 ATR) is a fixed-horizon close-to-close capture from the call bar, not the leg - a perfect caller entering at the pivot and holding 18 bars earns ~1.9-2.0 ATR on the same feeds. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
d82e7b8ed5 |
feat(training): money-weighted loss - cross-entropy could not see magnitude
The objective was never aligned with the book. Cross-entropy against the pivot label scores a 0.2 ATR pivot and a 3 ATR pivot identically, so the network optimises how OFTEN it is right and nothing tells it that being wrong on the big ones is what kills the P&L. That is exactly the measured pathology: precision 50.4-65.2% against a 13.4-14.5% base rate (23-52 sigma, unambiguous) while the book earns +0.002 to +0.045 ATR per call and loses to always-long on five of six charts. The label is NOT the problem. The oracle - a perfect caller of this exact label - earns +1.167 ATR @5 and +1.627 @19 over 3,986 calls. The target is rich and we capture ~2% of it. Per-sample weight = |forward return at the hold horizon| / mean(|r|), clamped to [0.25, 4.00]. Normalised by the mean, so an era trains on the same TOTAL gradient mass as before - redistributed from cheap bars to expensive ones - and the effective learning rate does not move. - Variables/MoneyWeight.mqh: the knob (MONEY_WEIGHT_VERSION 0 restores unweighted training and the pre-change fingerprint exactly), the transform, the fingerprint tag. Runtime predicate, not #if - MQL5 has no #if <expression>. - The floor is load-bearing: sampleWeight = 0 IS NOT A SKIP in this optimiser (Adam momentum and weight decay still apply, t and m_batchCount still advance), so weights must never decay toward zero. - Cached as the RAW |r| under the label's own validity flag, normalised at use: the mean keeps moving as bars resolve, so a cached weight would be stale for every bar but the last. - Peer pool rows stay at 1.0 - the pool record carries no forward return and deliberately does not name its source bar. Normalisation is what keeps the local:peer gradient ratio unchanged. - Not lookahead: measured from future bars exactly like the label, reaches the loss only, and has no route into the feature vector. - RETRAIN-FORCING by design. A weighted and an unweighted model have identical topology and identical weight-file shape and differ only in what they were taught to value; the fingerprint is the only thing that can tell them apart. - Reports the realised weight distribution once per cache, including the share pinned at each clamp - a saturated scheme is invisible in every downstream number. Compiled 0 errors, 0 warnings. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
724f8dce83 |
feat(features): fleet keep mask - 13 of 49 columns, the false-pivot lever
RETRAIN-FORCING. Re-keys every fingerprint (column count 49 -> 13, plus an explicit |FMASK: token). Decomposing the deployed fleet's payoff by distance-to-pivot says the whole edge is one number. A hit pays +0.94 to +1.57 ATR and is FLAT across d=-3..0; a miss costs -0.76 to -1.62; the miss rate is 34.5% to 49.4%. Remove the misses and every chart earns +0.97 to +1.22 instead of +0.05 to +0.32. One percentage point off that rate is worth +0.011 to +0.018 ATR per call. Two obvious levers are already spent, and the logs say so. Raising the vote threshold buys precision and hands the expectancy back - SP500's precision climbs 38.9% -> 57.3% across the six rungs while its book expectancy FALLS +0.39 -> +0.21 - and the rung search already takes the best rung available. Exits do not rescue it either: the 64-cell sweep fails Sidak on all six, p 0.19 to 1.00. Narrowing the label to drop d=+1 was checked and rejected: it is negative on SP500/USDCAD/XTIUSD, neutral on XAUUSD/USDJPY and strongly POSITIVE on EURUSD (+0.769, n=116). What is left is discrimination at fixed coverage, and 294 inputs against ~4,700 independent observations is 16 per first-layer weight - the overfit regime, where an out-of-sample overfit IS a false pivot. The keep-screen already measured which columns carry information, and the open question blocking a prune was whether the masks agree. They do. Intersected across all 24 members - six charts x four architectures - swing[5..8] and ma[0..1] are UNANIMOUS, ma[2..4] carry 23/23/20 votes, and 16 of 49 columns are kept by nobody. This mask is every column kept by at least half the fleet: swing[5..8], ma[0..4], crossasset[3,4], volume[1,3]. Width 294 -> 78, first-layer budget 16 -> ~60. Two of the votes are worth reading twice. swing[0..4] - the five CONFIRMED-PIVOT features - scored 4, 2, 0, 2 and 1 of 24: the label is a ZigZag pivot event and the ZigZag geometry carries almost nothing, while trend position carries everything. And alt[0..11] scored 4, 1, 1, 1 with atr at zero, so a whole rate-limited external pipeline buys nothing under this label. Alt stays ENABLED and merely unmasked, so the screen keeps reporting and the finding stays falsifiable. Implementation keeps one authority for block order and width: CFeatureBuilder::FeatureBlockTable, which FeatureSlotName, the kept count and the emit-time compaction all read. The mask is stated BLOCK-RELATIVE so toggling a block cannot silently re-point it, and applied once per bar at the existing sanitize seam rather than inside nine emitting blocks. m_neuronsCount is set from the counted kept columns, never from a hand-written subtraction, and a one-shot width check fires if the two ever disagree. FEATURE_MASK_VERSION 0 restores the full set exactly, including the pre-mask fingerprint. This is a measurement, not a conclusion: the screen tests each column's MARGINAL information and cannot see a column that is useless alone and useful in combination, so judge the retrain against the per-chart false-pivot rate and book expectancy above. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
2c2cdfed6a |
feat(exits): sweep the barrier grid on held-out rows; remove SL_Mode/TP_Mode
PHASE 1 of doing the trade-management search inside the EA instead of in the MT5 optimizer, for the one pair that cannot wait: the stop and the target. WHY NOT ALGLIB. The space is discrete - 8 stop widths x 8 target widths = 64 cells. At that size you do not search, you ENUMERATE. Exhaustive has no seed, no convergence question and no tuning of its own, and it answers the thing a search cannot: whether the good region is a broad plateau or one lucky cell. WHERE IT RUNS. Inside the era report's existing OOS threshold sweep - same held-out rows, same purge, same certified rung the deploy gate uses. No tester, no agents, no .set. ORDERING IS THE WHOLE POINT. MFE and MAE cannot say which barrier a trade hit FIRST, and "both were touched" is the common case, so a grid evaluated from the extremes alone would be guesswork dressed as measurement. MeasureBarPayoff() already walks the forward window bar by bar; it now records the first-touch bar offset for each grid level in each direction. One compare per level per bar on a loop that already runs. The sweep then costs 64 integer compares per fired row and reads no prices at all. Same-bar ties resolve to the STOP. Bar data cannot order two touches inside one bar, and assuming the target would be the optimistic half of an unknowable coin flip, on exactly the half that flatters the result. JUDGED AGAINST THE NULL OF THE MAXIMUM, NOT ZERO. Picking the best of 64 cells and reporting its own z is the best-of-N error this project has already made four times, including on the deploy decision. SidakFamilyP over the cells actually scored is the same correction the rung sweep and the baseline comparator use. A cell needs BARRIER_MIN_CALLS before it may win at all - a thin cell tops the grid on noise alone. The pair is STORED either way: a reader must be able to tell "swept and rejected" from "never swept", so the p travels with the pair (WSTA in .stats) and gates its use at read time, not its record. SL_Mode and TP_Mode inputs are REMOVED, and STOP_LOSS_MODE/TAKE_PROFIT_MODE with them - deleted rather than left dangling, per the RISK_REWARD_RATIO rule: a live enum with no input behind it is the shape of the 2026-07 incident where a saved .set kept feeding a deleted ordinal back in. With the inputs gone there is no ordinal left to feed, and ValidateTradeManagementInputs() loses two members. Three tiers at read time, most trusted first: the swept pair when it cleared its own family-wise test; else the mean excursions (cruder, but nothing was SELECTED to produce them, so they need no such test); else a fixed fallback that announces itself and is unreachable on a deployed chart, since a chart with no completed era cannot pass the deploy gate. Entry offset, expiration, trailing and the exit-vote flag stay inputs for the MT5 GA - they need entry-fill and path-stepping simulation this phase does not have. Compiled 0 errors, 0 warnings. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
f942cb2ef2 |
feat(exits): size the stop and target from measured excursions; gate indicators on a new bar
Two changes, both driven by the same single-backtest profile (6,607,688 ticks, 165 s, SP500 H4 2019-2026). INDICATORS, 39% OF THE PASS. m_indicators.Refresh() cost 61.4 s at 9.30 us per tick, running on every quote regardless of Expert_EveryTick. Every indicator this EA registers is computed from CLOSED bars - ATR, the MA, the ZigZag, the feature block - so their value cannot change between two ticks of the same bar and the refresh was recomputing a constant. It is now gated on a new bar, with its own CNewBar watermark: IsNewBar() consumes the transition and Refresh() runs before Processing(), so sharing m_newBar would have silently disabled the SetDirection gate. The exit invariant holds. What changes intrabar is price and position, and neither comes from an indicator buffer: m_symbol.RefreshRates() still runs on every tick and is what CheckClose/CheckTrailingStop/pending maintenance read. A trailing stop still moves on any tick; it now compares the live price against an ATR from the last closed bar, which is the ATR it should have been using. ProtectOpenPosition() keeps its ungated copy - it runs only on ticks Refresh() declined, and never appeared in the profile. SL_MEASURED / TP_MEASURED (both value 100), now the shipped defaults. The fixed pair was SL x2 / TP x6 against a measured SP500 MFE of 2.59 ATR and MAE of 2.71: the stop sat INSIDE the average adverse move and the target BEYOND the average favourable one, so the average trade was stopped out before reaching a target it does not reach. That converts winners into losers mechanically at any precision, and no model work can fix a barrier pair pointing the wrong way. These are NOT the removed SL_INTELLIGENT/TP_INTELLIGENT. Those scaled the barriers by the model's own CONFIDENCE - an over-confident model gave itself a tighter stop and a wider target, which is why they were deleted. These read a MEASUREMENT of what the market did on the bars this ensemble fired on: the mean adverse and favourable excursions in ATR at the hold horizon, already computed by the era report that certifies the deploy and previously printed and thrown away. Only the WIDTH is measured; the reward:risk that falls out is reported, never targeted - the ratio is policy, the width is what pays. A FRESH ENUM VALUE, never the vacated -1 the removed members held: MetaTrader does not validate enum inputs, so a .set saved by that build still feeds -1 in, and reusing it would silently give a stale file a new meaning. ValidateTradeManagementInputs() keeps rejecting -1 and now accepts 100. Plumbing follows the derived-threshold route exactly, because ExpertSignalCustom.mqh is the PARENT of the filter that owns g_ensBest* and cannot read them: stashed inside the isBetter block (so the widths describe the bars the CHECKPOINT fired on, never a later era's), persisted as WST9 in .stats (a deployed ensemble runs no further eras - the tier-ladder failure one layer along), and published per tick via PublishMeasuredBarriers(). The -1/-1 "not measured" state is published too, so a reset-weights cannot leave a stop sized off a dead ensemble. With no measurement it falls back to the shipped fixed presets and says so once per run. It deliberately does not substitute a plausible number: an invented width would be indistinguishable from a measured one in every log afterwards. Compiled 0 errors, 0 warnings. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
04cc345661 |
perf(tester): split the per-tick profile so the slow bucket names itself
Three turns of reasoning about where a 6.3-minute pass goes have produced three hypotheses and no measurement. The profiler that would answer it already existed and printed nothing: across 129 optimization passes on 12 agents the "tester pass profile" line appeared ZERO times, while OnInit's output appeared on every one. Two fixes, both aimed at ending the guessing rather than at being right. 1. The line is now built by WarriorTesterProfileLine() and printed from OnTester() as well as OnDeinit(). OnTester runs on the agent at the end of the pass, BEFORE OnDeinit. If the line appears there and not in OnDeinit, an optimization agent is discarding OnDeinit's Print; if it appears in neither, g_tpTicks is genuinely 0 and the instrumentation never ran. Those need different fixes. The zero-tick case now prints its own explicit line instead of staying silent, because a profiler that says nothing when it fails is indistinguishable from a fast pass. 2. "Expert.OnTick" was one bucket containing both halves of the question. CExpertCustom::Refresh() runs on EVERY tick whatever Expert_EveryTick says - correctly, since an open position must be manageable on any quote - and under a 1-minute-OHLC model that body executes millions of times per pass. Its three steps are now timed separately: TCHasEnoughHistory(), RefreshRates(), and m_indicators.Refresh(). They are reported as a SUBSET of Expert.OnTick, not as siblings, because double-counting a bucket is how a profile lies. ProtectOpenPosition()'s copy of the same refresh counts into the same bucket rather than hiding on the declined-tick path. The accumulators move to System\TesterProfile.mqh: ExpertCustom.mqh has to see them and is included long before Warrior_EA.mq5's own globals. Every bracket is guarded on g_tpActive, which OnInit sets only for MQL_TESTER/OPTIMIZATION/FORWARD - a live chart never reads the clock for it. No behavioural change to any path. Compiled 0 errors, 0 warnings. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
a9c21fdf10 |
feat(tester): gate input RANGES at OnTesterInit, not just input values
ValidateTradeManagementInputs() hard-gates input VALUES at OnInit because MetaTrader replays a saved .set without validating it. Nothing gated input RANGES, so an optimization would happily sweep a parameter that cannot affect a trade - or one that destroys the run - and score every pass as if it meant something. Measured on the 2026-08-31 session: Signal_ThresholdOpen was swept over 10..80, which MetaTrader expands to twelve PERCENTAGE_PRESETS members. Its own label reads "SEED (derived after era 1)": every pass restores the pinned rung from .stats into g_ensDerivedThreshold, and OnTick publishes it over m_threshold_open before the first bar closes. Twelve identical strategies, a 12x multiplier on a fifteen-hour estimate, and no warning anywhere. A wasted dimension does not only multiply runtime - it fills the GA's fitness landscape with plateaus, so crossover produces clones. 115 of 129 passes scored 0. WarriorGuardOptimizationRanges() runs once in the controlling terminal, the only place ParameterGetRange/ParameterSetRange are legal and the last moment a bad sweep is free to stop. Two classes: REFUSE - keys BuildModelFingerprint(), or the capacity budget upstream of the neuron count inside it (Use_Training_Pool). These select a different MODEL, not a different strategy: the agent resolves a .nnw filename that does not exist and either votes silently or trains inside the backtest, and the pass still lands in the results table looking real. Session stopped with INIT_PARAMETERS_INCORRECT. PIN - a seed, or inert on an inference-only pass: the training-only inputs, the DB ranking pair, presentation flags, and the news filter (calendar error 4806 over historical dates, so it fails open on every bar - the message says so, because a config optimised here trades through news live). The sweep is switched off and the operator's own value kept. tradingdirection is deliberately NOT in the table - it is a legitimate override, and WarriorEffectiveDirection() is explicit that the screen never overrules an operator who chose. A degenerate range spanning BOTH and the screened side is reported instead, from the controlling terminal where the drift screen has full history. An unresolved name warns rather than refuses: the table addresses inputs by string, so a rename unbinds it, but a false refusal would make the EA impossible to optimize at all. The summary line prints on every optimization, clean ones included, so a table that has come unstuck shows up on the first run - same doctrine as the drift screen reporting on charts it does not restrict. Compiled 0 errors, 0 warnings. Live paths untouched: nothing outside OnTesterInit() is reachable from this file. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
f7adb0bf16 |
feat(panel): Pause works on a deployed chart - it holds ONLINE LEARNING, and owes the bars back
Operator asked for a queue: hold online-training signals while an optimization
runs, then push them "as if the stop never happened". The queue already exists
and is better than a queue - it is a WATERMARK. m_onlineLearnedUpToTime
advances one LEARNED bar at a time inside the catch-up walk and is never
bulk-set at the end, so a step that returns early leaves it exactly where it
was; the next walk starts at the frontier, runs back to the watermark and
learns oldest -> newest in strict chronological order. It rides in .stats, so
it survives a restart. Nothing to serialize, nothing to drain.
What was missing was the switch. OnlineLearnStep() already gates on
m_trainingPaused - but HandleCpTogglePause() REFUSED on a deployed chart ("there
is no training run to pause"), which was true of the training loop and wrong
about the chart. Online continual learning runs only on deployed charts, so the
one state where pausing matters was the one state the button would not enter.
WHY IT MATTERS FOR THE GA: online learning re-saves the deployed .nnw, and every
tester agent re-seeds its optcache when the production file is newer than its
copy. A model that keeps learning mid-run makes later passes evaluate different
weights than earlier ones and the GA reads that as a parameter effect - and the
config you optimised was tuned against a model that no longer exists by the time
you deploy it. Pausing freezes the weights for the run and gives the bars back
afterwards.
Bounded by ONLINE_LEARN_MAX_CATCHUP (64 confirmed bars, ~10 days on H4) - the
walk stops there and older bars stay behind the watermark. Covers an
optimization run; not a way to park a chart for a month. Stated in the alert.
Trading is unaffected: nothing in the live signal path reads m_trainingPaused;
the only other deployed-state reader is the one-shot ladder/snapshot rebuild,
which defers. The alert names which thing was paused, because "training paused"
on a deployed chart reads as pausing something that was not running.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
eed262c416 |
fix(chart): the vote-arrow file stamped the threshold from INIT, not the one in force
An arrow is a claim about a threshold, so the store discards its cache when that threshold changes between sessions. Correct rule, stale input: the stamp was captured once in Configure() during OnInit and never refreshed, while the derived rung goes on MOVING for the rest of the session (it follows each era's rung until a checkpoint pins it). So the two files written in the same OnDeinit disagreed by construction: .stats saved the CURRENT threshold, .votearrows saved the init-time one. Next start compared them and threw away a perfectly good arrow set. Measured 2026-08-27 22:26 - four of six charts lost everything: XAUUSD stored 25 / now 20 XTIUSD stored 25 / now 5 SP500 stored 25 / now 5 USDJPY stored 15 / now 20 The two that survived, EURUSD (173 arrows) and USDCAD (13), were the two whose rung happened not to move. WarriorVoteArrowThreshold() is now the one expression, read live at the load comparison AND at every save; the cached member is gone, so it cannot go stale again. Only the close threshold is still passed to Configure(), because VOTE_EXIT_DISABLED_THRESHOLD genuinely cannot move. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
5304cfb876 |
refactor: delete the machinery I added that was not earning its place
Self-audit against YAGNI. Three things from this session's commits were
configuration nobody asked for, and one was a measurement report pasted into
source.
DELETED - SignalCooldownScopeOverride + WarriorSignalCooldownScope(). I added a
source override for the cooldown scope, then concluded in the very next commit
that source-overrides-operator is the wrong pattern and stood it down to -1.
What was left was a mechanism whose only state is "disabled" - the definition of
speculative generality. The scope is read straight from its input now. The thing
that actually solves "which value is running" is the init line that PRINTS the
resolved value, not a second place to set it.
DELETED - REQUIRE_BOTH_SIDES_PAY. A #define that was always true, added so the
new gate could be "reverted". Nobody asked for that switch and a gate condition
that is optional is not a gate condition. The rule is either right or it is not;
if it turns out wrong, git is the revert mechanism.
TRIMMED - 38 lines of comment in the ensemble verdict listing all six charts'
by-side payoff numbers. Those are a MEASUREMENT: they belong in the commit that
made the change and in the session record, not pinned in source where they go
stale on the next run and start actively misinforming. Two lines left saying
what the code does and when it can be false.
TRIMMED - the cooldown override comment from 11 lines to 4. Same reasoning.
KEPT, with the case for each:
CSignalDeclusterPolicy replaced FOUR copies of one rule that had silently
diverged into two different units and two different
windows. Net negative lines.
PivotLabelFillTarget replaced FIVE hand-rolled target vectors and is the
only place the sum-to-1.0 gradient invariant lives.
PivotLabel* window fns the -1 sentinel collision they fix was a live defect.
CRunningMean two meanings were sharing one accumulator, which is
how the purge came to be sized off the wrong one.
LABEL_RESOLUTION_CAP_BARS a bound on a MEASURED quantity feeding the training
split. The floor is load-bearing; the cap stops a
degenerate ZigZag eating the training set.
Test_PivotLabelWindow 161 of the 195 net added lines. Two of the three
things it pins were live bugs this session.
Net across the whole session, production code only, excluding tests:
+479 / -390 = +89 lines, and in exchange: decluster rule 4 copies -> 1,
target builders 5 -> 1, chart reconciliation 2 -> 1, twelve loose NMS members
-> two policy objects.
Compiles clean, 0 errors 0 warnings.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
eac8912dd0 |
fix(gate,cooldown): both sides must PAY, not just fire; stand down the source overrides
Three fixes off the 2026-08-27 session logs (six H4 charts, fresh from era 0).
1. THE DEPLOY GATE NEVER READ ITS OWN DRIFT-FREE TEST.
twoSided was `firedLong > 0 && firedShort > 0` - an anti-degeneracy check.
The by-side PAYOFF prints three lines below it under the heading "THE TEST
THAT IS DRIFT-FREE", and nothing consumed it. Measured at each chart's own
derived rung, hold horizon, ATR per call:
EURUSD long +0.073 short +0.115 <- the only one where both pay
USDJPY long +0.358 short -0.136
USDCAD long +0.177 short -0.070
XAUUSD long +0.076 short -0.379
XTIUSD long +0.222 short -0.065
SP500 long +0.657 short -0.273
Five of six were stamped DEPLOYABLE on 30-37% precision (chance 13-14%)
while their short book lost money on every call. XAUUSD's whole vote earned
-0.157 against its own always-long book at +0.436 and still passed.
Precision is measured against the LABEL: it says a turn was called, not that
the leg after it paid. On a trending instrument the down-legs are shorter, so
a symmetric caller with real skill still bleeds on one side - and the higher
rungs make it worse, because more member agreement means more agreement WITH
the drift (SP500 at the 15% rung fires short n=0).
tradeableOK now ANDs in bothSidesPay. A floor at zero, not a significance
test - no per-side variance is tracked, so "positive" is the honest claim.
A side that never fired is not judged; an era where neither side has a
forward window is not judged either. REQUIRE_BOTH_SIDES_PAY restores the old
behaviour. The verdict line now NAMES this when it withholds DEPLOYABLE,
because "not good enough yet" and "one side of the book loses money" are
different problems with different fixes.
Timing is deliberate: ENSEMBLE_CHECKPOINT_MIN_ERA is 20 and the fastest chart
is at era 17, so nothing has deployed yet and this costs nothing to land now.
2. THE SOURCE OVERRIDE WAS DISCARDING THE OPERATOR'S COOLDOWN.
SignalCooldownOverrideBars was 30. The charts are configured for 10. The
session log says "signal cooldown (30 bars)" 43 times across all six. That is
the same failure the override exists to fix, pointed the other way - source
defeating the operator instead of a stale profile defeating source - and it
is WORSE, because a stale profile is at least visible in the chart's own
inputs dialog and this is not. Stood down to 0, which is the exit condition
its own comment always described.
The scope override (added earlier this session) is stood down to -1 for
exactly the same reason rather than kept out of convenience.
The real fix for "which value is running" is not a second place to set it:
ConfigureAISignal now prints the RESOLVED window and scope once per chart
alongside the chart input and both overrides, and says so explicitly when
they disagree. One line kills the whole class.
3. The overlay sweep's hard-coded OVERLAY_NMS_WINDOW of 6 is gone with the
window/decluster work replayed onto main - it now uses the resolved cooldown
like every other layer.
MEASURED, and it corrects a number I reported earlier in the session: against
the 30 bars the override was forcing, 89.8% of consecutive vote-arrow pairs sat
inside the window. Against the 10 bars actually configured, the figure is 11.7%
(1038 pairs across six .votearrows sidecars, gaps counted in real H4 bars with
weekends excluded). The earlier figure was true of the window that was running
and NOT of the one the operator set; the 10-bar number is the one to judge the
fix against. USDJPY carries most of what remains (18.4%, and all 21 fleet pairs
below 7 bars) - it is also the least-trained chart, at era 3.
NOT CHANGED, and flagged rather than fixed: the source default is still SCB_30
and its comment argues 20 is the smallest defensible value, because a trade on
this label is held 5 + the median leg = 18-19 bars. A 10-bar cooldown re-announces
inside that hold. That is the operator's call, not a bug.
Compiles clean: EA and both test suites, 0 errors 0 warnings.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
184de331fc |
fix(label,signal): move the pivot window into the leg; one declustering rule, measured in bars
Two independent defects, both about a rule written down more than once.
THE LABEL WINDOW SAT ENTIRELY ON THE APPROACH. It labelled d=1..5 bars BEFORE
the turn. Measured payoff by distance, fleet-pooled over 102 era rows and 3
charts at the hold horizon: d=1 +2.095, d=2 +1.743, d=3 +1.300, d=4 +0.969,
d=5 +0.769 ATR/call. Slope -0.343 per bar, monotone on all three individually.
Pure hold truncation predicts d=5 earns 14/18 = 0.78x of d=1 (~1.63); it earns
0.769, which is 0.37x - so truncation explains less than half and the rest is
the ADVERSE APPROACH, the call sitting through bars of price moving against it.
Every bar of lead the window granted was paying that.
The window now straddles the turn: d=+1..-3 (LEAD 1, LAG 3). It keeps the
best-paying approach bar and spends the rest of its width on the far side,
where direction is confirmed. WIDTH IS UNCHANGED AT 5, deliberately - it sets
the class balance, and ApplyLogitAdjustment's tau is capped at 1.2 logits with
the cap binding at EVERY imbalance, so only the uncorrected residual moves
(1.9x at width 5). Simulated over three leg-length regimes the balance shifts
by <=0.2pp, so nothing downstream of it had to move. That is also why the
target stays FLAT across the window rather than Gaussian-tapered: a taper
shrinks directional mass ~40% and tau has no budget left to absorb it.
Finality is re-derived, not weakened. One committed pivot strictly newer than
the window's newest slot freezes the whole window - the labelling pivot's
finality and the negative verdict together - because ZigZag can only ever move
its newest pivot, and only toward newer bars. Verified: the witness is the next
pivot newer than the labelling pivot in 100% of simulated cases, so minMove
still measures the leg AFTER the turn.
TGT:PVT1 -> TGT:PVT2:lead:lag. Every .nnw re-keys; the fleet retrains from era 0.
-1 IS NO LONGER A VALID SENTINEL - it is one bar into the leg. Added
PIVOT_LABEL_NO_PIVOT and PivotLabelDistanceSlot(); the by-distance payoff report
bucketed the no-pivot population and the leg bars into the same slot otherwise,
and its "directional labels only" test (`d >= 0`) would have dropped every leg
bar it now exists to measure.
THE DECLUSTERING RULE WAS WRITTEN OUT FOUR TIMES and the copies disagreed.
Rules 0-3 lived in NmsLiveAccept, PruneDirectionalClusters, pass 3's OOS tally
and the overlay sweep, each carrying a comment insisting it must match the
others. It did not:
* the two sweeps compared bar INDICES; the live gates compared wall-clock
seconds. Elapsed time is always >= bars*period because weekends and session
breaks add time without adding bars, so the live rule was strictly the most
permissive and could only ever UNDER-suppress. On H1 a 30-bar window is 30
hours against a ~50 hour FX weekend: a Friday signal never blocked a Monday
one, once a week, per chart. On session-break instruments, daily.
* the overlay sweep had rules 1 and 2 only - no cooldown, no alternation -
against a hard-coded OVERLAY_NMS_WINDOW of 6 while the live gate required
30, and its comment claimed parity with m_signalClusterWindow. It drew
reconstructed arrows 7 bars apart onto a chart whose gate requires 30.
Now one CSignalDeclusterPolicy, measured in bars via WarriorBarsBetween(), with
an unresolvable frame SUPPRESSING rather than passing. The vote layer configures
it as a pure cooldown because that is what CheckOpenPosition actually applies.
Three more that let clusters through:
* THE CLOCK WAS NEVER SEEDED. MT5 restores arrows from the .chr profile but
not the state that spaced them, so every restart, recompile or timeframe
change let the next bar fire regardless. Seeded from the newest arrow on the
chart once the progressive restore completes.
* THE COOLDOWN WAS CONSUMED BEFORE THE TRADE EXISTED. VoteCooldownAccept both
tested and committed, ahead of order-parameter validation - and the failure
branch restores the vote for retry but could not un-burn 30 bars. Split into
a pure test plus VoteCooldownCommit() at the draw.
* SCOPE HAD NO SOURCE OVERRIDE. Bars did; scope did not, so a chart whose
profile held PER_DIRECTION silently dropped rule 0 with no way to correct it.
Also DRY: five hand-rolled 3-class target vectors -> PivotLabelFillTarget(), the
one place the slot order, the smoothing constants and the sum-to-1.0 invariant
live (the gradient is target_i - softmax_i, so a vector that does not sum to 1
adds a constant drift to all three logits). Two copies of the chart-wide
reconciliation -> WarriorReconcileVoteCooldown(). Twelve loose NMS members ->
two policy objects.
Documented at CNet::backProp: sampleWeight = 0 IS NOT A SKIP. Adam still applies
decayed momentum and decoupled weight decay, t still advances, and m_batchCount++
is unconditional - so masking a class-skewed subset that way shrinks every weight
rather than ignoring the example. Drop the bar before queueing instead.
Compiles clean (0 errors, 0 warnings). Window arithmetic verified by simulation;
the label's in-situ behaviour against real ZigZag output is NOT verified here.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
213b3aac15 |
fix(chart): the reconciliation never ran - it was hooked to a sweep deployed charts do not do
cooldown-recon put the chart-wide cooldown at the end of the overlay sweep. It executed ZERO times. This store's own header already said why: the sweep 're-arms only when an era ends. A DEPLOYED ensemble runs no further eras'. Five of six charts were deployed, so there were zero 'Filtered view: swept' lines in the entire session while the saved files still held 148 same-side pairs under 30 bars on XTIUSD. Moved to the completion of the progressive vote-arrow restore, which runs on every chart including deployed ones. The restore thinning alone was never going to be enough either: MT5 persists chart objects in profiles\Charts\*\chart*.chr, so arrows drawn under an older window are ALREADY on the chart when the process starts, and a freshly-thinned restore just adds to them. Two correctly-thinned sets still union into clusters. The chart is the only authority. Same construction as before: OBJ_TREND only (the line is the canonical half of a mark, matching Snapshot()), sorted by time first because object order is not time order, and the gap>0 guard so a mis-ordered set fails visibly by keeping rather than silently by deleting. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
4e2cdd3aff |
fix(chart): reconcile the cooldown over the CHART, not over one producer's record
The record-based prune was individually correct and still left clusters. It is not the only producer of a vote arrow: the persisted-arrow RESTORE thins its own list from its own state, the overlay sweep thins its own list from its own state, and the live gate marks the current bar from a third. Each spaces ITS OWN survivors 30 bars apart; interleaved on one chart the union sits 1 bar apart. Two independent thinning passes over one namespace produce a union, not an intersection. MEASURED from the saved .votearrows files, which is what the chart actually holds: XTIUSD 284 arrows, 180 gaps under 30 bars, 148 of them SAME-SIDE, min gap 0 XAUUSD 263 arrows, 177 gaps under 30 bars, 144 same-side, min gap 1 EURUSD 309 arrows, 176 gaps under 30 bars, 110 same-side, min gap 0 while every producer's own log reported it had thinned correctly. The sweep's 'drew' counter says what ONE producer drew; the chart is the union. Verifying on that counter is what let this stand through four builds. The authority is now the chart itself: after the sweep, walk every SIG_VOTE_PREFIX OBJ_TREND object, sort by time, enforce one window. Whatever drew an arrow, this runs last. OBJ_TREND only - a mark is a line AND an arrow and the line is canonical, the same test Snapshot() uses. Sorted first, because object order is not time order and an unsorted forward walk yields negative gaps, which is how a prune once deleted 272 of 273 arrows. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
844aac653a |
fix(train): the OOS final pass ran a full epoch at an undecayed rate
Capping the pass at m_etaCeiling was not enough. Measured on the first two live runs: USDCAD 0.00085 over 13,335 bars, EURUSD 0.00242 over 15,041 - a 3x spread across charts, because a chart whose plateau ladder reset recently still carries a high eta and the cap never bound. The slice turns out to be roughly HALF the data, not a tail, so one pass over it at the model's own rate is a full training epoch on a model that has already been selected and certified. That is materially more than the 'just a bit finer weights' this was asked for. OOS_FINAL_PASS_ETA_SCALE (0.25) now scales the rate. Scaling rather than shortening the pass keeps the whole slice in play - seeing the held-out bars at all is the point - while making the step proportionate to an already-selected model. USDCAD and EURUSD have already taken the unscaled pass; that is not reversible without a retrain. USDJPY has not converged yet and will get the corrected one. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
c5b9a1ad10 |
fix(signal): a changed input default cannot reach an already-attached EA
Raising Signal_CooldownBars from 10 to 30 changed nothing. All six live charts
kept reporting a 10-bar window, because MT5 stores an input PER CHART in
profiles\Charts\*\chart*.chr and an already-attached EA ignores a changed
default entirely. This codebase already documents that trap, in the derived-
threshold comment in Training.mqh - and converting SignalClusterWindow from a
const to an input reintroduced the exact problem the const existed to avoid.
SignalCooldownOverrideBars (const, 30) now wins over the input; 0 hands control
back to the panel. Tunability per chart is kept, source-correctability is back.
Not applied when the input says OFF: an operator who switched the cooldown off
meant it, and silently re-enabling it from source would be the same surprise
pointed the other way.
ALSO gates the per-model arrow restore on DrawUnfilteredSignals. DrawObject()
returns early when the raw view is off, but AdvanceChartSignalRestore called
WarriorPlotSignalLevel DIRECTLY and never checked - so every restart repainted up
to MAX_PERSISTED_ARROWS per-model opinions per member, four members per chart, on
top of the combined-vote arrows. Same shape as the vote-arrow restore bug in
|
||
|
|
09af7d5ee9 |
feat(signal): default the cooldown to 30 bars - 10 thinned almost nothing
Measured on the live fleet at 10 bars: the restore thinned 100 of 364 and the overlay prune 23-66 per chart, leaving 182-250 arrows over ~5000 bars. Every layer was working; the window was simply below the ~20-bar natural spacing between vote arrows, so it could only catch the tightest pairs. The floor is principled, not cosmetic. A trade on this label is held for 5 + the median ZigZag leg = 18-19 bars, so any second signal inside that window is the same trade being re-announced. 20 is the smallest defensible value and 30 is one step above it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
3218db4a38 |
feat(train): ONE pass over the held-out slice at deploy, on the restored checkpoint
The OOS slice is the newest history and the model never trains on it, while online learning adapts to every bar resolving AFTER deployment. That leaves a gap exactly at the handover, over the most regime-relevant data there is. This closes it: select on validation, then refit on everything, which is standard practice. Placed AFTER Net.RestoreWeights() and ResetOptimizerState() and BEFORE PersistDeployedModel(), so it refines the weights that were actually SELECTED rather than whatever the run happened to end on, and what it produces is what gets written down. THE COST IS REAL AND IS NOW STATED IN THE LOG. The deploy line promises "every model reverts to the weights it held at the era whose combined vote scored best, so the ensemble that trades is exactly the one that was measured". After this pass that is no longer literally true, so the pass prints that the certified numbers belong to the PRE-PASS weights and must be quoted that way. Set EnableOosFinalPass=false to keep certified == traded exactly. Guards: * ONE-SHOT PER RUN, and the flag is set BEFORE the loop so no early return inside it can leave the pass eligible to fire twice over bars it already trained on. Reset at m_trainRunActive=true, because a retrain is a fresh selection and earns a fresh pass. * THE CONVERGED RATE, never a plateau-boosted one: m_modelEta can still carry PLATEAU_RESTART_BOOST from an escape attempt, and this is a refinement of a selected model, not another warm restart. g_eta is what backProp reads, so that is what is capped and restored. * OLDEST -> NEWEST. Series indices count backwards, so decreasing i moves forward in time - the order the bars happened in. * A failed feedForward is never followed by backProp; the output layer would still hold the previous sample's activations and the update would be this bar's label against another bar's prediction. * m_oosFinalPassCutoff records the newest bar consumed and is deliberately NOT cleared on a new run, so a later run can say plainly that its out-of-sample window reaches back into bars this model has already seen. Expect the gain to come from CURRENCY rather than finer weights: OOS precision was measured flat from era 20 while in-sample error kept falling, so the data this model can already see is exhausted. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
322c052a65 |
fix(chart): the persisted vote-arrow restore put back every arrow the cooldown removed
Clutter remained after cooldown-v3 because there is a FOURTH producer of
SIG_VOTE_PREFIX arrows: CVoteArrowStore, which replays a .votearrows file and
draws into the SAME object names as the overlay. Replaying a file written before
the cooldown existed therefore resurrects exactly the arrows the prune deleted.
On a DEPLOYED chart that is the entire arrow set. The store's own header says
why: the overlay re-sweep "re-arms only when an era ends. A DEPLOYED ensemble
runs no further eras" - which is the reason this store exists at all, and also
the reason nothing would ever have removed those arrows again.
The restore now thins to the cooldown at Load(), before the progressive draw is
armed, so it is idempotent: a file already written from a cooled chart passes
through untouched, an older one is corrected once.
IT SORTS BY TIME FIRST, AND THAT IS NOT OPTIONAL. Snapshot() walks
ObjectsTotal(), so the record is in OBJECT order - its own comment says so, and
the existing MAX_KEPT trim already sorts a copy for exactly this reason. Applying
a spacing rule to an unsorted record yields negative gaps, and a negative gap is
inside any window: that is the bug that wiped 272 of 273 arrows in
|
||
|
|
2ca32e933f |
fix(signal): the overlay cooldown prune ran backwards and left ONE arrow per chart
cooldown-v2 suppressed 272 of 273 on SP500, 320 of 321 on EURUSD, 329 of 330 on USDCAD - one surviving arrow on every chart in the fleet. The record is OLDEST-FIRST. The prune walked it backwards, so every gap came out NEGATIVE, and a negative gap is always <= the window: everything after the first arrow was suppressed. The direction was taken from the member comment on m_overlayIndex, which reads "walking newest -> oldest" and is WRONG. The sweep DECREMENTS a SERIES index (0 = newest) from MathMin(span, barsAvail-150) down to m_overlayStopIndex, so it walks OLDEST -> NEWEST. The pre-existing overlay NMS at the draw site agrees - it tests (m_overlayNmsKeptIdx - idx) and expects that to be positive for later bars. A stale comment counts as a guess, and this one cost a build. Comment corrected at the declaration so the next reader is not misled the same way. Guard added: the gap must be > 0 as well as <= the window. A non-positive gap means the record is not in the order this loop assumes, and suppressing the whole chart is precisely what that looks like from the outside - so it now fails visibly by KEEPING rather than silently by deleting. Found only because the verification was the drawn arrow count rather than an assertion that the code was correct. Compiling clean said nothing about it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
110080dc86 |
fix(signal): the cooldown belonged at the VOTE layer, as a filter - not per member
cooldown-v1 extended NmsLiveAccept, which declusters each MEMBER's own signal. That is not what the charts show and not what trades. The combined vote in CExpertSignalCustom had NO spacing rule at all - grep found not one reference to the cluster window in that file - so four individually-declustered members were averaged into a vote that could fire on consecutive bars. Measured live: 2,970 voting bars becoming 299-328 vote arrows. Proof of the diagnosis, from the deployed fleet under cooldown-v1: SP500 273 and XAUUSD 212 arrows, unchanged from before the change. The member-level rule could not touch them. Gated where the vote becomes a trade - CheckOpenPosition, beside the open-prohibition and open-market-closed checks, tracing as "open-cooldown". That is the filter chain the request asked for from the start and it is where this should have gone first. Suppression there means no order AND no live arrow, honouring the same "no arrow, no vote, no position" contract the member rule already had. THE DRAWN HISTORY NEEDED A SECOND PASS, NOT AN INLINE TEST. The overlay sweep walks NEWEST->OLDEST and is chunked across ticks, so an inline cooldown would keep the NEWEST bar of a cluster while the live gate keeps the FIRST, and the drawn set would contradict the traded set - the exact defect the renderer's own comments warn about. The sweep now records what it drew and prunes it backwards over that record, which is forward in time. Direction() is a TRANSACTION that can run more than once on a bar, so the live accept is cached per bar time. Without that a second call flips the bar's verdict after it has already journaled one. One resolver, WarriorSignalCooldownBars(), now serves both layers so they can never disagree about the window. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
6308a19f27 |
feat(signal): make the signal cooldown tunable, and add a hard any-direction gate
The declustering the charts needed already existed - NmsLiveAccept, per-direction run-collapse plus cross-direction resolution plus strict alternation - and it was already set to 10 bars. It could not be TUNED: SignalClusterWindow was a compile- time const, so finding the right value needed a rebuild. That is the actual gap. Now three inputs, as enum dropdowns: Signal_CooldownScope per-direction, or a hard any-direction gate on top Signal_CooldownBars SCB_OFF..SCB_50, default 10 Signal_CooldownMinutes SCM_OFF..SCM_1440, overrides bars when set Minutes resolve against the CHART period and round UP, so a cooldown asked for in wall-clock is never silently shorter than requested and survives a timeframe change. SCB_/SCM_ prefixes are deliberately unique. M15/M30/M60 are ALREADY members of NF_LOOKBACK_PRESETS, and MQL5 binds a duplicated enum member to the first-declared enum silently - the obvious names would have compiled straight into the news filter's values. THE ANY-DIRECTION GATE IS ADDITIVE, NOT A REPLACEMENT, and the first cut of this had it backwards. Measured on the live log: the current rules draw 222 arrows over 4999 bars, while a BARE 10-bar cooldown permits up to 454 - because ALTERNATION is what declutters today, not the window. Swapping the rules out would have roughly doubled the clutter it was asked to remove. Layered, it can only ever suppress more. Suppressed bars still advance the per-direction last-SEEN cursors, so a run straddling the boundary does not restart as if it were fresh. Applied at all THREE sites that must agree - live inference, OOS pass-3 scoring and the chart renderer. Their own comments say why: an arrow set that does not obey the same rule as the traded set shows calls the EA would never take. Also corrects a stale comment that called this window "display only". It is not: when it suppresses, the live path zeroes the signal outright - no arrow, no vote, no position. Training never sees it, so these cost no retrain and are correctly absent from the fingerprint. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
8083a31754 |
diag(gate): move the conviction curve to the horizon that has value, and add mean-d per rung
The 5-bar conviction curve cannot answer the question it was built for. The oracle measures ~0 at 5 bars across three charts (+0.012, -0.054, +0.064), so PERFECT foresight earns nothing there and no rung can show payoff either. Every reading it produced was null by construction. It was placed at 5 bars for statistical power, before the oracle showed what that horizon is worth. Kept as a control; the hold-horizon curve is the one to read. Also adds MEAN DISTANCE-TO-PIVOT PER RUNG, which is the high-power form of the same question. Payoff falls ~0.34 ATR for every bar of distance to the pivot (fleet-pooled: d=1 +2.095, d=2 +1.743, d=3 +1.300, d=4 +0.969, d=5 +0.769, wrong calls -0.668). So a rung that selects NEARER pivots is worth more per call even at unchanged precision - and mean-d is a far tighter statistic than mean-payoff, because d spans five bars where payoff spans several ATR. That matters because it can REOPEN a lever I closed. Precision does not rise with the rung - every 15-vs-10 comparison across six charts sits below 0.71 sigma - so the threshold looked exhausted. But precision is not the only thing a threshold can select for. If conviction correlates with proximity to the pivot, raising it buys payoff without buying precision. Directional labels only: an incorrect call has no pivot and therefore no distance, and folding those in as zero would read as "this rung picks pivots that are imminent" when it means the opposite. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
0d9320cc87 |
diag(gate): the ORACLE - what a perfect caller of this label would earn
The ceiling on the target, and the measurement that decides where the work goes. Same payoff arithmetic, signed by the LABEL's direction instead of the vote's, over every directionally-labelled shared bar. If a model that got EVERY pivot right still earns nothing over the holding horizon then the target carries no money and no amount of model improvement reaches any - the label, not the network, is what has to change. If the oracle earns well the target is sound and the shortfall is the model's. Those are completely different programmes and nothing so far distinguishes them. It uses no forecast, so it is not a leak: it is the value of perfect foresight OF THIS LABEL, reported as a benchmark. Nothing may trade on it. Accumulated above the voter and direction-policy filters, like the zero-skill book, because it is a property of the bars and their labels rather than of what the vote did with them. A bar with no directional label offers a perfect caller nothing to take and is skipped rather than counted as zero - the benchmark is "every call it COULD make". Motivated by the first skill-by-distance row, which already reframes the day: correct calls earn +0.75 to +1.90 ATR against a spread of 0.005-0.042, and incorrect ones cost -0.66. That puts break-even precision near 32% against a measured 33-37% - thin, but on the right side, and utterly unlike the "no payoff" reading the confounded 5-bar window suggested. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
1a9b56e3b0 |
diag(label): expose bars-to-pivot - the confound the payoff test was missing
CORRECTION to what the payoff instrument was measuring. The 5-bar horizon looked like the powered test and it is confounded. SwingPivotDirectionLabel returns Buy when a swing LOW lands up to PIVOT_LABEL_TOLERANCE_BARS bars AHEAD, and says the quiet part itself: gating on where the pivot sits relative to entry "would drop exactly the bars where the turn has not finished coming to us", and how much adverse move remains before the turn "is a trade-management question". So on a CORRECT Buy call price is often still falling for d more bars. A window shorter than d measures the APPROACH, not the leg, and its negative contribution is expected on the calls that are RIGHT. The tight null at 5 bars (-0.012 +/- 0.074) is therefore not evidence of no payoff. Neither horizon is both clean and powered: 5 bars is powered and confounded, 18-19 is clean and has an SE of 0.277. (idx - P1) was computed in the label and thrown away. Now cached beside m_labelResolveAge under the same validity flag, and bucketed in the era verdict. DELIBERATELY NOT USED AS A PER-CALL HORIZON, which is the trap sitting right next to this: d exists only on bars the label found a pivot for, so a horizon that varied with d would hand correct and incorrect calls different windows and bias the comparison outright. The horizon stays fixed; d only buckets. The bucket for "the label called no pivot here" is reported by name rather than folded in, because it is the control the others are read against. Buckets 1..N condition on the label, so they describe the MECHANISM, not what a book earns. Reads: rising with d means the edge is in EARLY calls and the tolerance window is spending it - fixable by reweighting the loss, not by a new label. Flat means that hypothesis dies. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
3372b82dfa |
diag(gate): the conviction curve - does payoff rise with vote magnitude?
The practical question behind "can I just trade the strongest signals" is whether payoff rises with vote magnitude. The threshold sweep already visits every rung, so the whole curve costs four arrays and no extra pass. Reported as the DRIFT-FREE statistic per rung - long plus short, both sign corrected - with the two halves alongside. The halves alone invite reading a drift-fed long side as skill, which is exactly the error the zero-skill book caught at the certified rung: an always-long book earns MORE than the vote on two of three charts. Taken at the SHORT horizon, which is the one with the power. Pooled across the three training charts the certified rung reads -0.012 +/- 0.074 ATR - a tight null, 95% interval [-0.16, +0.13], with the long/short pattern (+0.030 against -0.041) being the drift signature exactly. The hold horizon agrees and is 3.7x noisier, so the answer is not a horizon artifact. Precision is already known not to rise significantly with the rung (every 15-vs-10 comparison across six charts sits below 0.71 sigma). If payoff rises anyway that is a surprise worth having; if it does not, the two agree and the threshold lever is closed on both counts. Still gates nothing. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
feaadd80a2 |
diag(gate): split the payoff by side at the horizon that can actually resolve it
The by-side test is the one that separates directional skill from drift, but at
the HOLD horizon it cannot answer: payoff overlap is the horizon itself, so an
18-bar window leaves ~65 independent observations per chart and a standard error
of 0.25-0.45 ATR against an effect that would matter at 0.1.
The 5-bar window carries ~3.8x the independent observations and roughly half the
standard error. It buys that power by risking a window that ends before the
pivot has committed - which is exactly why the horizon was widened in
|
||
|
|
ce4f74fe2c |
diag(ensemble): measure how much the four members actually disagree
The ensemble beats its best single member by +2.2 to +6.8pp on all six charts - sign-stable across six instruments, so the ensemble is doing real work rather than diluting. How much MORE is available depends entirely on how decorrelated the members are: the variance of an m-member average scales as (1+(m-1)r)/m, so at r=0.8 four models are worth about 1.2 independent ones and at r=0.3 nearly 3. Nothing measured that, so the obvious next lever - different feature subsets per member, or a fifth architecture - could not be costed. Both force a full retrain of 24 models, which is not a price to pay on a guess. Measured on the SIGNED VOTE, which is what actually gets averaged: not accuracy, not raw confidence. Two members can agree on direction almost always and still contribute independently through magnitude. Accumulated over every SHARED row rather than fired ones - restricting to fired rows would measure agreement only where the members already agreed enough to fire, which is the sample most biased toward agreement. A member whose signed vote never varies (all abstentions, a dead tier) is SKIPPED rather than counted as r=0, which would drag the mean toward "decorrelated" using a member carrying no information at all. Reported as an effective member count, which is the honest way to say what four models are worth. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
7500e08e17 |
feat(gate): the payoff number needed a zero-skill book and a by-side split
payoff-v1 reported what a call was worth and nothing to compare it against. A
positive mean R is not a finding on its own: if the instrument drifts, an
ALWAYS-LONG book earns a positive mean too, and drift is the one anomaly family
this project has found that survives cost - so the vote would be reporting the
market's own move as if it were its own.
Two comparisons, and the second is the one that decides it:
ZERO-SKILL BOOK - the same forward move accumulated with a fixed long sign over
every SHARED row, not only fired ones. Accumulated above the voter and
direction-policy filters deliberately: restricting it to bars the vote fired on
would compare the vote against a baseline the vote itself selected. Always-short
is exactly its negative, so one pass covers both.
BY SIDE - the vote's own payoff split by the direction it took, still sign
corrected, at the rung the live signal is actually trading:
both sides positive -> directional skill, it pays going either way
one positive, one negative
and roughly cancelling -> it found the drift, and the pooled mean is
saying nothing about skill
This is drift-free BY CONSTRUCTION - drift enters both sides with opposite sign
after the correction, so it cannot manufacture a two-sided positive. That is
precisely what a pooled mean cannot tell you and what no baseline subtraction
fully recovers.
The split is taken at the CHECKPOINTED rung, not this era's derived one: the
derived rung is not known until after the row loop that accumulates the split,
and the checkpointed rung is the operating point the question is actually about.
Still gates nothing.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
43c1b27654 |
feat(gate): measure what a call was WORTH, not only how often it was right
The ensemble deploy gate certifies PRECISION against a chance rate and has never known whether a correct call pays for its own spread. Every verdict this project has recorded - 33% precision against a 14% chance rate, an edge that clears its exact-binomial bar comfortably - is silent on the one question that decides whether any of it is tradeable, and the cost boundary is exactly where several earlier edges died with their precision already believed. Adds a per-row payoff measurement, taken once per ROW (a chart property, not a member one) at the same time the label is written: * forward close move over K = round(SwingLifespanEstimate()) bars, * the up and down extreme excursions over the same window, each divided by the bar's own ATR. K is deliberately the label lifespan the effective-sample-size deflation already uses, so precision and payoff describe the same window and can be read in one sentence. POLICY-FREE: no stop, no target, no trailing rule. It measures the SIGNAL, not a trade-management choice layered on top - exit shaping moves payoff around without creating any, so mixing the two would hide which was responsible. Stored unsigned by direction; the sign comes from the vote at verdict time, and a short's excursions SWAP rather than negate - negating them would report a short's worst case as a negative best case. The newest K bars of the OOS slice have no forward window and are dropped from the tally with their own denominator, never counted as a zero move: that is the leading-edge trap that made the lag profile's first run a false positive. The era verdict now prints mean R, MFE and MAE at the certified rung against the spread in the same ATR units. It GATES NOTHING - wiring a policy to an unvalidated payoff number is how a measurement becomes a decision before anyone has checked it. Build tag payoff-v1. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
0e1e952b96 |
feat(topology): cap the input window at 6 bars for capacity - 588 inputs -> 294
Three charts (SP500, XAUUSD, XTIUSD) sat on the FIRST_LAYER_MIN_WIDTH floor
even after pooling took SP500 from 4.1 to 1.8 weights per independent
observation. ComputeFirstLayerWidth needs width <= ~331 to clear it; 49 columns
x 12 bars = 588.
TWO QUESTIONS, AND THE WINDOW IS NOW THE SMALLER ANSWER. The ZigZag ladder
answers "how far back is a swing worth looking" and says 12. The capacity
budget answers "how far back can this much data support" and says 6. Taking the
min stops the first writing a cheque the second cannot cover.
WHY THE LAG AXIS AND NOT THE COLUMN AXIS - the choice was between this and a
per-column mask (designed, parked on feature/column-mask):
- On the LAG axis there is a measured null. The corrected lag profile finds no
linear structure at any lag within +/-50, on all six charts, family-wise
p=1.0000, argmax scattered across different columns and lags per chart.
- On the COLUMN axis the two measures that would justify a mask - marginal MI
retention and variance share - are explicitly blind to joint and temporal
structure, and the columns they would delete include the entire price core,
which is the one place such structure would plausibly live.
Cutting where there is a measured null beats cutting where the instrument
cannot see. Corroborating: PAI/CONV/LSTM/HYBRID score within ~1pp of each
other, so the temporal machinery is not visibly earning the deeper lags.
THE CAP IS A FLEET CONSTANT, NOT A PER-CHART DERIVATION. Pool rows are keyed on
`bars x columns`, so a capacity cap computed from a chart's own observation
count would differ across the fleet by construction and hand every chart its
own layout, its own fingerprint and its own pool of one - exactly what orphaned
SP500. Set from the most starved chart; every chart shares it.
Conv survives: CONV_RECEPTIVE_FIELD_BARS is 3, so a 6-bar window still leaves 4
sliding positions. LSTM sequence length becomes 6.
RETRAIN-FORCING and POOL-INVALIDATING: width changes, so old .nnw and old
TrainPool rows are both incompatible. Wipe both - which puts the fleet back in
the cold-start condition
|