forked from MrBaro75/Warrior_EA
747 commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
213b3aac15 |
fix(chart): the reconciliation never ran - it was hooked to a sweep deployed charts do not do
cooldown-recon put the chart-wide cooldown at the end of the overlay sweep. It executed ZERO times. This store's own header already said why: the sweep 're-arms only when an era ends. A DEPLOYED ensemble runs no further eras'. Five of six charts were deployed, so there were zero 'Filtered view: swept' lines in the entire session while the saved files still held 148 same-side pairs under 30 bars on XTIUSD. Moved to the completion of the progressive vote-arrow restore, which runs on every chart including deployed ones. The restore thinning alone was never going to be enough either: MT5 persists chart objects in profiles\Charts\*\chart*.chr, so arrows drawn under an older window are ALREADY on the chart when the process starts, and a freshly-thinned restore just adds to them. Two correctly-thinned sets still union into clusters. The chart is the only authority. Same construction as before: OBJ_TREND only (the line is the canonical half of a mark, matching Snapshot()), sorted by time first because object order is not time order, and the gap>0 guard so a mis-ordered set fails visibly by keeping rather than silently by deleting. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
4e2cdd3aff |
fix(chart): reconcile the cooldown over the CHART, not over one producer's record
The record-based prune was individually correct and still left clusters. It is not the only producer of a vote arrow: the persisted-arrow RESTORE thins its own list from its own state, the overlay sweep thins its own list from its own state, and the live gate marks the current bar from a third. Each spaces ITS OWN survivors 30 bars apart; interleaved on one chart the union sits 1 bar apart. Two independent thinning passes over one namespace produce a union, not an intersection. MEASURED from the saved .votearrows files, which is what the chart actually holds: XTIUSD 284 arrows, 180 gaps under 30 bars, 148 of them SAME-SIDE, min gap 0 XAUUSD 263 arrows, 177 gaps under 30 bars, 144 same-side, min gap 1 EURUSD 309 arrows, 176 gaps under 30 bars, 110 same-side, min gap 0 while every producer's own log reported it had thinned correctly. The sweep's 'drew' counter says what ONE producer drew; the chart is the union. Verifying on that counter is what let this stand through four builds. The authority is now the chart itself: after the sweep, walk every SIG_VOTE_PREFIX OBJ_TREND object, sort by time, enforce one window. Whatever drew an arrow, this runs last. OBJ_TREND only - a mark is a line AND an arrow and the line is canonical, the same test Snapshot() uses. Sorted first, because object order is not time order and an unsorted forward walk yields negative gaps, which is how a prune once deleted 272 of 273 arrows. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
844aac653a |
fix(train): the OOS final pass ran a full epoch at an undecayed rate
Capping the pass at m_etaCeiling was not enough. Measured on the first two live runs: USDCAD 0.00085 over 13,335 bars, EURUSD 0.00242 over 15,041 - a 3x spread across charts, because a chart whose plateau ladder reset recently still carries a high eta and the cap never bound. The slice turns out to be roughly HALF the data, not a tail, so one pass over it at the model's own rate is a full training epoch on a model that has already been selected and certified. That is materially more than the 'just a bit finer weights' this was asked for. OOS_FINAL_PASS_ETA_SCALE (0.25) now scales the rate. Scaling rather than shortening the pass keeps the whole slice in play - seeing the held-out bars at all is the point - while making the step proportionate to an already-selected model. USDCAD and EURUSD have already taken the unscaled pass; that is not reversible without a retrain. USDJPY has not converged yet and will get the corrected one. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
0dda31534e |
fix(chart): purge stale member arrows even when there is no sidecar to restore
The purge was gated on m_arrowRestorePending, so it only ran when the .arrows sidecar had something to load. Stale arrows do not come from the sidecar - MT5 persists chart objects in profiles\Charts\*\chart*.chr independently of it. A chart carrying member arrows from a session where DrawUnfilteredSignals was ON would therefore never be cleaned, because the sidecar it would have needed is gone. Currently inert: there are zero .arrows sidecars on disk. Which is also the correction to the diagnosis in c5b9a1a's message - the '264 arrows restored' that prompted it was a MISREAD of 'queued 264 combined-vote arrows', so member arrows were never the cause of the reported clutter. The gate remains correct as defence; it was not the fix. The window was. Also silences the line when there is nothing to skip and nothing to clear: it runs on every chart on every start, and six 'removed 0' lines are noise. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
cc45451a22 |
fix(signal): the source override was applied before the minutes input, which could defeat it
Signal_CooldownMinutes is a per-chart input like Signal_CooldownBars, and it ran AFTER SignalCooldownOverrideBars - so a chart carrying a stored minutes value would have silently defeated the override, which is precisely the problem the override was added to solve. Latent rather than live: every chart currently holds SCM_OFF. That is the kind of 'works today' that stops working the first time someone sets the input, and it would have looked like the override simply not functioning. Minutes resolve first; the const now has the last word. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
c5b9a1ad10 |
fix(signal): a changed input default cannot reach an already-attached EA
Raising Signal_CooldownBars from 10 to 30 changed nothing. All six live charts
kept reporting a 10-bar window, because MT5 stores an input PER CHART in
profiles\Charts\*\chart*.chr and an already-attached EA ignores a changed
default entirely. This codebase already documents that trap, in the derived-
threshold comment in Training.mqh - and converting SignalClusterWindow from a
const to an input reintroduced the exact problem the const existed to avoid.
SignalCooldownOverrideBars (const, 30) now wins over the input; 0 hands control
back to the panel. Tunability per chart is kept, source-correctability is back.
Not applied when the input says OFF: an operator who switched the cooldown off
meant it, and silently re-enabling it from source would be the same surprise
pointed the other way.
ALSO gates the per-model arrow restore on DrawUnfilteredSignals. DrawObject()
returns early when the raw view is off, but AdvanceChartSignalRestore called
WarriorPlotSignalLevel DIRECTLY and never checked - so every restart repainted up
to MAX_PERSISTED_ARROWS per-model opinions per member, four members per chart, on
top of the combined-vote arrows. Same shape as the vote-arrow restore bug in
|
||
|
|
09af7d5ee9 |
feat(signal): default the cooldown to 30 bars - 10 thinned almost nothing
Measured on the live fleet at 10 bars: the restore thinned 100 of 364 and the overlay prune 23-66 per chart, leaving 182-250 arrows over ~5000 bars. Every layer was working; the window was simply below the ~20-bar natural spacing between vote arrows, so it could only catch the tightest pairs. The floor is principled, not cosmetic. A trade on this label is held for 5 + the median ZigZag leg = 18-19 bars, so any second signal inside that window is the same trade being re-announced. 20 is the smallest defensible value and 30 is one step above it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
3218db4a38 |
feat(train): ONE pass over the held-out slice at deploy, on the restored checkpoint
The OOS slice is the newest history and the model never trains on it, while online learning adapts to every bar resolving AFTER deployment. That leaves a gap exactly at the handover, over the most regime-relevant data there is. This closes it: select on validation, then refit on everything, which is standard practice. Placed AFTER Net.RestoreWeights() and ResetOptimizerState() and BEFORE PersistDeployedModel(), so it refines the weights that were actually SELECTED rather than whatever the run happened to end on, and what it produces is what gets written down. THE COST IS REAL AND IS NOW STATED IN THE LOG. The deploy line promises "every model reverts to the weights it held at the era whose combined vote scored best, so the ensemble that trades is exactly the one that was measured". After this pass that is no longer literally true, so the pass prints that the certified numbers belong to the PRE-PASS weights and must be quoted that way. Set EnableOosFinalPass=false to keep certified == traded exactly. Guards: * ONE-SHOT PER RUN, and the flag is set BEFORE the loop so no early return inside it can leave the pass eligible to fire twice over bars it already trained on. Reset at m_trainRunActive=true, because a retrain is a fresh selection and earns a fresh pass. * THE CONVERGED RATE, never a plateau-boosted one: m_modelEta can still carry PLATEAU_RESTART_BOOST from an escape attempt, and this is a refinement of a selected model, not another warm restart. g_eta is what backProp reads, so that is what is capped and restored. * OLDEST -> NEWEST. Series indices count backwards, so decreasing i moves forward in time - the order the bars happened in. * A failed feedForward is never followed by backProp; the output layer would still hold the previous sample's activations and the update would be this bar's label against another bar's prediction. * m_oosFinalPassCutoff records the newest bar consumed and is deliberately NOT cleared on a new run, so a later run can say plainly that its out-of-sample window reaches back into bars this model has already seen. Expect the gain to come from CURRENCY rather than finer weights: OOS precision was measured flat from era 20 while in-sample error kept falling, so the data this model can already see is exhausted. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
322c052a65 |
fix(chart): the persisted vote-arrow restore put back every arrow the cooldown removed
Clutter remained after cooldown-v3 because there is a FOURTH producer of
SIG_VOTE_PREFIX arrows: CVoteArrowStore, which replays a .votearrows file and
draws into the SAME object names as the overlay. Replaying a file written before
the cooldown existed therefore resurrects exactly the arrows the prune deleted.
On a DEPLOYED chart that is the entire arrow set. The store's own header says
why: the overlay re-sweep "re-arms only when an era ends. A DEPLOYED ensemble
runs no further eras" - which is the reason this store exists at all, and also
the reason nothing would ever have removed those arrows again.
The restore now thins to the cooldown at Load(), before the progressive draw is
armed, so it is idempotent: a file already written from a cooled chart passes
through untouched, an older one is corrected once.
IT SORTS BY TIME FIRST, AND THAT IS NOT OPTIONAL. Snapshot() walks
ObjectsTotal(), so the record is in OBJECT order - its own comment says so, and
the existing MAX_KEPT trim already sorts a copy for exactly this reason. Applying
a spacing rule to an unsorted record yields negative gaps, and a negative gap is
inside any window: that is the bug that wiped 272 of 273 arrows in
|
||
|
|
2ca32e933f |
fix(signal): the overlay cooldown prune ran backwards and left ONE arrow per chart
cooldown-v2 suppressed 272 of 273 on SP500, 320 of 321 on EURUSD, 329 of 330 on USDCAD - one surviving arrow on every chart in the fleet. The record is OLDEST-FIRST. The prune walked it backwards, so every gap came out NEGATIVE, and a negative gap is always <= the window: everything after the first arrow was suppressed. The direction was taken from the member comment on m_overlayIndex, which reads "walking newest -> oldest" and is WRONG. The sweep DECREMENTS a SERIES index (0 = newest) from MathMin(span, barsAvail-150) down to m_overlayStopIndex, so it walks OLDEST -> NEWEST. The pre-existing overlay NMS at the draw site agrees - it tests (m_overlayNmsKeptIdx - idx) and expects that to be positive for later bars. A stale comment counts as a guess, and this one cost a build. Comment corrected at the declaration so the next reader is not misled the same way. Guard added: the gap must be > 0 as well as <= the window. A non-positive gap means the record is not in the order this loop assumes, and suppressing the whole chart is precisely what that looks like from the outside - so it now fails visibly by KEEPING rather than silently by deleting. Found only because the verification was the drawn arrow count rather than an assertion that the code was correct. Compiling clean said nothing about it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
110080dc86 |
fix(signal): the cooldown belonged at the VOTE layer, as a filter - not per member
cooldown-v1 extended NmsLiveAccept, which declusters each MEMBER's own signal. That is not what the charts show and not what trades. The combined vote in CExpertSignalCustom had NO spacing rule at all - grep found not one reference to the cluster window in that file - so four individually-declustered members were averaged into a vote that could fire on consecutive bars. Measured live: 2,970 voting bars becoming 299-328 vote arrows. Proof of the diagnosis, from the deployed fleet under cooldown-v1: SP500 273 and XAUUSD 212 arrows, unchanged from before the change. The member-level rule could not touch them. Gated where the vote becomes a trade - CheckOpenPosition, beside the open-prohibition and open-market-closed checks, tracing as "open-cooldown". That is the filter chain the request asked for from the start and it is where this should have gone first. Suppression there means no order AND no live arrow, honouring the same "no arrow, no vote, no position" contract the member rule already had. THE DRAWN HISTORY NEEDED A SECOND PASS, NOT AN INLINE TEST. The overlay sweep walks NEWEST->OLDEST and is chunked across ticks, so an inline cooldown would keep the NEWEST bar of a cluster while the live gate keeps the FIRST, and the drawn set would contradict the traded set - the exact defect the renderer's own comments warn about. The sweep now records what it drew and prunes it backwards over that record, which is forward in time. Direction() is a TRANSACTION that can run more than once on a bar, so the live accept is cached per bar time. Without that a second call flips the bar's verdict after it has already journaled one. One resolver, WarriorSignalCooldownBars(), now serves both layers so they can never disagree about the window. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
6308a19f27 |
feat(signal): make the signal cooldown tunable, and add a hard any-direction gate
The declustering the charts needed already existed - NmsLiveAccept, per-direction run-collapse plus cross-direction resolution plus strict alternation - and it was already set to 10 bars. It could not be TUNED: SignalClusterWindow was a compile- time const, so finding the right value needed a rebuild. That is the actual gap. Now three inputs, as enum dropdowns: Signal_CooldownScope per-direction, or a hard any-direction gate on top Signal_CooldownBars SCB_OFF..SCB_50, default 10 Signal_CooldownMinutes SCM_OFF..SCM_1440, overrides bars when set Minutes resolve against the CHART period and round UP, so a cooldown asked for in wall-clock is never silently shorter than requested and survives a timeframe change. SCB_/SCM_ prefixes are deliberately unique. M15/M30/M60 are ALREADY members of NF_LOOKBACK_PRESETS, and MQL5 binds a duplicated enum member to the first-declared enum silently - the obvious names would have compiled straight into the news filter's values. THE ANY-DIRECTION GATE IS ADDITIVE, NOT A REPLACEMENT, and the first cut of this had it backwards. Measured on the live log: the current rules draw 222 arrows over 4999 bars, while a BARE 10-bar cooldown permits up to 454 - because ALTERNATION is what declutters today, not the window. Swapping the rules out would have roughly doubled the clutter it was asked to remove. Layered, it can only ever suppress more. Suppressed bars still advance the per-direction last-SEEN cursors, so a run straddling the boundary does not restart as if it were fresh. Applied at all THREE sites that must agree - live inference, OOS pass-3 scoring and the chart renderer. Their own comments say why: an arrow set that does not obey the same rule as the traded set shows calls the EA would never take. Also corrects a stale comment that called this window "display only". It is not: when it suppresses, the live path zeroes the signal outright - no arrow, no vote, no position. Training never sees it, so these cost no retrain and are correctly absent from the fingerprint. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
8083a31754 |
diag(gate): move the conviction curve to the horizon that has value, and add mean-d per rung
The 5-bar conviction curve cannot answer the question it was built for. The oracle measures ~0 at 5 bars across three charts (+0.012, -0.054, +0.064), so PERFECT foresight earns nothing there and no rung can show payoff either. Every reading it produced was null by construction. It was placed at 5 bars for statistical power, before the oracle showed what that horizon is worth. Kept as a control; the hold-horizon curve is the one to read. Also adds MEAN DISTANCE-TO-PIVOT PER RUNG, which is the high-power form of the same question. Payoff falls ~0.34 ATR for every bar of distance to the pivot (fleet-pooled: d=1 +2.095, d=2 +1.743, d=3 +1.300, d=4 +0.969, d=5 +0.769, wrong calls -0.668). So a rung that selects NEARER pivots is worth more per call even at unchanged precision - and mean-d is a far tighter statistic than mean-payoff, because d spans five bars where payoff spans several ATR. That matters because it can REOPEN a lever I closed. Precision does not rise with the rung - every 15-vs-10 comparison across six charts sits below 0.71 sigma - so the threshold looked exhausted. But precision is not the only thing a threshold can select for. If conviction correlates with proximity to the pivot, raising it buys payoff without buying precision. Directional labels only: an incorrect call has no pivot and therefore no distance, and folding those in as zero would read as "this rung picks pivots that are imminent" when it means the opposite. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
0d9320cc87 |
diag(gate): the ORACLE - what a perfect caller of this label would earn
The ceiling on the target, and the measurement that decides where the work goes. Same payoff arithmetic, signed by the LABEL's direction instead of the vote's, over every directionally-labelled shared bar. If a model that got EVERY pivot right still earns nothing over the holding horizon then the target carries no money and no amount of model improvement reaches any - the label, not the network, is what has to change. If the oracle earns well the target is sound and the shortfall is the model's. Those are completely different programmes and nothing so far distinguishes them. It uses no forecast, so it is not a leak: it is the value of perfect foresight OF THIS LABEL, reported as a benchmark. Nothing may trade on it. Accumulated above the voter and direction-policy filters, like the zero-skill book, because it is a property of the bars and their labels rather than of what the vote did with them. A bar with no directional label offers a perfect caller nothing to take and is skipped rather than counted as zero - the benchmark is "every call it COULD make". Motivated by the first skill-by-distance row, which already reframes the day: correct calls earn +0.75 to +1.90 ATR against a spread of 0.005-0.042, and incorrect ones cost -0.66. That puts break-even precision near 32% against a measured 33-37% - thin, but on the right side, and utterly unlike the "no payoff" reading the confounded 5-bar window suggested. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
1a9b56e3b0 |
diag(label): expose bars-to-pivot - the confound the payoff test was missing
CORRECTION to what the payoff instrument was measuring. The 5-bar horizon looked like the powered test and it is confounded. SwingPivotDirectionLabel returns Buy when a swing LOW lands up to PIVOT_LABEL_TOLERANCE_BARS bars AHEAD, and says the quiet part itself: gating on where the pivot sits relative to entry "would drop exactly the bars where the turn has not finished coming to us", and how much adverse move remains before the turn "is a trade-management question". So on a CORRECT Buy call price is often still falling for d more bars. A window shorter than d measures the APPROACH, not the leg, and its negative contribution is expected on the calls that are RIGHT. The tight null at 5 bars (-0.012 +/- 0.074) is therefore not evidence of no payoff. Neither horizon is both clean and powered: 5 bars is powered and confounded, 18-19 is clean and has an SE of 0.277. (idx - P1) was computed in the label and thrown away. Now cached beside m_labelResolveAge under the same validity flag, and bucketed in the era verdict. DELIBERATELY NOT USED AS A PER-CALL HORIZON, which is the trap sitting right next to this: d exists only on bars the label found a pivot for, so a horizon that varied with d would hand correct and incorrect calls different windows and bias the comparison outright. The horizon stays fixed; d only buckets. The bucket for "the label called no pivot here" is reported by name rather than folded in, because it is the control the others are read against. Buckets 1..N condition on the label, so they describe the MECHANISM, not what a book earns. Reads: rising with d means the edge is in EARLY calls and the tolerance window is spending it - fixable by reweighting the loss, not by a new label. Flat means that hypothesis dies. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
3372b82dfa |
diag(gate): the conviction curve - does payoff rise with vote magnitude?
The practical question behind "can I just trade the strongest signals" is whether payoff rises with vote magnitude. The threshold sweep already visits every rung, so the whole curve costs four arrays and no extra pass. Reported as the DRIFT-FREE statistic per rung - long plus short, both sign corrected - with the two halves alongside. The halves alone invite reading a drift-fed long side as skill, which is exactly the error the zero-skill book caught at the certified rung: an always-long book earns MORE than the vote on two of three charts. Taken at the SHORT horizon, which is the one with the power. Pooled across the three training charts the certified rung reads -0.012 +/- 0.074 ATR - a tight null, 95% interval [-0.16, +0.13], with the long/short pattern (+0.030 against -0.041) being the drift signature exactly. The hold horizon agrees and is 3.7x noisier, so the answer is not a horizon artifact. Precision is already known not to rise significantly with the rung (every 15-vs-10 comparison across six charts sits below 0.71 sigma). If payoff rises anyway that is a surprise worth having; if it does not, the two agree and the threshold lever is closed on both counts. Still gates nothing. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
feaadd80a2 |
diag(gate): split the payoff by side at the horizon that can actually resolve it
The by-side test is the one that separates directional skill from drift, but at
the HOLD horizon it cannot answer: payoff overlap is the horizon itself, so an
18-bar window leaves ~65 independent observations per chart and a standard error
of 0.25-0.45 ATR against an effect that would matter at 0.1.
The 5-bar window carries ~3.8x the independent observations and roughly half the
standard error. It buys that power by risking a window that ends before the
pivot has committed - which is exactly why the horizon was widened in
|
||
|
|
ce4f74fe2c |
diag(ensemble): measure how much the four members actually disagree
The ensemble beats its best single member by +2.2 to +6.8pp on all six charts - sign-stable across six instruments, so the ensemble is doing real work rather than diluting. How much MORE is available depends entirely on how decorrelated the members are: the variance of an m-member average scales as (1+(m-1)r)/m, so at r=0.8 four models are worth about 1.2 independent ones and at r=0.3 nearly 3. Nothing measured that, so the obvious next lever - different feature subsets per member, or a fifth architecture - could not be costed. Both force a full retrain of 24 models, which is not a price to pay on a guess. Measured on the SIGNED VOTE, which is what actually gets averaged: not accuracy, not raw confidence. Two members can agree on direction almost always and still contribute independently through magnitude. Accumulated over every SHARED row rather than fired ones - restricting to fired rows would measure agreement only where the members already agreed enough to fire, which is the sample most biased toward agreement. A member whose signed vote never varies (all abstentions, a dead tier) is SKIPPED rather than counted as r=0, which would drag the mean toward "decorrelated" using a member carrying no information at all. Reported as an effective member count, which is the honest way to say what four models are worth. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
7500e08e17 |
feat(gate): the payoff number needed a zero-skill book and a by-side split
payoff-v1 reported what a call was worth and nothing to compare it against. A
positive mean R is not a finding on its own: if the instrument drifts, an
ALWAYS-LONG book earns a positive mean too, and drift is the one anomaly family
this project has found that survives cost - so the vote would be reporting the
market's own move as if it were its own.
Two comparisons, and the second is the one that decides it:
ZERO-SKILL BOOK - the same forward move accumulated with a fixed long sign over
every SHARED row, not only fired ones. Accumulated above the voter and
direction-policy filters deliberately: restricting it to bars the vote fired on
would compare the vote against a baseline the vote itself selected. Always-short
is exactly its negative, so one pass covers both.
BY SIDE - the vote's own payoff split by the direction it took, still sign
corrected, at the rung the live signal is actually trading:
both sides positive -> directional skill, it pays going either way
one positive, one negative
and roughly cancelling -> it found the drift, and the pooled mean is
saying nothing about skill
This is drift-free BY CONSTRUCTION - drift enters both sides with opposite sign
after the correction, so it cannot manufacture a two-sided positive. That is
precisely what a pooled mean cannot tell you and what no baseline subtraction
fully recovers.
The split is taken at the CHECKPOINTED rung, not this era's derived one: the
derived rung is not known until after the row loop that accumulates the split,
and the checkpointed rung is the operating point the question is actually about.
Still gates nothing.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
98f485b901 |
fix(gate): the payoff horizon ended before the pivot it was measuring
The first cut measured payoff over SwingLifespanEstimate() bars. That is
PIVOT_LABEL_TOLERANCE_BARS - a constant of the TARGET describing how many bars
share one pivot event - and it is the wrong horizon for what a call is worth.
The label fires when a pivot lands WITHIN that window. So at that horizon the
pivot may only just have committed, and a perfectly correct call can still show
a negative forward move because the turn it predicted has not had one bar to
run. Measuring only there would understate the payoff of a signal working
exactly as designed, and could inflip its sign.
Measures two horizons and reports both:
SHORT = PIVOT_LABEL_TOLERANCE_BARS "has the pivot arrived" - a control
HOLD = that + the median ZigZag leg the pivot PLUS the leg it opens,
which is how long a trade on this
call would actually be held
Adds CTopology::SwingLegMedianBars(). It is deliberately NOT the same thing as
SwingLifespanEstimate() and the declaration says so: the lifespan is a constant
of the target and is what the effective-sample-size deflation divides by, while
the leg median is a measurement of the chart and is how long the move runs.
Conflating them is what produced the wrong horizon in the first place.
Non-const and lazily measured, because a model that adopted its .cfg never
walked the chart and would otherwise report HISTORY_BARS_FALLBACK as if it were
a measurement - the same lazy pattern DeriveHistoryBars() already uses.
Reporting both horizons is also the guard against picking one and calling it
the truth. A break-even conclusion here has already been overturned once purely
by getting a horizon wrong.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
43c1b27654 |
feat(gate): measure what a call was WORTH, not only how often it was right
The ensemble deploy gate certifies PRECISION against a chance rate and has never known whether a correct call pays for its own spread. Every verdict this project has recorded - 33% precision against a 14% chance rate, an edge that clears its exact-binomial bar comfortably - is silent on the one question that decides whether any of it is tradeable, and the cost boundary is exactly where several earlier edges died with their precision already believed. Adds a per-row payoff measurement, taken once per ROW (a chart property, not a member one) at the same time the label is written: * forward close move over K = round(SwingLifespanEstimate()) bars, * the up and down extreme excursions over the same window, each divided by the bar's own ATR. K is deliberately the label lifespan the effective-sample-size deflation already uses, so precision and payoff describe the same window and can be read in one sentence. POLICY-FREE: no stop, no target, no trailing rule. It measures the SIGNAL, not a trade-management choice layered on top - exit shaping moves payoff around without creating any, so mixing the two would hide which was responsible. Stored unsigned by direction; the sign comes from the vote at verdict time, and a short's excursions SWAP rather than negate - negating them would report a short's worst case as a negative best case. The newest K bars of the OOS slice have no forward window and are dropped from the tally with their own denominator, never counted as a zero move: that is the leading-edge trap that made the lag profile's first run a false positive. The era verdict now prints mean R, MFE and MAE at the certified rung against the spread in the same ATR units. It GATES NOTHING - wiring a policy to an unvalidated payoff number is how a measurement becomes a decision before anyone has checked it. Build tag payoff-v1. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
0e1e952b96 |
feat(topology): cap the input window at 6 bars for capacity - 588 inputs -> 294
Three charts (SP500, XAUUSD, XTIUSD) sat on the FIRST_LAYER_MIN_WIDTH floor
even after pooling took SP500 from 4.1 to 1.8 weights per independent
observation. ComputeFirstLayerWidth needs width <= ~331 to clear it; 49 columns
x 12 bars = 588.
TWO QUESTIONS, AND THE WINDOW IS NOW THE SMALLER ANSWER. The ZigZag ladder
answers "how far back is a swing worth looking" and says 12. The capacity
budget answers "how far back can this much data support" and says 6. Taking the
min stops the first writing a cheque the second cannot cover.
WHY THE LAG AXIS AND NOT THE COLUMN AXIS - the choice was between this and a
per-column mask (designed, parked on feature/column-mask):
- On the LAG axis there is a measured null. The corrected lag profile finds no
linear structure at any lag within +/-50, on all six charts, family-wise
p=1.0000, argmax scattered across different columns and lags per chart.
- On the COLUMN axis the two measures that would justify a mask - marginal MI
retention and variance share - are explicitly blind to joint and temporal
structure, and the columns they would delete include the entire price core,
which is the one place such structure would plausibly live.
Cutting where there is a measured null beats cutting where the instrument
cannot see. Corroborating: PAI/CONV/LSTM/HYBRID score within ~1pp of each
other, so the temporal machinery is not visibly earning the deeper lags.
THE CAP IS A FLEET CONSTANT, NOT A PER-CHART DERIVATION. Pool rows are keyed on
`bars x columns`, so a capacity cap computed from a chart's own observation
count would differ across the fleet by construction and hand every chart its
own layout, its own fingerprint and its own pool of one - exactly what orphaned
SP500. Set from the most starved chart; every chart shares it.
Conv survives: CONV_RECEPTIVE_FIELD_BARS is 3, so a 6-bar window still leaves 4
sliding positions. LSTM sequence length becomes 6.
RETRAIN-FORCING and POOL-INVALIDATING: width changes, so old .nnw and old
TrainPool rows are both incompatible. Wipe both - which puts the fleet back in
the cold-start condition
|
||
|
|
1e6d00e602 |
fix(vote): a resumed converged model passed the eligibility test and then abstained on every bar
Found by restarting the terminal against three charts that had just deployed - the exact scenario |
||
|
|
bbe26a09fb |
fix(diag): the lag profile's first-ever run produced a spectacular false positive
Resurrecting the diagnostic in
|
||
|
|
4113afd293 |
fix(diag): the linear lag profile never ran once - it started at the newest bar
Every chart, every era, on every run in the logs: "linear lag profile skipped - only 0 contiguous OOS bars." That reads as "not enough data". It was not. The walk never started. ReportLinearLagProfile walks newest-first from r=0 and breaks on the first row without a label, to avoid splicing across a hole. But r=0 IS the newest bar, and a forward-looking swing-pivot label cannot be resolved there by construction - the opposite pivot has not committed yet. So HasLabel(0) is false, the loop breaks on its first iteration, and n=0. Permanently. The leading gap is SYSTEMATIC (always about the label resolution), not a hole in the middle of the series, so stepping over it splices nothing. Contiguity is still enforced from the first labelled row onward. WHY THIS MATTERS BEYOND THE DIAGNOSTIC: the input window is 12 bars, and the capacity budget divides by width = columns x bars. Cutting the window is the largest lever left for the three charts still pinned to the 16-unit first-layer floor, and there has been no measurement of whether the deeper lags carry anything - because the one diagnostic that would answer it has never produced a number. The old lag verdict in memory predates the pivot-event label. The skip message now reports where the walk ran out, so "0 from r=0" (never started) is distinguishable from "0 from r=37" (genuinely short window). Build tag -> lagprofile-v1. NOT a feature-layout change: no retrain, models resume from their weights and the training pool stays valid. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
6c2959dc32 |
feat(topology): re-derive capacity once when the training pool appears
A cold fleet start sizes every model BEFORE any chart has published a pool
file, so the first layer is budgeted as if the chart trains alone and then
pinned to .cfg. This is not a rare race - it is what happens EVERY time the
feature layout changes, because that invalidates the pool and forces a wipe.
Correcting it by hand needs a two-phase start: run the fleet to fill the pool,
stop, wipe the weights while KEEPING the pool, restart so derivation sees it.
That is not something an unattended fleet can do for itself, and getting it
wrong is silent - the models simply stay narrow.
TuneIndicatorsAndTrain now notices that the pool has appeared and re-derives
once, reusing ResetWeights() - the existing tested path that re-measures all
four sizes, rebuilds and rewrites the .cfg. No second copy of that logic.
Bounded on every axis that could make it a loop:
- once per model (the flag is set BEFORE the reset, because ResetWeights
zeroes m_eraCount and the model would otherwise re-qualify forever)
- only while era <= CAPACITY_RESIZE_MAX_ERA, so the discarded eras are worth
nothing
- only on CAPACITY_RESIZE_MIN_GROWTH real growth
- only if the recomputed width actually differs; if it does not, the check
settles itself rather than re-running the census every era
Safe against the one thing that would make it self-defeating: the derived width
is NOT part of BuildModelFingerprint, so a model that resizes does not leave
the pool it resized for.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
00699e1af8 |
fix(chart): stale combined-vote arrows survived every wipe, because two files lived outside Warrior_EA\
Operator report: arrows labelled as restored from a previous session on a fleet training from era 0. Confirmed - all six charts restored 115-431 combined-vote arrows drawn by models that no longer exist. TWO INDEPENDENT DEFECTS, either of which alone causes it. 1. CVoteArrowStore::Discard() HAD NO CALLER. The member-scoped .arrows file is cleared by ClearPersistedChartSignals on a fresh topology. The CHART-scoped .votearrows store has an equivalent Discard(), written for exactly this, and nothing ever called it. The store is keyed on the DB config fingerprint, which does not move when a model is wiped, so it reloaded across any reset - fresh topology, panel weight reset, or a model-file wipe. A vote is a claim made by a specific set of members. If any member rebuilt from scratch this run, the whole stored history is void, so g_warriorFreshTopologyThisRun is now raised wherever a member discards weights or builds a fresh topology, and the store Discards instead of Loads. 2. TWO WARRIOR FILES LIVED OUTSIDE Warrior_EA\. .sigvis and .votearrows were written to the ROOT of Common\Files, outside the one directory that "wipe the Warrior EA files" has always meant. Two consecutive wipes this session left them standing untouched, and neither wipe was as fresh as reported. Both now live under Warrior_EA\ChartState\. A wipe that does not remove all of a program's state is not a wipe, and nothing in the log told the operator which files were missed. NOTE for anyone re-running the wipe: pre-existing WarriorVote_*.votearrows and Warrior_EA_*.sigvis in the Common\Files ROOT are orphaned by this change and should be deleted once. Build tag -> fleet-pool-v3. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
afe1038d11 |
fix(topology): stop a training-alone size becoming permanent, and stop the keep-screen latching underpowered
1. THE POOL FIX WAS LANDING ON A TOPOLOGY THAT COULD NOT SEE IT. ComputeFirstLayerWidth budgets against EstimatedInSampleBars, which counts this chart's own bars PLUS the training pool. On a COLD fleet start every chart derives and pins its topology BEFORE any chart has published a pool file - measured on the 18:13 start, model creation at 18:13:21 against a first publish at 18:13:48. All six sized as if training alone, wrote that into .cfg, and adopted it back on every later start even with the pool full. SP500 ran a first layer floored to 16 while adopting 30229 peer rows. Adopt-don't-compare exists to protect weights shaped by those sizes. It was also running for a model with NO .nnw, where there is nothing to protect and the .cfg is just a record of one unlucky moment. The four derived sizes are now re-measured when no weights exist. Safe on all three counts that matter: free (nothing to discard), cannot loop (once weights exist the .cfg is authoritative again), and cannot fragment the pool - the derived width is NOT in BuildModelFingerprint, which keys only on the FEATURE layout. Verified: field 2 of the fingerprint is LEGACY_HISTORY_BARS_SLOT, not the first-layer width. TO TAKE EFFECT the weights must be wiped while the TrainPool is KEPT - the census has to be non-empty at derivation time. A full wipe empties the pool and reproduces the original condition exactly. 2. THE KEEP-SCREEN LATCHED ON AN UNDERPOWERED SAMPLE. MI_MIN_SAMPLES is a floor for "can this be computed", and it was being used as the bar for "is this answer final". The screen fired on the first era clearing 200 rows and latched, measuring at 202-773 samples where a warm chart gives ~2065. Columns kept then tracked SAMPLE SIZE rather than information - EURUSD kept 0 of 49 at n=202, SP500 kept 15 at n=773, and the ordering across all six charts was very nearly monotone in n. A thin sample is still measured and printed, but it no longer closes the question: below MI_GOOD_SAMPLE_FRACTION of the target the result is labelled underpowered and a later era supersedes it, bounded by the same attempt budget. An underpowered screen that latches is worse than one that waits, because it looks like a result. Build tag -> fleet-pool-v2. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
a970405042 |
feat(pool,mi): one feature layout fleet-wide, and the keep-screen stops self-disabling on a cold start
TWO CHANGES, BOTH RETRAIN-FORCING BY INTENT.
1. SP500 was training alone, and one alt-data column was the reason.
The alt block's width joins the model fingerprint, and the pool reader only
adopts peer rows whose fingerprint and width match. The exporter gives each
instrument the series that apply to it - FX 15 columns, metals/oil 14, SP500
13 - so the fleet ran as three incompatible pools:
EURUSD/USDJPY/USDCAD adopt ~57-60k peer rows each
XAUUSD/XTIUSD adopt 6.4k / 20.3k
SP500 "EVERY peer file was REJECTED, so this chart is
training alone" - 0 rows
SP500 therefore trained on 2279 independent observations against a 600-wide
input with its first layer floored at 16, printing its own "expect
overfitting" warning. It is the one chart with no pool and the worst
capacity ratio in the fleet by a factor of three.
Fresh models now pin ALTDATA_FLEET_COLUMNS - the 12-column intersection -
instead of their own file header. An existing model still adopts its .cfg
pin, so this re-keys nothing that is already trained.
Intersection rather than union: filling an absent series with its median
makes that column constant per instrument, which lets a pooled model
identify the source instrument and stop learning the shared mechanism. It
is also 6 columns narrower. Cost is six columns whose retained information
is UNMEASURED - the keep-screen reports a bitmask nothing has mapped back
to names.
2. The MI keep-screen disabled itself for the whole run on any cold start.
ReportFeatureLabelInformation set m_miReportDone on ENTRY. On a cold start
the label cache is allocated before it is filled, so BuildMiSample finds no
row carrying a resolved label and returns 0 - a sixth exit, and the only
one the
|
||
|
|
d9092a2408 |
fix(vote): persist the member's skill verdict - a converged model was ruled no-skill on every restart
SP500 resumed converged at era 136 with its tier ladder correctly restored and still swept 4999 bars reporting "0 had a snapshot, drew 0 arrow(s)" while the other five charts drew 221-312. HasDemonstratedEdge() - added with the no-skill exclusion - compares m_eraStatPrecPct against m_eraStatChancePct. Both are written once per era by EnsembleStashEraStats. A converged model runs no eras, so after a restart both sat at their -1 ctor defaults, every member was ruled no-skill, ReconstructionWeight() returned 0 for all four, and the overlay divisor was zero on every bar. Exactly the failure the WST7 ladder persistence fixed one level down: the ladder says how much a member votes, this says whether it may. RankTiersFromOos already computes the pair (pooled holdout precision and the zero-skill reference rate) and now records it as the CERTIFIED edge. That path is reached by the era end AND by the deployed replay, which is the only measurement a converged model will ever make. Persisted as WST8; HasDemonstratedEdge() prefers the era pair and falls back to it. The census line also had to be fixed: it reported "NOT ONE of those bars had a single member snapshot ... no enrolled member has published m_overlaySigSnap" for a condition that was purely a skill verdict. The snapshots were there. It now counts the two causes separately and names the one that fired. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
15b028450b |
fix(vote): follow the derived rung until a checkpoint exists, pin thereafter
A LIVE DEFECT from combining today's two changes. The threshold pins ON CHECKPOINT ( |
||
|
|
326e314b3e |
diag(features): emit the keep-set as a comparable hex mask
The keep-screen answered whether pruning is worth doing - consistently, across
all six charts:
chart kept width first-layer budget
EURUSD 17/52 624 -> 204 11.3 -> 34.3
USDCAD 19/52 624 -> 228 10.0 -> 27.2
USDJPY 18/52 624 -> 216 11.2 -> 32.4
XAUUSD 17/51 612 -> 204 7.8 -> 23.2
SP500 18/50 600 -> 216 3.8 -> 10.5
XTIUSD 16/51 612 -> 192 3.8 -> 12.1
~1 column in 3 carries the association and the rate is stable across six
independent charts - noise would not reproduce that tightly. Pruning nearly
triples the capacity budget and lifts XAUUSD off the 16-wide floor. SP500 and
XTIUSD (the two pool-poor charts) improve ~2.8x and still miss it; they need the
12-bar window cut as well, which is a separate lever costing nothing in feature
semantics and not touching pool compatibility.
Headline MI is strong everywhere under the pivot-event label: 0.008-0.0099 nats
against a ~0.002 null, strongest column 0.047-0.077 against a ~0.006 null-max
(8-13x).
WHAT THIS COMMIT ADDS is the last fact needed before a mask can be built: WHICH
columns, as a hex bitmask, so two charts' masks can be compared by eye and by
grep. Identical masks across the fleet mean ONE fleet-wide mask keeps every chart
in a single pool group; divergent masks would split six charts into six groups of
one, and pooling is the only thing currently holding the FX charts above the
capacity floor - so a per-chart prune could cost more capacity than it buys.
Still report-only. No fingerprint change, no retrain forced.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
8c1266db0b |
diag(mi): name which BuildMiSample exit abandoned the sample
The MI screen collapsed to "-1.00000 nats/feature over 0 permutations" on the
first COLD start after a wipe, taking the new per-column keep-screen with it. On
the same chart seconds earlier the auto-tuner had scored the same function fine:
auto-tune complete - 12 candidates scored, mutual information 0.00843 nats
feature/label information - -1.00000 nats/feature ... over 0 permutations
So the data exists and something between the two collapses the sample window.
Cold-start only - every successful report today came from a warm start where the
models loaded from disk, and wiping is what exposed it.
I formed three explanations (label-cache invalidation by the tuner, a shift pad
scaled off an unmeasured label resolution, a zero feature width) and each failed
against the log. Three failed explanations is the point where guessing stops and
instrumenting starts.
BuildMiSample has five distinct -1 exits and the caller can only observe the
collapsed result. Each now names itself and prints the terms that would explain
it: bars, lo/hi, MI_MIN_SAMPLES, OOS split, history window, shift pad and the
measured label resolution the pad scales from. Throttled via TCLog.
Deliberately NOT also "fixing" the latch that makes this stick
(ReportFeatureLabelInformation sets m_miReportDone at ENTRY regardless of
outcome, and the first member then sets g_ensembleChartMiReportDone, so one
failed attempt disables the screen for every member on the chart for the whole
run). If the cause is a genuine cold-start ordering problem, making it retry
would paper over it - the instrumentation decides which fix is correct.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
a9e941d7ee |
feat(features): per-column MI keep-screen (report only)
Step 1 of the prune, stopping deliberately short of pruning - two blockers make
an immediate mask the wrong move, and this is the measurement that decides
whether pruning is worth doing at all.
WHY NOT PRUNE YET:
* the screen runs with cross-asset ABSENT - its own log line says the numbers
"describe a NARROWER vector than training will use". A mask built from it
would have no evidence either way about the cross-asset block.
* a per-chart mask FRAGMENTS THE POOL. The mask must participate in the
fingerprint, and the pool only accepts peers with an identical feature
layout. Pooling is currently the only thing keeping the FX trio off the
capacity floor - the three pool-poor charts (SP500, XAUUSD, XTIUSD) are
exactly the three still floored. Six per-chart masks = six pool groups of
one, and pruning could cost more capacity than it buys.
WHAT THIS ADDS: the per-column MI was always computed inside ScoreMiSample and
thrown away except for the sum and the max. It is retained now, and the same
permutation draws that build the headline null also accumulate a PER-COLUMN null,
which is what a per-column p-value needs - distinct from the null-of-the-max,
which answers the single family-wise question "is the strongest column real".
Selection uses Benjamini-Hochberg at q=0.10, NOT the family-wise bar. FWER
controls the chance of one false positive, which is right for a verdict and far
too conservative for selection - it would discard every genuinely weak-but-useful
feature. BH bounds the expected SHARE of kept columns that are noise, which is
what a feature set cares about.
The report prints the decision in capacity units: columns kept, the resulting
input width, and the first-layer budget before and after against the 16-wide
floor. 3 of 52 is not a feature set; 45 of 52 is not worth a fingerprint re-key.
The cross-asset caveat prints itself when it applies.
Context that makes this worth doing at all: under the pivot-event label the MI
screen now reads "above the noise floor - a real association" - mean 4x the null
(p=0.005), strongest column 7.7x the null-max, excess 0.80% of label entropy,
against 1.3x / 1.15x / ~0.1% under the old label. The noise-floor verdict that
closed several earlier directions was a property of the OLD label.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
32eb5c5f58 |
feat(vote): edge-over-chance currency, no-skill exclusion, checkpoint burn-in
RETRAIN-FORCING and deliberately so. Two independent fixes for the same symptom - charts that go quiet while others overtrade. 1. THE VOTE CURRENCY IS NOW EDGE OVER CHANCE, not an absolute win rate. A tier weight is a raw win rate and a raw win rate means nothing without the chance rate behind it: 30% is strong under a 14% base rate and catastrophic under 50%, yet both entered the mean as "30". That is why the threshold needed re-tuning every time the label changed - 25 was permissive at ~70% win rates under the old direction label and a near-unanimity rule at ~30% under the pivot-event one - and why one chart's 25% was never the same statement as another's. Subtracting the member's own chance rate makes the units percentage points of demonstrated edge, comparable across charts, labels and regimes. Clamped at zero: a below-chance tier is anti-informative, and contributing negatively would act on a broken model as an inverted oracle rather than discarding it. 2. A NO-SKILL MEMBER IS NOW ABSENT, NOT ABSTAINING. Measured on XTIUSD: a Perceptron collapsed to B97/S6/N3, pooled win rate 11.5% against a 14% chance rate - worse than guessing - and still voting. Three healthy members voting Sell scored -21.06/0.77 = -27.4 and cleared; with the dead one voting Buy it became (-21.06+1.44)/0.89 = -22.0 and was BLOCKED. It vetoed its own ensemble on ~95% of bars, and that WAS the chart's 3.3% coverage. Neither existing guard caught it: it IS self-ranked and its tier weights were 11-14. The fix has to remove it from the DIVISOR, not just the sum - an abstainer contributes weight by design, so zeroing only the contribution makes the dilution worse. VoteCapableWeight() already means exactly "may this member's weight sit in the denominator", so the skill test belongs there. ReconstructionWeight() and the OOS scorer's divisor move with it or the scorer certifies a vote live does not cast. The skill test reads the PREVIOUS era's measurement - gating this era's vote on this era's own outcome would be circular. 3. CHECKPOINT BURN-IN (ENSEMBLE_CHECKPOINT_MIN_ERA 20). XAUUSD deployed the checkpoint from ERA 2, XTIUSD from ERA 4, each after 69 and 65 further eras failed to beat it. Ensemble coverage measures AGREEMENT, and four models that have barely moved off their initialisation agree almost by construction - so coverage is inflated exactly when the models know least and decays as they differentiate (XAUUSD 6.6% at era 8 -> 0.4% at era 75). Since selectionScore is precision discounted by coverage, an early era outscores every mature one and the ladder freezes on it. INTENDED CONSEQUENCE: a chart whose MATURE coverage cannot clear the floor now refuses to deploy rather than shipping era-2 weights. Fewer deploys, honest ones. Burn-in eras are also kept out of g_ensCandidateEras (they could not have won, so counting them inflates the family-wise N and raises the bar for nothing) and out of g_ensErasSinceBest (or the run reaches "no better vote for N eras" with no best to beat, exhausting the escalation ladder before the first era may compete). Every pinned threshold and .stats record is in the OLD currency and is now meaningless - this forces a fresh start on its own. Nothing needs re-tuning because the threshold is DERIVED: the sweep re-picks the rung by itself. Compiled clean; NOT yet run. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
d059780c22 |
fix(io): stage atomic writes to a PER-CHART temp, not a shared one
A DATA-INTEGRITY BUG, pre-existing, surfaced by the clearer failure message in
|
||
|
|
869cd1b40c |
fix(pool): defer the atomic promotion to the timer instead of spinning on the tick
REPLACES the in-line retry from
|
||
|
|
72dba892cc |
fix(chart): configure the vote-arrow layer with the PINNED threshold
Second instance of the same regression |
||
|
|
b1c3a898aa |
fix(persist): adopt the pinned threshold on load; trim the accuracy label
THE REGRESSION, mine, from |
||
|
|
ad4ae58814 |
feat(vote): exit-on-reversal boolean, pin the threshold, retry the atomic rename
THE EXIT KNOB. Exit_On_Reversal_Vote (default false) replaces the deleted Signal_ThresholdClose with one boolean: false pins the close threshold to an arithmetically unreachable 101, true pins it to the SAME threshold the entry uses - the seed at first, then the derived value, republished together whenever it moves. A second threshold was always redundant; "the bot now says the other way" is one question. It also arms CExpertSignalCustom::m_holdToBarrier, which was DEAD CODE: HoldToBarrier(bool) had no caller anywhere in the build, so the flag had been permanently false and the disabled close threshold was carrying the whole hold-to-barrier policy alone. Both halves now move together. Default stays false because the reason is statistical: the gate certifies P(label agrees | vote fired) against a label that runs to the barrier, so an early close trades something never measured. Turning it on is a different strategy, not a tightening of this one. THE PIN. The live threshold now moves only when an era's weights become the checkpoint, and freezes once g_ensDeployApproved. Every era still derives its own rung - that is how the best one is found - but the rung that TRADES belongs to the checkpoint, exactly as the weights do. Two reasons, one measured and one structural: the per-era rung moves on 6-34% of steps (the live run flapped SP500 15 -> 10 -> 15 within a minute of starting), and without the pin a later era's rung could end up applied to an earlier era's deployed model. A ladder restart releases the pin, since clearing the checkpoint clears what it pinned. The era line now prints the rung its own numbers came from, so it stays honest when that differs from the pinned one. THE ATOMIC RENAME retried zero times. Six charts share the TrainPool and AltData directories, so a publish regularly lands while a peer chart holds the destination open and FileMove returns 5004 - 27 times in one day on the live fleet. Nothing was lost (the temp keeps the new content, the old file stays intact) but the row did not update until the next publish. Now four attempts at 25ms, on the FAILURE PATH ONLY - a successful rename never sleeps - and skipped in the tester, where the contention cannot happen and Sleep would distort a pass. A rescued retry is logged, so worsening contention is visible. Retrain-neutral. Compiled clean; NOT yet run. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
34f1e09372 |
feat(magic): assign the magic number once, then remember it
Expert_MagicNumber = 0 (the new default) means "draw one and write it down".
On first attach the EA picks a random magic in a distinctive band, persists it
to MQL5\Files\Warrior_<symbol>_<period>.magic, and reads that same value back on
every later start. Unique without anyone typing it, and STABLE.
Stability is the whole point. The magic is how the EA recognises its own
positions - a fresh one per start would leave every open position invisible to
the scheduled close-all, the risk-budget flatten and the journal's MAE/MFE walk:
trades still running that no code would ever manage again. So the value is
persisted before it is ever used to trade.
Stored TERMINAL-LOCAL rather than in Common\Files\Warrior_EA, on purpose: that
folder is the one wiped for a retrain, and positions outlive retrains. It also
gives two terminals on the same symbol different magics, which a chart-identity
hash could not.
Fallbacks, both of which stay stable without a file:
* tester/optimizer/forward use a magic derived from chart identity, so two
identical passes cannot differ.
* an unwritable file falls back to that same derived value, and says so.
Books occupy EVEN slots only, so one chart's short book (base+1) can never land
on another chart's long book.
WarriorOwnsMagic() now also recognises the legacy 2024/2025 pair permanently.
Without it, switching an existing chart to 0 while a position was open would
orphan that position. Every caller also matches the symbol, so claiming those
values can only reach positions on this EA's own chart.
Existing charts are untouched: MT5 stores inputs per chart, so the six live
charts keep the 2024 they already have and keep managing what they hold.
Compiled clean; NOT yet run.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
17270ab308 |
feat(trade): two books per symbol, and delete the vote exit
Allow_Hedging (default ON, live only on a RETAIL_HEDGING account) gives the EA an independent long book and short book on its symbol: at most one long and at most one short, each opened on its own side's vote and each held to its own barrier. On a netting account, or with the input off, the original single-position path runs bit-for-bit unchanged and init says which one is live. WHY THIS INSTEAD OF A VOTE EXIT. The deploy gate certifies P(label agrees | vote fired) and the label runs to the barrier, so closing early on a reversal makes the realised outcome stop being the labelled one - the certified precision no longer describes what is traded. Opening the other side acts on the new signal and leaves the old position's certification intact, and costs no more than reversing: both pay the new side's spread, the difference is only that the existing position runs on to a barrier already measured as positive-expectancy. So Signal_ThresholdClose is DELETED rather than tuned, along with its SIGNAL_CLOSE_PRESETS enum; the threshold is pinned to an arithmetically unreachable 101 (the stock default of 100 is reachable by a weighted mean of values capped at 100). Note the two books can never both fill from one signal: CheckOpenLong and CheckOpenShort test opposite signs of the same m_direction, so at most one clears per tick. A hedge only forms when a LATER opposite vote fires - which is what keeps it from being a guaranteed-loss wash pair. The mechanism is a SelectPosition() override keyed on the active book's magic; every inherited close/trail path then operates on that book untouched. The long book keeps Expert_MagicNumber, so no existing position, journal row or risk-budget state file is re-addressed. Short book is +1. Four ownership filters had to widen from "== m_magic" to WarriorOwnsMagic(), or the short book would have been invisible to the code that must reach it: the scheduled close-all (positions and orders), the risk budget's emergency flatten, and the journal's MAE/MFE walk. WarriorOwnsMagic() is deliberately NOT gated on Allow_Hedging - turning the input off while a short-book position is open would otherwise orphan it with nothing left to close it. Risk sizing needed no change: CapRiskAmount already subtracts OpenRiskAtStops(), which counts every position regardless of magic, so the second book is sized inside what the first one left. Conservative for a hedged pair, which cannot lose both stops - the safe direction. Retrain-neutral: neither input is in BuildModelFingerprint() or ComputeDbConfigFingerprint(). Compiled clean; NOT yet run. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
c6eb9085d5 |
feat(vote): derive the threshold instead of configuring it
Signal_ThresholdOpen becomes a seed. The era verdict now picks the HIGHEST
sweep rung whose vote still clears the whole deploy gate - coverage floor,
exact-binomial precision bar and two-sidedness together - computes the era's
verdict AT that rung, and publishes it to the live signal's m_threshold_open
so the bar the gate certifies is the bar the EA trades.
Measured on 619 era verdicts across all six live charts:
* every era on every symbol had at least one rung clearing the full gate.
At the fixed 25% the fleet was actually running, four of six symbols had
none, ever. The threshold, not the models, was the blocker.
* walk-forward (rung derived on era N, scored on era N+1): 10.2% coverage /
31.8% precision, against an oracle re-picking on N+1 of 10.3% / 31.7%.
Near-zero shrinkage - a measurement, not a fit. It holds because the
binding constraint is COVERAGE, a near-deterministic step function of the
vote distribution, not precision.
* vs a fixed 15% (best global value): +0.6pp precision, 3.4pp less coverage.
vs a fixed 20%: deployable on all six rather than four of six.
Selection on the highest PASSING rung, never on the best-precision rung - that
is a best-of-6 on a noisy statistic and this project has crowned noise that way
four times. The multiplicity that remains is paid for: nTried in
EnsembleSurvivesSelection is now eras x rungs. Costs nothing - all six charts
clear it by 6.5-12 sigma even forming z on effective rather than raw calls.
Also fixes, in the same path: the direction-policy gate is hoisted above the
per-rung tally so every rung is scored on the population the gate certifies.
Retrain-neutral: not in BuildModelFingerprint(), no .nnw re-keyed.
Compiled clean; NOT yet run.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
51620fe9e0 |
fix(vote): Signal_ThresholdOpen 25 -> 15, from the sweep's own numbers
The quorum model that said 20 was wrong. It reasoned from the vote's quantisation - 4 members at tier weight ~30, so 3-of-4 agreeing gives 22.5 and PCT_20 admits it - and simultaneous agreement turns out to be rarer than a per-member coverage of ~27% implies. Measured on real rows by the threshold sweep added in the previous commit: symbol floor 15% cov/prec 20% cov/prec 25% cov/prec (active) EURUSD 6.7% 15.7 / 32.6 12.5 / 33.7 4.1 / 35.0 SP500 6.9% 9.8 / 32.7 4.0 / 34.4 X 1.0 / 34.7 X USDCAD 7.2% 18.9 / 32.2 12.8 / 32.6 4.0 / 34.7 X XAUUSD 6.7% 12.8 / 29.2 4.4 / 33.9 X 1.0 / 26.5 X XTIUSD 6.9% 11.4 / 34.3 7.8 / 35.7 1.7 / 38.4 X 15 clears the coverage floor on every symbol; 20 fails SP500 and XAUUSD; 25 fails all of them. The precision surrendered is about 2pp, because precision is nearly flat across these rungs while coverage moves 10-20x - the high thresholds were buying almost nothing for the coverage they cost. The rule this encodes is: take the cheapest rung whose coverage clears the floor, not the best precision. Precision above the bar earns nothing extra; coverage below the floor makes the era undeployable regardless. Not in BuildModelFingerprint(), so every trained .nnw survives. DOES NOT MOVE THE RUNNING FLEET. MT5 stores input values per chart in profiles\Charts\*\chart*.chr, so this only takes effect on a fresh attach; the live charts have to be changed in each one's EA properties. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
b9da557e4e |
diag(gate): report what the vote would score at every threshold rung
The gate could say "coverage too low" but never "and here is what it would be one rung down", so the single parameter most responsible for a refusal was the one its own output said least about. Working it out by hand needed a model of the vote's quantisation (a weighted mean of member tier weights, so the threshold is really a quorum) and that model could not be checked: MT5 stores the input PER CHART in profiles\Charts\* \chart*.chr, so an already-attached EA ignores a changed source default - confirmed by a full close/recompile/relaunch after which the log still read "fired at vote>=25%". There was no cheap A/B available. Each era now reports coverage and precision at every PERCENTAGE_PRESETS rung from 5% to 30%, measured on the same rows the verdict just scored, marking the active rung and any rung that clears the coverage floor. It is accumulated before the live threshold test so the sweep sees every scored row, and gated by the same direction policy so its numbers are comparable with what the gate certifies. Nothing reads it to decide anything. Motivation, measured overnight across 534 eras with zero runtime errors: every symbol clears its precision bar and every symbol fails on coverage (0.0-3.3% against a ~6.7-7.2% floor), while the members stay healthy throughout at 22-27% precision against a 13-14% chance rate on 25-38% of bars. Only the aggregation fails. SP500 was DEPLOYABLE at era 5 with 7.5% coverage and sits at 0.7% by era 536 with precision unchanged - more training is proven not to help, because a 25% threshold against ~30 tier weights demands unanimity and the models diverge as they specialise. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
1533365a85 |
diag(gate): the coverage refusal contradicted itself
The message I added one commit ago printed, verbatim: "Precision was 28.4% against a 39.1% bar, so the calls it DID make were NOT good enough: the vote is too selective, not too weak." Those two clauses say opposite things. Only the "NOT" was conditional; the diagnosis after the colon was hardcoded, so whenever precision missed its bar the line asserted and denied the same thing in one sentence. The two cases are opposite diagnoses and must not share a sentence: - Precision CLEARED its bar -> the calls were good and there were too few of them. The vote is too selective. - Precision MISSED its bar -> this is still not "the model is weak", because the exact-binomial floor is computed from the INDEPENDENT call count, so thin coverage inflates the very bar it is judged against. Reporting that as a second, separate failure sends a reader off to fix the model when coverage is what moved the target. Caught by reading the diagnostic's own first live firing rather than by review - the same way the two regressions before it were found. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
0f756faf2a |
diag(gate): name the deployability condition that actually failed
The stage-3 refusal read "no era's combined vote ever cleared the deployability floor" and then listed all three conditions in one parenthesis - fires on a quarter of the base rate, both directions alive, precision above the reference by 2 sigma - without saying which one fired. The three have nothing in common as fixes, so the list was not a diagnosis. It cost real time to work out by hand tonight, and the answer was coverage every time. Keeps the best era's coverage, its floor and its precision bar alongside the win rate already retained, and names the failing condition. The coverage branch also states whether the calls it DID make cleared the precision bar, because "too selective" and "too weak" are opposite problems that the old message could not distinguish, and points at Signal_ThresholdOpen being a quorum rather than at the models. Cleared at both existing reset sites so a refusal can never describe an era that is no longer the best. Context: SP500 reached stage 3 at era 67 and was refused on coverage 0.5% against a 6.9% floor while its precision was 62.5% against a 61.4% bar - i.e. the vote was too selective, not too weak. Same doctrine as CTrainPoolReader::Announce's reject list. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
2e22e714c6 |
fix(pool): length-prefix the fingerprint - the cross-instrument pool was inert
STrainPoolHeader wrote its fingerprint into a FILE_BIN stream as FileWriteString(h, fingerprint + "\n") and read it back with FileReadString(h) - no length argument. In binary mode FileWriteString emits the characters raw: no length prefix, no terminator, and "\n" is just another character rather than a delimiter anything honours. The reader had nothing to stop at, over-read into the float rows that follow, and returned the fingerprint plus a few bytes of binary garbage - so `fingerprint != wantFp` could never succeed between two genuinely identical models. Verified in the bytes rather than inferred: xxd on a v1 file shows three ints then the fingerprint starting immediately at offset 12 with no count in front of it, and EURUSD/USDJPY/USDCAD all stored width 624 with byte-identical fingerprints while each one's log rejected the other two as "different model fingerprint". The StringReplace on "\n" is the tell that a delimiter was intended. Cross-asset-class peers really are incompatible and always will be - FX majors carry XA:6, indices/metals/oil carry XA:6:IDX2, giving widths 600/612/624 - which is why the reject list looked plausible and this went unread. The three FX majors were always poolable and never pooled. Length-prefixes the string, bounds-checks the count before sizing a read from it, and bumps TRAINPOOL_RECORD_VERSION 1 -> 2 so existing files are refused by the version gate with a reason instead of being misread. Also documents, without changing, why Signal_ThresholdOpen is now a unanimity rule: the vote is a weighted mean of tier weights, those fell from ~70 to ~30 with the pivot-event label, so PCT_25 went from ~36% of the reachable ceiling to ~83%. Measured: all 6 symbols clear their precision bar, 4 of 6 fail only on coverage, and coverage decays 6.8% -> 2.2% over 35 eras as the models specialise - which shrinks effN and so RAISES the deploy bar at flat precision. PCT_20 (a 3-of-4 quorum) is the indicated change but is left unmade: MT5 stores input values per chart in profiles\Charts\*\chart*.chr, so an already-attached EA ignores this default entirely - confirmed by a full close/recompile/relaunch cycle after which the log still read "fired at vote>=25%". Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
994fe3899c |
feat(label): pivot-EVENT target replaces direction-to-next-pivot
The old target asked "which way is the next pivot", which every bar of a
~13-20 bar leg answers identically - so the net could not tell a fresh turn
from mid-trend and learned the prevailing direction instead. Its own
zero-skill reference showed it: chance sat at 56/44, i.e. the label WAS the
drift, and the gate's standing warning ("a model that only reproduces it has
found the drift, not an edge") applied to the target itself.
Buy now means a swing LOW commits within PIVOT_LABEL_TOLERANCE_BARS bars,
Sell a swing HIGH, Neutral no turn that close. Pivot type is read from
ZigZagBuffer[p] == Low[p], exact by construction in ZigZag.mq5. The existing
P1-final-once-P2-commits rule is kept and now also settles the NEGATIVE
verdict, so the Neutral majority is permanent rather than provisional.
Measured on a full fresh run, all 6 charts:
class balance 56/44/~0 -> 13.7/13.7/72.6 (imbalance 5.3:1)
label overlap ~31 bars -> 5 bars
independent obs 368-1086 -> 2331-7032
weights/obs 9.2-26.2 -> 1.1-4.2
coverage 100% of bars -> 17-48%
23 of 24 models fire all three classes at precision 18-32% vs 13-15%
chance; SP500's ensemble reaches DEPLOYABLE (32.3% vs a 24.0% bar).
Two bindings had to move with the label:
- The capacity deflator. m_swingLifespan fed EstimatedInSampleBars() as
raw/31, measured from the legs. Overlap is now a property of the LABEL -
one turn is callable by exactly the tolerance window - so it is the
window, not a leg measurement. Missing this would have kept every model
sized for a sixth of its real evidence.
- A dormant cold-start seed. Labels.mqh seeds the output bias toward the
dominant class above COLD_START_SEED_MIN_DOMINANCE (0.70); at 56/44 it
never armed, at 72.6% Neutral it does - writing a fixed +-3.0 against a
true prior spread of ~1.75, which would start every net predicting Neutral
~95% of the time. Now seeds the measured log-prior, zero-centred and
capped by the same guard rail the logit adjustment uses (Lin et al. 2017).
TGT:SWG1 -> TGT:PVT1:<tolerance>, with the window in the token because it is
part of the label: every .nnw is invalidated and the fleet retrains.
Depth is still gated, and now for a precise reason: the first dense layer
stays at FIRST_LAYER_MIN_WIDTH because budget = effN/(inputWidth+1) is 11.2
at input 624. Reaching the next rung needs inputWidth <= ~218, i.e. feature
pruning - not architecture.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
6adb710a79 | fix(binomial): correct tail calculation in BinomialUpperTailP and add tests for accuracy |