refactor: split CExpertSignalAIBase implementation by responsibility
ExpertSignalAIBase.mqh was 8216 lines: the class declaration followed by 87
method bodies covering training, labelling, feature extraction, persistence,
chart drawing, online learning, the GA auto-tuner and inference, all in one
file. Train() alone is 1492 lines; a change to arrow drawing meant scrolling
past the era loop.
Moved the bodies into Expert\AIBase\, included at the bottom of the original
after the class declaration:
Training.mqh 1607 era loop, plateau ladder, checkpoint select, deploy
Features.mqh 1093 indicator creation + per-bar input feature vector
ChartUI.mqh 634 arrows, arrow persistence, status panel, cleanup
Persistence.mqh 492 .stats/.cfg sidecars, CPU-inference validation, copy
OnlineLearning.mqh 461 live continual learning, EMA shadow, OOS simulator
Labels.mqh 309 ZigZag pivot labels, async label-cache prebuild
AutoTune.mqh 275 genetic tuner (population, crossover, halving)
Inference.mqh 235 softmax, prior calibration, class priors
ExpertSignalAIBase.mqh 8216 -> 3131 (declaration + topology build only)
This is a pure relocation - verified mechanically, not by eye: HEAD's file
reconstructed from the eight partials plus the surviving remainder is
byte-identical to HEAD, span for span (scratchpad verify_split.py). No
declaration moved, no signature changed, no code rewritten, so behaviour is
unchanged by construction.
Compiles 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 00:42:45 -04:00
|
|
|
//+------------------------------------------------------------------+
|
|
|
|
|
//| Warrior_EA |
|
|
|
|
|
//| AnimateDread |
|
|
|
|
|
//| |
|
feat(ai): triple-barrier labels replace exact-pivot ZigZag targets
The 31:1 class imbalance was self-inflicted by the TARGET, not a property
of the market. Labelling only the exact bar where a ZigZag pivot confirms
gave Buy 1164 / Sell 1164 / Neutral 35841, and every correction mechanism
this codebase accumulated sits downstream of that one choice: the
logit-adjusted loss and its range cap, the prior EMA, the +-3.0 output-bias
seed, balanced-accuracy-then-precision selection with its coverage floor,
the recall floor and its catch-22, the alternation gate, NMS, and the four
oversampling designs that collapsed before them.
The reference this engine is built on (references/neuronetworksbook.pdf
ch. 3.1/3.3) also uses ZigZag, but targets the DIRECTION TO THE NEXT
EXTREMUM on every bar - ~50/50 by construction, with no imbalance to
correct at all. It never had this problem because it never asked "is this
the pivot bar".
Labels are now the triple barrier (Lopez de Prado ch. 3), using the EA's
OWN SL_Mode/TP_Mode: does a trade opened at this bar's close reach its
target before its stop, within a horizon. Buy = long resolves, Sell =
short resolves, Neutral = neither. Consequences:
- dir-precision in the era line stops being a proxy and becomes the win
rate of the strategy under its own exit rules.
- Expected balance ~25/25/50 at the shipped 1:3 (gambler's ruin), i.e.
~2:1 instead of 31:1. Measured and logged at the end of the prebuild.
- Spread is charged on both legs, so it is a NET win rate.
- Intrabar ambiguity resolves to the STOP. OHLC cannot order two touches
inside one bar and the optimistic reading is how a backtested edge
becomes a live loss.
ZigZag stays as input features (EnableSwingContext) and now also supplies
the vertical barrier: the horizon is the median confirmed leg length,
snapped to a coarse ladder. Derived, not configured, and deliberately kept
out of the filename fingerprint - a filename keyed on a measured quantity
orphans a trained model the moment the measurement moves.
Removed, because the premise died with the old target:
- the alternation gate. Correct for pivot labels (a ZigZag cannot emit two
same-type pivots in a row, so a repeat was provably a false fire), and
wrong for barrier labels, which answer each bar independently. It also
took its worst consequence with it: a one-sided model previously got ONE
trade per backtest, a hard blocker on marketplace validation.
- SignalClusterWindow now defaults off - it de-duplicated repeats that are
now real trades. Kept as an opt-in display control.
- LABEL_WINDOW_BARS, the pivot-widening pass, ConfirmedZigZagLabel.
- the era-0 output-bias seed now needs a genuinely dominant class (0.70)
rather than 0.40; at ~50% Neutral a +-3.0 seed is a distortion, not a
correction.
Also fixed, both found while wiring the above:
1. RefreshConvergedSignal sized its buffers from a date delta
(Bars(sym, period, dtStudied, TimeCurrent())). dtStudied is a training
watermark; in the tester it is loaded from a live-chart save AHEAD of
the simulated date, so the interval inverted, Bars() returned ~0, and
the buffer came out at exactly m_historyBars - deep enough for the OHLC
window and far too shallow for the Donchian-50 / 20-bar-return / SMA
extension behind it. Inference silently computed DIFFERENT features
from the ones training learned on, live as well as in the tester. Now
sized from what the feature builder actually needs.
2. The barrier horizon is resolved on the deployed path too. A deployed
model never enters Train(), so it never reached the prebuild, and
OnlineLearnStep reads the horizon as its confirmation delay - left at
the fallback it would have backpropped bars whose barriers had not
resolved. Silent lookahead in the one place that writes to a live model.
SL_Mode/TP_Mode join the weights fingerprint: they define the labels now,
so a model trained at 1:3 must never be silently reused at 1:1. This
re-keys every pre-existing model by design - none were trained on this task.
Inference census extended with the vote gate. LongCondition/ShortCondition
open with a readiness check the refresh counters never see; in the tester it
reduces to "the seeded _optcache.nnw must have LOADED", and if it did not,
every vote is hard-zeroed while the model still answers Buy. The old three
counters would have read that as "the model says Neutral" - false, and a
completely different fix. This is the leading candidate for the
zero-direction backtest and the census can now name it in one run.
Both builds compile 0 errors / 0 warnings. Forces a full retrain.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 20:39:49 -04:00
|
|
|
//| Triple-barrier labelling and the async label-cache prebuild. |
|
refactor: split CExpertSignalAIBase implementation by responsibility
ExpertSignalAIBase.mqh was 8216 lines: the class declaration followed by 87
method bodies covering training, labelling, feature extraction, persistence,
chart drawing, online learning, the GA auto-tuner and inference, all in one
file. Train() alone is 1492 lines; a change to arrow drawing meant scrolling
past the era loop.
Moved the bodies into Expert\AIBase\, included at the bottom of the original
after the class declaration:
Training.mqh 1607 era loop, plateau ladder, checkpoint select, deploy
Features.mqh 1093 indicator creation + per-bar input feature vector
ChartUI.mqh 634 arrows, arrow persistence, status panel, cleanup
Persistence.mqh 492 .stats/.cfg sidecars, CPU-inference validation, copy
OnlineLearning.mqh 461 live continual learning, EMA shadow, OOS simulator
Labels.mqh 309 ZigZag pivot labels, async label-cache prebuild
AutoTune.mqh 275 genetic tuner (population, crossover, halving)
Inference.mqh 235 softmax, prior calibration, class priors
ExpertSignalAIBase.mqh 8216 -> 3131 (declaration + topology build only)
This is a pure relocation - verified mechanically, not by eye: HEAD's file
reconstructed from the eight partials plus the surviving remainder is
byte-identical to HEAD, span for span (scratchpad verify_split.py). No
declaration moved, no signature changed, no code rewritten, so behaviour is
unchanged by construction.
Compiles 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 00:42:45 -04:00
|
|
|
//| |
|
|
|
|
|
//| PARTIAL IMPLEMENTATION FILE - not standalone. |
|
|
|
|
|
//| This holds CExpertSignalAIBase method BODIES only. The class |
|
|
|
|
|
//| declaration lives in Expert\ExpertSignalAIBase.mqh, which |
|
|
|
|
|
//| #includes this file at the bottom, after the declaration. Do not |
|
|
|
|
|
//| include it anywhere else and do not compile it on its own. |
|
|
|
|
|
//| |
|
|
|
|
|
//| Split out purely to make the 8216-line original navigable; the |
|
|
|
|
|
//| code inside was moved verbatim, not rewritten. |
|
|
|
|
|
//+------------------------------------------------------------------+
|
|
|
|
|
#ifndef WARRIOR_AIBASE_LABELS_MQH
|
|
|
|
|
#define WARRIOR_AIBASE_LABELS_MQH
|
|
|
|
|
//+------------------------------------------------------------------+
|
|
|
|
|
//| (Re)sizes the label AND feature caches and clears them if `bars` |
|
|
|
|
|
//| (or the now-relative index frame) has changed since the last |
|
|
|
|
|
//| build - see the member declaration comments for why this is the |
|
|
|
|
|
//| correct invalidation trigger. Returns true if a rebuild happened. |
|
|
|
|
|
//+------------------------------------------------------------------+
|
|
|
|
|
bool CExpertSignalAIBase::EnsureBarCachesCapacity(int bars)
|
|
|
|
|
{
|
|
|
|
|
if(bars == m_labelCacheBars && m_Time.GetData(0) == m_labelCacheAnchorTime)
|
|
|
|
|
return false;
|
|
|
|
|
ArrayResize(m_labelCacheBuy, bars);
|
|
|
|
|
ArrayResize(m_labelCacheSell, bars);
|
diag: is "optimal SL/TP" learnable? Score the features against excursions
Proposed direction: train the net to predict entry/SL/TP that maximise return
and minimise drawdown, rather than to classify direction. Before rebuilding a
head, measure whether the target is learnable at all.
That question splits into two that behave nothing alike:
HOW FAR price travels (MFE/MAE) - essentially volatility, and volatility
clustering is about the most robust regularity in markets.
WHICH WAY it goes first (the asymmetry) - direction, which is what every
noise-floor verdict in this project has been about.
Expectancy comes ONLY from the second. The first buys position sizing and
drawdown control - worth having under prop-firm limits, but not an edge: exit
management on RANDOM entries already moved the payoff ratio 0.92 -> 5.72 with
expectancy FLAT.
Crucially this is NOT already answered. Every MI figure here scored the
triple-barrier label, i.e. one specific question at one fixed geometry. A
noise-floor result there says nothing about whether excursion MAGNITUDE is
learnable - different target, different answer.
Four targets, and the verdict is the CONTRAST, printed explicitly because the
dangerous misreading of "UP clears" is "we can predict profitable trades":
RANGE (up+dn) - realised volatility, included as a POSITIVE CONTROL that
SHOULD clear. Every prior verdict here lacked a control
expected to pass; a range target at the floor indicts the
measurement, not the market.
UP / DOWN - MFE / MAE.
ASYMMETRY - up-dn, the only one that can pay.
Collected inside the walk the label already does (one max, one min per bar).
The early-out when both barriers resolved is GONE: it would have truncated the
excursions at whichever bar tripped the last barrier, making the measurement a
function of the CURRENT SL/TP - the circularity this is trying to escape. The
loop was already bounded by the horizon, so only the average cost moves.
Discretised into 3 EQUAL-FREQUENCY bins, so every downstream piece (block
permutation, null, p-value) is reused unchanged. Equal-frequency because MFE is
fat-tailed and fixed-width bins would put nearly every row in bin 0; it also
pins H(Y) at ln(3)=1.099 for all four, making them comparable to each other and
to the barrier label's ~1.02 instead of confounded by class balance.
Two bugs fixed in this code before it ever ran, both of which would have
produced a plausible quiet wrong answer rather than an error:
- TripleBarrierLabel early-returns on invalid ATR/close BEFORE the point the
accumulators were reset, so one bar's excursions would be cached under
another bar's index. Cleared at the top now, ahead of every return.
- An unresolvable bar is still flagged as labelled but carries excursions of
exactly 0. Under equal-frequency binning a block of identical zeros drags
the lowest cut onto zero and a third of the sample lands in one
uninformative bin - a depressed score that reads as "not predictable", a
false negative in the direction that would wrongly kill the idea. Rows
where both excursions are zero are dropped; price cannot travel zero both
ways over a whole horizon.
Read-only diagnostic. No topology or label change: no retrain of its own.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-07 10:22:41 -04:00
|
|
|
//--- Sized with the label caches they share a validity flag with, so the three can never disagree
|
|
|
|
|
//--- about how many bars they cover.
|
|
|
|
|
ArrayResize(m_excUpCache, bars);
|
|
|
|
|
ArrayResize(m_excDownCache, bars);
|
feat: first-passage ladder + expectancy scan - price every geometry, not just the chosen one
Corrects the premise of the previous plan. Break-even is NOT a ceiling.
If the model shifts the win probability on the bars it selects from
p0 = m/(m+k) to p0 + d, then
EV = (p0+d)*k - (1-p0-d)*m = d*(k+m)
because p0*k - (1-p0)*m is zero by construction. The stop:target RATIO
is expectancy-neutral - a punishing break-even is exactly repaid by the
payoff - and only the real edge d and the TOTAL WIDTH (k+m) move EV.
Width matters because the spread is charged once per trade however wide
the barriers are, so a narrow barrier spends much of its own range on
costs. DeriveBarrierGeometry's own comment already said the ratio buys
nothing; the objective just never followed from it.
Blocker this had to solve first: m_excUpCache/m_excDownCache hold only
MAXIMUM travel each way, and a maximum cannot say which side was
reached FIRST - so any geometry other than the walked one was
undecidable on precisely the bars where both barriers were touched,
~28% of the sample.
- BARRIER_LADDER: per bar, the first-touch AGE for 8 travel distances
in each direction, filled during the walk the labels already run.
Cursors keep it O(1) amortised per walked bar rather than 16
comparisons. Levels are travel FROM ENTRY, not barrier prices, so one
ladder serves both directions and the spread is applied analytically
when a level converts back to an SL/TP multiple - storing prices
would need four ladders and bake today's spread into the cache.
Sized, invalidated and validity-gated with the label caches.
- ReportGeometryExpectancyScan: every ladder pair priced exactly off
that cache - width in ATR and in SPREADS (cost efficiency, knowable
without knowing d), break-even, both base rates, the share of bars
resolved inside the horizon, and EV per unit of edge. Compares the
widest resolvable pair against the quantile rule's pick.
MEASUREMENT ONLY - the quantile rule still chooses. Nothing here can
measure d, and width buys nothing if the wider target is less
predictable. Base rates are printed beside each break-even because a
persistent gap is DRIFT and must not be credited to the model.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-10 12:59:18 -04:00
|
|
|
//--- Same lifetime and the same validity flag as the excursion caches beside them - see BARRIER_LADDER.
|
|
|
|
|
ArrayResize(m_ladderUpAt, bars * BARRIER_LADDER_COUNT);
|
|
|
|
|
ArrayResize(m_ladderDownAt, bars * BARRIER_LADDER_COUNT);
|
|
|
|
|
ArrayInitialize(m_ladderUpAt, 0);
|
|
|
|
|
ArrayInitialize(m_ladderDownAt, 0);
|
fix: the deploy gate was benchmarking a win rate against a label frequency
The gate rests on an invariant stated at ExpertSignalAIBase.mqh:199 - under a
driftless walk P(touch +k before -m) is m/(m+k), and break-even for a k:m trade
is ALSO m/(m+k), so "beats chance" and "is profitable" are the same test.
That invariant needs reward >= risk, and the measured geometry no longer
satisfies it. With target 1.62*ATR and stop 3.33*ATR, break-even is 67.3%, but
both-won bars were stripped out of Buy and Sell so the label base rate read
37.5%. chancePrecPct is max(BuyTotal,SellTotal)/bars, so the gate was clearing
models nearly 30pp short of break-even: 42% "directional precision" is +4 sigma
against 37.5% and loses money on every single trade. Live since 217b9bc.
Root cause is that label agreement stopped being the same question as trade
profitability. Buy implies winLong, but the converse fails on every both-won
bar, and the label can only name one of two directions that both pay.
So stop asking the model whether it matched a label and start asking whether
its trade paid:
- cache winLong/winShort per bar beside the label, under the same validity
flag; published from the barrier walk before the collapse to 3 classes
- dirPrecPct now counts wins on the side actually called
- chancePrecPct is max(P(winLong), P(winShort)), MEASURED - the textbook
m/(m+k) would credit SP500's drift to the model
- the NMS "what would I have made" pair, the live-fired precision, and the
IS/OOS cumulative win rates all move to the same test. IS and OOS are read
side by side as the overfitting signal, so measuring one in wins and the
other in agreement would put a fixed gap between them that has nothing to do
with generalization
- the confidence threshold is FITTED on wins too, so the operating point
maximises what the gate grades
- per-class label-agreement precision is still computed and logged; it is the
right diagnostic for class separation, just not for a deploy decision
- era line renamed dir-precision -> win-rate, chance -> chance=break-even
Both build variants compile 0 errors, 0 warnings.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 16:00:51 -04:00
|
|
|
ArrayResize(m_winLongCache, bars);
|
|
|
|
|
ArrayResize(m_winShortCache, bars);
|
feat(gate): grade OOS calls on the exit policy actually in force, and move vote combining out of the members and into the orchestrator
Option (a) from the exit-policy question: the certified number must be the traded
number. Plus the modularity correction the user called for on 778b6c0.
1. VOTE COMBINING BELONGS TO THE ORCHESTRATOR, NOT TO A MEMBER.
778b6c0 fixed the last-writer-wins bug on g_LiveAISignedConfidence by having a
member average its siblings through g_warriorEnsemble. That trades a scheduling bug
for a coupling bug, and it is the wrong shape for this EA: every signal runs in its
own instance, minds its own state, and VOTES to the orchestrator, which is the only
thing allowed to combine opinions.
Replaced with a publish/aggregate pair in Variables\ConfidenceBridge.mqh whose split
is enforced by shape rather than by convention:
- PublishAIVote(slot, conf) - a member writes ONLY its own slot, reads nobody's;
- AggregateAIVotes() - called by CExpertSignalCustom::LiveSignedConfidence.
CExpertSignalAIBase::EnsembleLiveSignedConfidence is gone. The orchestrator also
republishes the aggregate into g_LiveAISignedConfidence, because the intelligent
trailing reads that global directly and must act on the same number the exit route
does rather than on a leftover from whichever member ticked last. A solo AI signal
owns slot 0, so the non-ensemble path is unchanged.
2. THE GATE NOW REPLAYS THE REAL EXIT RULE.
SimulateTradeOutcome() walks the same price series with the same fill/barrier/spread
convention as ComputeLabelForBar - deliberately by copy, so a disagreement between
the two can only be a policy effect and never a discrepancy between two pieces of our
own arithmetic - and terminates at the FIRST of stop / target / vote reversal /
horizon. Barriers are tested before the vote on the same bar: intrabar we cannot know
which came first, and the barrier is what the broker executes automatically, so
checking the vote first would credit the exit policy with escapes a real stop would
have taken out of its hands.
It runs AFTER pass 3, not inside it. A vote-flip exit for a trade entered at bar r is
decided by the model's output at bars r-1, r-2, ... - NEWER bars - and pass 3 walks
oldest-to-newest, so at the moment r is graded its own exit does not exist yet. Only
once m_oosDecisionSeries is complete over the whole OOS window can the replay run.
In ensemble mode that series carries the member's adjusted decision and the live exit
reads the ensemble aggregate, which is the coupling the user identified: an LSTM entry
really can be closed by the ensemble turning against it.
3. THE STATISTIC HAS TO CHANGE WITH THE POLICY, AND THAT IS THE REAL FINDING.
A barrier exit pays a fixed R. A vote exit pays whatever the close happens to be. So
the moment vote exits are enabled the payoff is CONTINUOUS, and "win rate vs
break-even" stops being a meaningful test - there is no fixed break-even for a
variable payoff. SimulateTradeOutcome therefore returns R rather than a bool, and the
replay reports expectancy in R with its SE taken from the R distribution (overlap-
deflated on the same EffectiveSampleSize doctrine as every other SE here), not from a
binomial.
This is the same class of error as win-based scoring in 2026-08-09: measuring a
variable-payoff process with a fixed-payoff statistic. Naming it now, while vote exits
are still off, is much cheaper than discovering it after they go on.
4. WHY THIS IS SAFE TO SHIP TODAY.
Min_Vote_Close ships Disabled (101, unreachable on both scales it drives) and reaches
the AI signal through the new ExitPolicy() setter as 1.01, which the setter turns into
0.0 = no vote exit. Under that policy every replayed trade resolves at a barrier and
the simulation is arithmetically the same trade the deploy gate already certifies -
they cannot drift. The report says so explicitly, and prints ONCE per run in that
state; when vote exits are on it prints every era, because then the divergence is the
thing to watch. Nothing about today's numbers moves.
The gate switchover is wired but dormant by construction: it becomes exit-aware the
moment the input is enabled, which is exactly what "the certified number is the traded
number" has to mean.
KNOWN LIMIT, stated rather than hidden: only the AI early-exit route is replayed. The
rule-based averaged-vote close (m_threshold_close) depends on every other filter's
live vote, which pass 3 does not reproduce, so a position the classic filters would
have closed is held to its barrier here. The replay therefore holds LONGER than live
and overstates barrier-reached outcomes. Faithful only while the AI is the dominant
vote - which is the configuration this is being built for.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 16:25:57 -04:00
|
|
|
//--- Sized with the caches and zeroed, so a bar the scan has not reached yet reads as "no opinion"
|
|
|
|
|
//--- (0.0 = abstain) rather than as last era's decision - which would let a stale vote close a trade
|
|
|
|
|
//--- in the simulation that nothing would have closed live.
|
|
|
|
|
ArrayResize(m_oosDecisionSeries, bars);
|
|
|
|
|
ArrayInitialize(m_oosDecisionSeries, 0.0);
|
refactor: split CExpertSignalAIBase implementation by responsibility
ExpertSignalAIBase.mqh was 8216 lines: the class declaration followed by 87
method bodies covering training, labelling, feature extraction, persistence,
chart drawing, online learning, the GA auto-tuner and inference, all in one
file. Train() alone is 1492 lines; a change to arrow drawing meant scrolling
past the era loop.
Moved the bodies into Expert\AIBase\, included at the bottom of the original
after the class declaration:
Training.mqh 1607 era loop, plateau ladder, checkpoint select, deploy
Features.mqh 1093 indicator creation + per-bar input feature vector
ChartUI.mqh 634 arrows, arrow persistence, status panel, cleanup
Persistence.mqh 492 .stats/.cfg sidecars, CPU-inference validation, copy
OnlineLearning.mqh 461 live continual learning, EMA shadow, OOS simulator
Labels.mqh 309 ZigZag pivot labels, async label-cache prebuild
AutoTune.mqh 275 genetic tuner (population, crossover, halving)
Inference.mqh 235 softmax, prior calibration, class priors
ExpertSignalAIBase.mqh 8216 -> 3131 (declaration + topology build only)
This is a pure relocation - verified mechanically, not by eye: HEAD's file
reconstructed from the eight partials plus the surviving remainder is
byte-identical to HEAD, span for span (scratchpad verify_split.py). No
declaration moved, no signature changed, no code rewritten, so behaviour is
unchanged by construction.
Compiles 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 00:42:45 -04:00
|
|
|
ArrayResize(m_labelCacheHasValue, bars);
|
|
|
|
|
ArrayInitialize(m_labelCacheHasValue, false);
|
|
|
|
|
ArrayResize(m_featureCache, bars * m_neuronsCount);
|
|
|
|
|
ArrayResize(m_featureCacheHasValue, bars);
|
|
|
|
|
ArrayResize(m_featureCacheValid, bars);
|
|
|
|
|
ArrayInitialize(m_featureCacheHasValue, false);
|
|
|
|
|
m_labelCacheBars = bars;
|
|
|
|
|
m_labelCacheAnchorTime = m_Time.GetData(0);
|
|
|
|
|
return true;
|
|
|
|
|
}
|
|
|
|
|
//+------------------------------------------------------------------+
|
|
|
|
|
//| Lazy cache-miss fallback for a bar the eager prebuild pass (see |
|
feat(ai): triple-barrier labels replace exact-pivot ZigZag targets
The 31:1 class imbalance was self-inflicted by the TARGET, not a property
of the market. Labelling only the exact bar where a ZigZag pivot confirms
gave Buy 1164 / Sell 1164 / Neutral 35841, and every correction mechanism
this codebase accumulated sits downstream of that one choice: the
logit-adjusted loss and its range cap, the prior EMA, the +-3.0 output-bias
seed, balanced-accuracy-then-precision selection with its coverage floor,
the recall floor and its catch-22, the alternation gate, NMS, and the four
oversampling designs that collapsed before them.
The reference this engine is built on (references/neuronetworksbook.pdf
ch. 3.1/3.3) also uses ZigZag, but targets the DIRECTION TO THE NEXT
EXTREMUM on every bar - ~50/50 by construction, with no imbalance to
correct at all. It never had this problem because it never asked "is this
the pivot bar".
Labels are now the triple barrier (Lopez de Prado ch. 3), using the EA's
OWN SL_Mode/TP_Mode: does a trade opened at this bar's close reach its
target before its stop, within a horizon. Buy = long resolves, Sell =
short resolves, Neutral = neither. Consequences:
- dir-precision in the era line stops being a proxy and becomes the win
rate of the strategy under its own exit rules.
- Expected balance ~25/25/50 at the shipped 1:3 (gambler's ruin), i.e.
~2:1 instead of 31:1. Measured and logged at the end of the prebuild.
- Spread is charged on both legs, so it is a NET win rate.
- Intrabar ambiguity resolves to the STOP. OHLC cannot order two touches
inside one bar and the optimistic reading is how a backtested edge
becomes a live loss.
ZigZag stays as input features (EnableSwingContext) and now also supplies
the vertical barrier: the horizon is the median confirmed leg length,
snapped to a coarse ladder. Derived, not configured, and deliberately kept
out of the filename fingerprint - a filename keyed on a measured quantity
orphans a trained model the moment the measurement moves.
Removed, because the premise died with the old target:
- the alternation gate. Correct for pivot labels (a ZigZag cannot emit two
same-type pivots in a row, so a repeat was provably a false fire), and
wrong for barrier labels, which answer each bar independently. It also
took its worst consequence with it: a one-sided model previously got ONE
trade per backtest, a hard blocker on marketplace validation.
- SignalClusterWindow now defaults off - it de-duplicated repeats that are
now real trades. Kept as an opt-in display control.
- LABEL_WINDOW_BARS, the pivot-widening pass, ConfirmedZigZagLabel.
- the era-0 output-bias seed now needs a genuinely dominant class (0.70)
rather than 0.40; at ~50% Neutral a +-3.0 seed is a distortion, not a
correction.
Also fixed, both found while wiring the above:
1. RefreshConvergedSignal sized its buffers from a date delta
(Bars(sym, period, dtStudied, TimeCurrent())). dtStudied is a training
watermark; in the tester it is loaded from a live-chart save AHEAD of
the simulated date, so the interval inverted, Bars() returned ~0, and
the buffer came out at exactly m_historyBars - deep enough for the OHLC
window and far too shallow for the Donchian-50 / 20-bar-return / SMA
extension behind it. Inference silently computed DIFFERENT features
from the ones training learned on, live as well as in the tester. Now
sized from what the feature builder actually needs.
2. The barrier horizon is resolved on the deployed path too. A deployed
model never enters Train(), so it never reached the prebuild, and
OnlineLearnStep reads the horizon as its confirmation delay - left at
the fallback it would have backpropped bars whose barriers had not
resolved. Silent lookahead in the one place that writes to a live model.
SL_Mode/TP_Mode join the weights fingerprint: they define the labels now,
so a model trained at 1:3 must never be silently reused at 1:1. This
re-keys every pre-existing model by design - none were trained on this task.
Inference census extended with the vote gate. LongCondition/ShortCondition
open with a readiness check the refresh counters never see; in the tester it
reduces to "the seeded _optcache.nnw must have LOADED", and if it did not,
every vote is hard-zeroed while the model still answers Buy. The old three
counters would have read that as "the model says Neutral" - false, and a
completely different fix. This is the leading candidate for the
zero-direction backtest and the census can now name it in one run.
Both builds compile 0 errors / 0 warnings. Forces a full retrain.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 20:39:49 -04:00
|
|
|
//| AdvanceBarrierLabelState()) didn't cover - e.g. a new candle that |
|
|
|
|
|
//| closed after prebuild already completed. Such a bar sits inside |
|
|
|
|
|
//| the unresolved horizon: its triple-barrier outcome needs |
|
|
|
|
|
//| m_barrierHorizonBars more closes before it is knowable at all. |
|
|
|
|
|
//| Rather than guess, this always labels Neutral; the sequential |
|
|
|
|
|
//| prebuild scan is what assigns Buy/Sell once the forward window |
|
|
|
|
|
//| this bar's verdict depends on has actually closed. |
|
refactor: split CExpertSignalAIBase implementation by responsibility
ExpertSignalAIBase.mqh was 8216 lines: the class declaration followed by 87
method bodies covering training, labelling, feature extraction, persistence,
chart drawing, online learning, the GA auto-tuner and inference, all in one
file. Train() alone is 1492 lines; a change to arrow drawing meant scrolling
past the era loop.
Moved the bodies into Expert\AIBase\, included at the bottom of the original
after the class declaration:
Training.mqh 1607 era loop, plateau ladder, checkpoint select, deploy
Features.mqh 1093 indicator creation + per-bar input feature vector
ChartUI.mqh 634 arrows, arrow persistence, status panel, cleanup
Persistence.mqh 492 .stats/.cfg sidecars, CPU-inference validation, copy
OnlineLearning.mqh 461 live continual learning, EMA shadow, OOS simulator
Labels.mqh 309 ZigZag pivot labels, async label-cache prebuild
AutoTune.mqh 275 genetic tuner (population, crossover, halving)
Inference.mqh 235 softmax, prior calibration, class priors
ExpertSignalAIBase.mqh 8216 -> 3131 (declaration + topology build only)
This is a pure relocation - verified mechanically, not by eye: HEAD's file
reconstructed from the eight partials plus the surviving remainder is
byte-identical to HEAD, span for span (scratchpad verify_split.py). No
declaration moved, no signature changed, no code rewritten, so behaviour is
unchanged by construction.
Compiles 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 00:42:45 -04:00
|
|
|
//+------------------------------------------------------------------+
|
|
|
|
|
void CExpertSignalAIBase::ComputeLabelForBar(int i, int bars, bool &buy, bool &sell)
|
|
|
|
|
{
|
|
|
|
|
buy = false;
|
|
|
|
|
sell = false;
|
|
|
|
|
}
|
|
|
|
|
//+------------------------------------------------------------------+
|
feat(ai): triple-barrier labels replace exact-pivot ZigZag targets
The 31:1 class imbalance was self-inflicted by the TARGET, not a property
of the market. Labelling only the exact bar where a ZigZag pivot confirms
gave Buy 1164 / Sell 1164 / Neutral 35841, and every correction mechanism
this codebase accumulated sits downstream of that one choice: the
logit-adjusted loss and its range cap, the prior EMA, the +-3.0 output-bias
seed, balanced-accuracy-then-precision selection with its coverage floor,
the recall floor and its catch-22, the alternation gate, NMS, and the four
oversampling designs that collapsed before them.
The reference this engine is built on (references/neuronetworksbook.pdf
ch. 3.1/3.3) also uses ZigZag, but targets the DIRECTION TO THE NEXT
EXTREMUM on every bar - ~50/50 by construction, with no imbalance to
correct at all. It never had this problem because it never asked "is this
the pivot bar".
Labels are now the triple barrier (Lopez de Prado ch. 3), using the EA's
OWN SL_Mode/TP_Mode: does a trade opened at this bar's close reach its
target before its stop, within a horizon. Buy = long resolves, Sell =
short resolves, Neutral = neither. Consequences:
- dir-precision in the era line stops being a proxy and becomes the win
rate of the strategy under its own exit rules.
- Expected balance ~25/25/50 at the shipped 1:3 (gambler's ruin), i.e.
~2:1 instead of 31:1. Measured and logged at the end of the prebuild.
- Spread is charged on both legs, so it is a NET win rate.
- Intrabar ambiguity resolves to the STOP. OHLC cannot order two touches
inside one bar and the optimistic reading is how a backtested edge
becomes a live loss.
ZigZag stays as input features (EnableSwingContext) and now also supplies
the vertical barrier: the horizon is the median confirmed leg length,
snapped to a coarse ladder. Derived, not configured, and deliberately kept
out of the filename fingerprint - a filename keyed on a measured quantity
orphans a trained model the moment the measurement moves.
Removed, because the premise died with the old target:
- the alternation gate. Correct for pivot labels (a ZigZag cannot emit two
same-type pivots in a row, so a repeat was provably a false fire), and
wrong for barrier labels, which answer each bar independently. It also
took its worst consequence with it: a one-sided model previously got ONE
trade per backtest, a hard blocker on marketplace validation.
- SignalClusterWindow now defaults off - it de-duplicated repeats that are
now real trades. Kept as an opt-in display control.
- LABEL_WINDOW_BARS, the pivot-widening pass, ConfirmedZigZagLabel.
- the era-0 output-bias seed now needs a genuinely dominant class (0.70)
rather than 0.40; at ~50% Neutral a +-3.0 seed is a distortion, not a
correction.
Also fixed, both found while wiring the above:
1. RefreshConvergedSignal sized its buffers from a date delta
(Bars(sym, period, dtStudied, TimeCurrent())). dtStudied is a training
watermark; in the tester it is loaded from a live-chart save AHEAD of
the simulated date, so the interval inverted, Bars() returned ~0, and
the buffer came out at exactly m_historyBars - deep enough for the OHLC
window and far too shallow for the Donchian-50 / 20-bar-return / SMA
extension behind it. Inference silently computed DIFFERENT features
from the ones training learned on, live as well as in the tester. Now
sized from what the feature builder actually needs.
2. The barrier horizon is resolved on the deployed path too. A deployed
model never enters Train(), so it never reached the prebuild, and
OnlineLearnStep reads the horizon as its confirmation delay - left at
the fallback it would have backpropped bars whose barriers had not
resolved. Silent lookahead in the one place that writes to a live model.
SL_Mode/TP_Mode join the weights fingerprint: they define the labels now,
so a model trained at 1:3 must never be silently reused at 1:1. This
re-keys every pre-existing model by design - none were trained on this task.
Inference census extended with the vote gate. LongCondition/ShortCondition
open with a readiness check the refresh counters never see; in the tester it
reduces to "the seeded _optcache.nnw must have LOADED", and if it did not,
every vote is hard-zeroed while the model still answers Buy. The old three
counters would have read that as "the model says Neutral" - false, and a
completely different fix. This is the leading candidate for the
zero-direction backtest and the census can now name it in one run.
Both builds compile 0 errors / 0 warnings. Forces a full retrain.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 20:39:49 -04:00
|
|
|
//| SL/TP ATR multiples for the triple-barrier label, taken from the |
|
|
|
|
|
//| EA's own SL_Mode/TP_Mode (m_sl_mode/m_tp_mode, protected members |
|
|
|
|
|
//| of CExpertSignalCustom, set in Warrior_EA.mq5's per-topology |
|
|
|
|
|
//| setup block). Using the traded values is the entire point: it is |
|
|
|
|
|
//| what makes the era line's dir-precision a real win rate instead |
|
|
|
|
|
//| of a proxy for one. |
|
|
|
|
|
//| |
|
|
|
|
|
//| The INTELLIGENT modes scale with AI confidence, which does not |
|
|
|
|
|
//| exist when a label is computed - and must not, or the target |
|
|
|
|
|
//| would depend on the model's own output and the whole thing would |
|
|
|
|
|
//| be circular. Both therefore fall back to their ZERO-CONFIDENCE |
|
|
|
|
|
//| base (the trade the EA would place knowing nothing), which is |
|
|
|
|
|
//| also the widest stop and tightest target either mode can pick, so |
|
|
|
|
|
//| the label is the conservative member of the family it stands for. |
|
|
|
|
|
//| TP_INTELLIGENT is risk-relative by design, so its multiple is |
|
|
|
|
|
//| expressed against the resolved stop rather than against ATR. |
|
refactor: split CExpertSignalAIBase implementation by responsibility
ExpertSignalAIBase.mqh was 8216 lines: the class declaration followed by 87
method bodies covering training, labelling, feature extraction, persistence,
chart drawing, online learning, the GA auto-tuner and inference, all in one
file. Train() alone is 1492 lines; a change to arrow drawing meant scrolling
past the era loop.
Moved the bodies into Expert\AIBase\, included at the bottom of the original
after the class declaration:
Training.mqh 1607 era loop, plateau ladder, checkpoint select, deploy
Features.mqh 1093 indicator creation + per-bar input feature vector
ChartUI.mqh 634 arrows, arrow persistence, status panel, cleanup
Persistence.mqh 492 .stats/.cfg sidecars, CPU-inference validation, copy
OnlineLearning.mqh 461 live continual learning, EMA shadow, OOS simulator
Labels.mqh 309 ZigZag pivot labels, async label-cache prebuild
AutoTune.mqh 275 genetic tuner (population, crossover, halving)
Inference.mqh 235 softmax, prior calibration, class priors
ExpertSignalAIBase.mqh 8216 -> 3131 (declaration + topology build only)
This is a pure relocation - verified mechanically, not by eye: HEAD's file
reconstructed from the eight partials plus the surviving remainder is
byte-identical to HEAD, span for span (scratchpad verify_split.py). No
declaration moved, no signature changed, no code rewritten, so behaviour is
unchanged by construction.
Compiles 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 00:42:45 -04:00
|
|
|
//+------------------------------------------------------------------+
|
feat: derive the ATR multiples from measured excursions - no hardcoded geometry
The barrier was still two constants. SL_Mode/TP_Mode left the Inputs tab in
3482b6c, but the fallback was a hardcoded 2:6 and the geometry scan only ever
chose from a hardcoded grid {2,3} x {2,3,4,6,8,10}. Picking the least-bad of
eleven guesses is not deriving anything.
WHY THE SCAN WAS THE WRONG INSTRUMENT, now measurable rather than argued. It
ranks pairings by how predictable their OUTCOME is - a question about direction.
The excursion test (2c78f3b) ran on SP500 H1 and direction is the one thing
absent: ASYMMETRY p=0.0846, against RANGE/UP/DOWN all at p=0.0050, with RANGE
scoring 0.01345 vs a 0.00343 null - 4x, where the barrier label sits at 1.01x.
Hence the scan failing its own gate on every run, and its "winner" wandering
2:8 -> 3:8 -> 2:8 -> 2:4 across four runs of the same data. Excursion SIZE is
strongly measurable, so derive the geometry from that instead.
stop = q25 of measured ADVERSE travel (ordinary noise does not reach it)
target = q50 of measured FAVOURABLE travel (reached ~half the time, by
construction, inside the horizon)
Continuous, in ATR units, superseding the enum multiples. Reachability ("target
on X% of bars, stop on Y%") and the implied break-even are printed so the choice
is auditable rather than trusted.
FIXED-POINT ITERATION, not one-shot. ComputeBarrierHorizonBars scales the
horizon with the target (first-passage time grows with the band) and the
excursions are measured OVER the horizon, so target -> horizon -> excursions ->
target is a real loop - deriving once sizes the target from travel measured
under the PREVIOUS horizon. Re-measures until the multiples move <5%, capped at
3 passes, and says so if it does not settle.
Does NOT create expectancy, and the log says as much: chance precision equals
break-even at every geometry (m/(m+k) on both sides). It buys a target the
market reaches and a stop that survives noise. Where Min_Risk_Reward_Ratio
forces a target the market rarely reaches, it WARNS rather than overriding -
the ratio is the user's risk policy, so the honest move is to state its cost.
That is the collision that once rejected 100% of setups.
Pinned in the .cfg as doubles appended AFTER this morning's two ints, so .cfg
files written earlier today still load (their length guard finds no doubles) and
a model that carries them was trained on them and never re-derives.
Also fixes a message from e5ceed6 that claimed "this model resumed from disk"
unconditionally - it printed above a "seeding era 0" line on a brand-new model,
because the branch fires whenever the cache is not built, which is equally true
before a fresh model's first prebuild. A diagnostic that misreports its own
trigger is worse than one that says nothing: it gets quoted back as evidence.
FORCES A FULL RETRAIN (labels change).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-07 12:06:25 -04:00
|
|
|
//+------------------------------------------------------------------+
|
|
|
|
|
//| DERIVE THE BARRIER FROM WHAT PRICE ACTUALLY DOES, not from an |
|
|
|
|
|
//| enum. Reads the measured MFE/MAE distribution collected by the |
|
|
|
|
|
//| label prebuild and sets the ATR multiples from its quantiles. |
|
|
|
|
|
//| |
|
|
|
|
|
//| WHY THIS AND NOT THE GEOMETRY SCAN. The scan ranks candidate SL:TP |
|
|
|
|
|
//| pairings by how predictable their OUTCOME is, which is a question |
|
|
|
|
|
//| about direction - and direction is the one thing measured absent |
|
|
|
|
|
//| here (ASYMMETRY p=0.0846 on SP500 H1, against RANGE/UP/DOWN all at |
|
|
|
|
|
//| p=0.0050). That is why its winner fails its own gate on every run |
|
|
|
|
|
//| and why its "best" wanders 2:8 -> 3:8 -> 2:8 -> 2:4. Excursion |
|
|
|
|
|
//| SIZE, by contrast, clears at 4x its null. So derive the geometry |
|
|
|
|
|
//| from the quantity that is actually measurable. |
|
|
|
|
|
//| |
|
|
|
|
|
//| WHAT THIS DOES NOT DO: create expectancy. Under a driftless walk |
|
|
|
|
|
//| the probability of touching +k*ATR before -m*ATR is m/(m+k), which |
|
|
|
|
|
//| is ALSO the break-even win rate for that payoff - so no choice of |
|
|
|
|
|
//| geometry has an edge, and this one does not either. What it buys |
|
|
|
|
|
//| is a target that is actually reachable inside the horizon and a |
|
|
|
|
|
//| stop wide enough to survive ordinary noise, both read off the |
|
|
|
|
|
//| data instead of guessed. The reachability figures are printed so |
|
|
|
|
|
//| the choice can be audited rather than trusted. |
|
|
|
|
|
//+------------------------------------------------------------------+
|
feat: first-passage ladder + expectancy scan - price every geometry, not just the chosen one
Corrects the premise of the previous plan. Break-even is NOT a ceiling.
If the model shifts the win probability on the bars it selects from
p0 = m/(m+k) to p0 + d, then
EV = (p0+d)*k - (1-p0-d)*m = d*(k+m)
because p0*k - (1-p0)*m is zero by construction. The stop:target RATIO
is expectancy-neutral - a punishing break-even is exactly repaid by the
payoff - and only the real edge d and the TOTAL WIDTH (k+m) move EV.
Width matters because the spread is charged once per trade however wide
the barriers are, so a narrow barrier spends much of its own range on
costs. DeriveBarrierGeometry's own comment already said the ratio buys
nothing; the objective just never followed from it.
Blocker this had to solve first: m_excUpCache/m_excDownCache hold only
MAXIMUM travel each way, and a maximum cannot say which side was
reached FIRST - so any geometry other than the walked one was
undecidable on precisely the bars where both barriers were touched,
~28% of the sample.
- BARRIER_LADDER: per bar, the first-touch AGE for 8 travel distances
in each direction, filled during the walk the labels already run.
Cursors keep it O(1) amortised per walked bar rather than 16
comparisons. Levels are travel FROM ENTRY, not barrier prices, so one
ladder serves both directions and the spread is applied analytically
when a level converts back to an SL/TP multiple - storing prices
would need four ladders and bake today's spread into the cache.
Sized, invalidated and validity-gated with the label caches.
- ReportGeometryExpectancyScan: every ladder pair priced exactly off
that cache - width in ATR and in SPREADS (cost efficiency, knowable
without knowing d), break-even, both base rates, the share of bars
resolved inside the horizon, and EV per unit of edge. Compares the
widest resolvable pair against the quantile rule's pick.
MEASUREMENT ONLY - the quantile rule still chooses. Nothing here can
measure d, and width buys nothing if the wider target is less
predictable. Base rates are printed beside each break-even because a
persistent gap is DRIFT and must not be credited to the model.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-10 12:59:18 -04:00
|
|
|
//+------------------------------------------------------------------+
|
|
|
|
|
//| Expectancy of every ladder pair, measured exactly off the |
|
|
|
|
|
//| first-passage cache (see BARRIER_LADDER). |
|
|
|
|
|
//| |
|
|
|
|
|
//| WHY THIS OBJECTIVE. Let the model shift the win probability on the |
|
|
|
|
|
//| bars it selects from the base rate p0 = m/(m+k) to p0 + d. Then |
|
|
|
|
|
//| EV = (p0+d)*k - (1-p0-d)*m = d*(k+m), |
|
|
|
|
|
//| because p0*k - (1-p0)*m is zero by construction. So the stop:target|
|
|
|
|
|
//| RATIO is expectancy-neutral - a punishing break-even is exactly |
|
|
|
|
|
//| repaid by the payoff - and only two things move EV: the real edge |
|
|
|
|
|
//| d, and the TOTAL BARRIER WIDTH (k+m). Width matters because the |
|
|
|
|
|
//| spread is charged once per trade however wide the barriers are, so |
|
|
|
|
|
//| a narrow barrier spends a large share of its own range on costs. |
|
|
|
|
|
//| That is why every row below reports width in SPREADS as well as in |
|
|
|
|
|
//| ATR: it is the cost efficiency of the geometry, and it is knowable |
|
|
|
|
|
//| without knowing d. |
|
|
|
|
|
//| |
|
|
|
|
|
//| WHAT IT DOES NOT DO: measure d. Nothing here can - d is a property |
|
|
|
|
|
//| of the model and the features, not of the barrier - so this cannot |
|
|
|
|
|
//| and must not be read as "this geometry is profitable". It answers |
|
|
|
|
|
//| the narrower question the previous rule never asked: GIVEN an edge,|
|
|
|
|
|
//| which geometry converts the most of it into money, and what does |
|
|
|
|
|
//| each pair cost in spread and in trade frequency. |
|
|
|
|
|
//| |
|
|
|
|
|
//| The base rates are printed beside each break-even deliberately. On |
|
|
|
|
|
//| a driftless walk they coincide; a persistent gap is DRIFT (being |
|
|
|
|
|
//| long pays on an index) and must never be credited to the model - |
|
|
|
|
|
//| see chancePrecPct, which is measured for exactly that reason. |
|
|
|
|
|
//+------------------------------------------------------------------+
|
|
|
|
|
void CExpertSignalAIBase::ReportGeometryExpectancyScan(void)
|
|
|
|
|
{
|
|
|
|
|
int bars = m_labelCacheBars;
|
|
|
|
|
if(bars <= 0 || ArraySize(m_ladderUpAt) < bars * BARRIER_LADDER_COUNT)
|
|
|
|
|
return;
|
|
|
|
|
//--- IS region only, matching DeriveBarrierGeometry and BuildMiSample: a geometry chosen with the
|
|
|
|
|
//--- holdout in view has used the holdout for selection, and it stops being a holdout.
|
|
|
|
|
int oosCutoff = (int)(MathMax(0, MathMin(100, m_oosSplitPct)) / 100.0
|
|
|
|
|
* MathMax(bars - MathMax(m_historyBars, 0), 0));
|
|
|
|
|
int from = MathMax(oosCutoff, 0);
|
|
|
|
|
double spread = (double)m_symbol.Spread() * m_symbol.Point();
|
|
|
|
|
if(!MathIsValidNumber(spread) || spread < 0.0)
|
|
|
|
|
spread = 0.0;
|
|
|
|
|
//--- Spread expressed in ATR, averaged over the same bars the ladder covers - the ladder is in ATR
|
|
|
|
|
//--- units, so the cost has to be converted into the same units before it can be netted off a leg.
|
|
|
|
|
double spreadAtrSum = 0.0;
|
|
|
|
|
int atrN = 0;
|
|
|
|
|
for(int i = from; i < bars; i++)
|
|
|
|
|
{
|
|
|
|
|
if(i >= ArraySize(m_labelCacheHasValue) || !m_labelCacheHasValue[i])
|
|
|
|
|
continue;
|
|
|
|
|
double a = m_ATR.Main(i);
|
|
|
|
|
if(!MathIsValidNumber(a) || a <= 0.0)
|
|
|
|
|
continue;
|
|
|
|
|
spreadAtrSum += spread / a;
|
|
|
|
|
atrN++;
|
|
|
|
|
}
|
|
|
|
|
if(atrN < BARRIER_DERIVE_MIN_SAMPLES)
|
|
|
|
|
return;
|
|
|
|
|
double spreadAtr = spreadAtrSum / atrN;
|
fix(geometry): the target was small BY CONSTRUCTION - ratio is now policy, scale is measured, ladder ceiling removed
The derivation read the stop from q75 of ADVERSE travel and the target from q50
of FAVOURABLE travel. Over one horizon those distributions are broadly the same
shape, so q75 > q50 MECHANICALLY - the target came out smaller than the stop no
matter what the market did. SP500 H4 shipped stop 3.07 / target 1.70: a 0.55:1
payoff needing 64.3%. That was never a measurement, it was two mismatched
constants.
The reachability line printed beside it - "target on 50.0% of bars, stop on
25.0%" - is exactly 1-q50 and 1-q75. Tautological. It cannot disconfirm
anything, and it read as validation.
WIDTH AND RATIO ARE INDEPENDENT AND ONLY ONE PAYS. EV = edge x width;
ratio is EV-neutral (a driftless walk reaches +m before -k with probability
k/(k+m), which IS break-even). Width is what buys cost efficiency: the spread
is a fixed 0.047*ATR here, so the shipped 4.77*ATR width paid it 21 times per
unit of travel. So:
RATIO = policy. BARRIER_TARGET_RR = 2.0 (user's 1:2). Break-even 33.3%.
SCALE = measured. The stop quantile is chosen from a ladder, WIDEST FIRST,
taking the first rung whose implied 2x target is still reached often
enough to be a trainable class.
That last clause is the difference from the min-reward:risk raise removed in
2026-08-09, which forced target = 2 x stop with NO reachability test, landed on
6.66*ATR reachable on 3.3% of bars, and trained the model to predict something
that essentially never happened. Same ratio; the scale now retreats until the
data says the target is attainable. Every rung is logged.
LADDER CEILING REMOVED. BARRIER_LADDER stopped at 5.00 and the expectancy scan's
"best resolvable pair on width alone" came back as stop 5.05 / target 4.95 - it
pinned to the top rung. A recommendation landing exactly on the edge of its own
search space is a boundary, not a finding: it cannot tell "5 ATR is optimal"
from "5 ATR is all we allowed". Extended to 20*ATR (8 -> 14 rungs). Nothing else
needs editing - every consumer is parameterised by BARRIER_LADDER_COUNT - and
the horizon constraints (decided >= 60%, reachability floor) now bind instead of
a constant.
THE SCAN COULD NOT SEE THE SHIPPED GEOMETRY. ReportBarrierGeometryScan looked
the configured pair up in its integer grid, and DeriveBarrierGeometry produces
CONTINUOUS multiples (3.07/1.70) that can never equal a grid point - so
cfgExcess stayed at its -1.0 sentinel and the report printed "configured 3:2
scores -1.00000", which reads as a catastrophic score and actually means "never
evaluated". Worse, the grid skipped target<stop entirely because it "inverts the
trade's whole premise" - while the derivation was shipping exactly that. The
incumbent is now always scored as a peer (never crowned; it is already in force
and is not an enum pairing the scan could adopt).
BREAK-EVEN NOW INCLUDES THE SPREAD. Every report quoted the frictionless
SL/(SL+TP). On SP500 H4 that read 64.3% while the MEASURED zero-skill rate was
62.1% - a 2.2pp gap that IS the cost, and that made every model look 2.2pp
better than it was. CostAdjustedBreakEvenPct() prices a win at (TP - spread) and
a loss at (SL + spread), matching the expectancy scan's convention exactly so
the two reports cannot disagree.
It also feeds FitDirConfThreshold, which is the correctness half: the operating
point subtracts break-even from precision, so the frictionless figure made every
candidate threshold look better by the width of the spread - 2.2pp against a
measured edge of 2.3pp, i.e. very nearly all of it.
Era line now carries both: "break-even 64.3% frictionless, 66.6% AFTER SPREAD".
Forces a full relabel and retrain. Requested.
NOT COMPILED - user compiles.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 00:34:32 -04:00
|
|
|
//--- Published so CostAdjustedBreakEvenPct() can price the cost into every break-even the run quotes.
|
|
|
|
|
m_spreadAtr = spreadAtr;
|
feat: first-passage ladder + expectancy scan - price every geometry, not just the chosen one
Corrects the premise of the previous plan. Break-even is NOT a ceiling.
If the model shifts the win probability on the bars it selects from
p0 = m/(m+k) to p0 + d, then
EV = (p0+d)*k - (1-p0-d)*m = d*(k+m)
because p0*k - (1-p0)*m is zero by construction. The stop:target RATIO
is expectancy-neutral - a punishing break-even is exactly repaid by the
payoff - and only the real edge d and the TOTAL WIDTH (k+m) move EV.
Width matters because the spread is charged once per trade however wide
the barriers are, so a narrow barrier spends much of its own range on
costs. DeriveBarrierGeometry's own comment already said the ratio buys
nothing; the objective just never followed from it.
Blocker this had to solve first: m_excUpCache/m_excDownCache hold only
MAXIMUM travel each way, and a maximum cannot say which side was
reached FIRST - so any geometry other than the walked one was
undecidable on precisely the bars where both barriers were touched,
~28% of the sample.
- BARRIER_LADDER: per bar, the first-touch AGE for 8 travel distances
in each direction, filled during the walk the labels already run.
Cursors keep it O(1) amortised per walked bar rather than 16
comparisons. Levels are travel FROM ENTRY, not barrier prices, so one
ladder serves both directions and the spread is applied analytically
when a level converts back to an SL/TP multiple - storing prices
would need four ladders and bake today's spread into the cache.
Sized, invalidated and validity-gated with the label caches.
- ReportGeometryExpectancyScan: every ladder pair priced exactly off
that cache - width in ATR and in SPREADS (cost efficiency, knowable
without knowing d), break-even, both base rates, the share of bars
resolved inside the horizon, and EV per unit of edge. Compares the
widest resolvable pair against the quantile rule's pick.
MEASUREMENT ONLY - the quantile rule still chooses. Nothing here can
measure d, and width buys nothing if the wider target is less
predictable. Base rates are printed beside each break-even because a
persistent gap is DRIFT and must not be credited to the model.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-10 12:59:18 -04:00
|
|
|
Print(ID + StringFormat(": barrier expectancy scan - spread averages %.3f*ATR over %d bars. EV per "
|
|
|
|
|
"trade = edge x width, so the ratio is EV-neutral and WIDTH is what pays; "
|
|
|
|
|
"'spreads' is width/spread (cost efficiency), 'decided' is the share of bars "
|
|
|
|
|
"the long side resolved inside the %d-bar horizon. No row here demonstrates "
|
|
|
|
|
"an edge - it prices one.", spreadAtr, atrN, m_barrierHorizonBars));
|
|
|
|
|
double bestWidth = -1.0;
|
|
|
|
|
int bestT = -1, bestS = -1;
|
|
|
|
|
for(int tL = 0; tL < BARRIER_LADDER_COUNT; tL++)
|
fix(shutdown): make ExitPolicy public, and stop every long loop the moment MT5 asks
Two things, one of which was a compile error.
1. ExitPolicy() was declared in the protected block but is pushed in from
Warrior_EA.mq5:770. Moved to public beside the other EA-facing setters.
2. Chart objects surviving OnDeinit. The 4,500 ms teardown budget is measured
from the STOP REQUEST, not from OnDeinit's first line, and OnDeinit cannot
begin until whatever is in flight returns - so a scan still running after
_StopFlag is raised does not delay the cleanup, it SPENDS it, and the purge
never gets its turn.
New CExpertSignalAIBase::ShutdownRequested() = IsStopped() || m_shutdownInProgress.
Deliberately NOT m_trainingStopRequested: that latches, and a latched flag
would permanently disable scans that must run again on the next Start.
Guarded, longest first:
- TuneIndicatorsByFilter - per candidate, restoring the OPERATOR's settings
on the way out (best[] is mutated in place; the tuner otherwise keeps the
last trial's parameters, which nothing chose).
- ReportBarrierGeometryScan - per pairing, breaking to ONE restore point so
m_barrierScanLiveLabels can never be left true (that makes ComputeLabelForBar
read the last candidate's multiples as the configured geometry).
- ReportFeatureLabelInformation / ReportExcursionInformation / lag profile -
nulls ABANDON rather than truncate: fewer draws is not a smaller null, it
is a wrong one, and p shifts toward significance. m_dirEvidence staying
false is the safe direction.
- SimulateExitPolicyOutcomes - zeroes its accumulators so the divergence line
is dropped instead of latching a partial expectancy as the run's only report.
- ReportGeometryExpectancyScan - per ladder rung.
- HttpGet - one choke point for up to a dozen blocking WebRequests per
first-pass Update(). An in-flight request cannot be cancelled; refusing to
start another is the whole remedy.
- PollTraining, OnChartEventHandler's study event, TuneIndicatorsAndTrain -
entry points, so a queued event cannot open an era during teardown.
TuneIndicatorsAndTrain's guard is the first statement, ahead of the
m_tuneFilterDone / g_ensembleChartTuneDone latches.
- OnTick / OnTimer / OnChartEvent.
Training's own bar loops already honoured this (pass 1 per bar, passes 2/2.5/3
yield on a 120 ms budget); the warm-up scans did not, and they are the longest
uninterruptible stretches the EA has.
StopTraining() is unchanged: the operator's Stop still finalises synchronously.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 17:03:11 -04:00
|
|
|
{
|
|
|
|
|
//--- Every ladder pairing walks the whole labelled window again. Purely a pricing report - it
|
|
|
|
|
//--- adopts nothing - so leaving with the rows printed so far costs only information.
|
|
|
|
|
if(ShutdownRequested())
|
|
|
|
|
return;
|
feat: first-passage ladder + expectancy scan - price every geometry, not just the chosen one
Corrects the premise of the previous plan. Break-even is NOT a ceiling.
If the model shifts the win probability on the bars it selects from
p0 = m/(m+k) to p0 + d, then
EV = (p0+d)*k - (1-p0-d)*m = d*(k+m)
because p0*k - (1-p0)*m is zero by construction. The stop:target RATIO
is expectancy-neutral - a punishing break-even is exactly repaid by the
payoff - and only the real edge d and the TOTAL WIDTH (k+m) move EV.
Width matters because the spread is charged once per trade however wide
the barriers are, so a narrow barrier spends much of its own range on
costs. DeriveBarrierGeometry's own comment already said the ratio buys
nothing; the objective just never followed from it.
Blocker this had to solve first: m_excUpCache/m_excDownCache hold only
MAXIMUM travel each way, and a maximum cannot say which side was
reached FIRST - so any geometry other than the walked one was
undecidable on precisely the bars where both barriers were touched,
~28% of the sample.
- BARRIER_LADDER: per bar, the first-touch AGE for 8 travel distances
in each direction, filled during the walk the labels already run.
Cursors keep it O(1) amortised per walked bar rather than 16
comparisons. Levels are travel FROM ENTRY, not barrier prices, so one
ladder serves both directions and the spread is applied analytically
when a level converts back to an SL/TP multiple - storing prices
would need four ladders and bake today's spread into the cache.
Sized, invalidated and validity-gated with the label caches.
- ReportGeometryExpectancyScan: every ladder pair priced exactly off
that cache - width in ATR and in SPREADS (cost efficiency, knowable
without knowing d), break-even, both base rates, the share of bars
resolved inside the horizon, and EV per unit of edge. Compares the
widest resolvable pair against the quantile rule's pick.
MEASUREMENT ONLY - the quantile rule still chooses. Nothing here can
measure d, and width buys nothing if the wider target is less
predictable. Base rates are printed beside each break-even because a
persistent gap is DRIFT and must not be credited to the model.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-10 12:59:18 -04:00
|
|
|
for(int sL = 0; sL < BARRIER_LADDER_COUNT; sL++)
|
|
|
|
|
{
|
|
|
|
|
//--- Ladder levels are TRAVEL from the entry close; converting back to the SL/TP multiples that
|
|
|
|
|
//--- would actually be pinned puts the spread where the fill puts it - see BARRIER_LADDER.
|
|
|
|
|
double reward = BARRIER_LADDER[tL] - spreadAtr;
|
|
|
|
|
double risk = BARRIER_LADDER[sL] + spreadAtr;
|
|
|
|
|
if(reward <= 0.0 || risk <= 0.0)
|
|
|
|
|
continue; // target inside the spread - not a tradeable geometry at any hit rate
|
|
|
|
|
long nLong = 0, nShort = 0, nDecided = 0, nSeen = 0;
|
|
|
|
|
for(int i = from; i < bars; i++)
|
|
|
|
|
{
|
|
|
|
|
if(i >= ArraySize(m_labelCacheHasValue) || !m_labelCacheHasValue[i])
|
|
|
|
|
continue;
|
|
|
|
|
int b = i * BARRIER_LADDER_COUNT;
|
|
|
|
|
int tUpT = m_ladderUpAt[b + tL]; // long target / short stop reference
|
|
|
|
|
int tDnS = m_ladderDownAt[b + sL]; // long stop
|
|
|
|
|
int tDnT = m_ladderDownAt[b + tL]; // short target
|
|
|
|
|
int tUpS = m_ladderUpAt[b + sL]; // short stop
|
|
|
|
|
//--- 0 means "never touched inside the horizon". A smaller age is the EARLIER touch, and a
|
|
|
|
|
//--- tie goes to the stop - the same pessimistic convention the label walk uses, so these
|
|
|
|
|
//--- numbers describe the same game the training target does.
|
|
|
|
|
if(tUpT > 0 && (tDnS == 0 || tUpT < tDnS))
|
|
|
|
|
nLong++;
|
|
|
|
|
if(tDnT > 0 && (tUpS == 0 || tDnT < tUpS))
|
|
|
|
|
nShort++;
|
|
|
|
|
if(tUpT > 0 || tDnS > 0)
|
|
|
|
|
nDecided++;
|
|
|
|
|
nSeen++;
|
|
|
|
|
}
|
|
|
|
|
if(nSeen < BARRIER_DERIVE_MIN_SAMPLES)
|
|
|
|
|
continue;
|
|
|
|
|
double pL = 100.0 * nLong / nSeen;
|
|
|
|
|
double pS = 100.0 * nShort / nSeen;
|
|
|
|
|
double be = 100.0 * risk / (risk + reward);
|
|
|
|
|
double width = risk + reward;
|
|
|
|
|
double decided = 100.0 * nDecided / nSeen;
|
|
|
|
|
PrintFormat("%s: stop %.2f target %.2f | width %.2f*ATR = %.1f spreads | break-even %.1f%% |"
|
|
|
|
|
" base long %.1f%% short %.1f%% | decided %.1f%% | EV at a 1pp edge %.4f*ATR",
|
|
|
|
|
ID, risk, reward, width, (spreadAtr > 0.0 ? width / spreadAtr : 0.0), be,
|
|
|
|
|
pL, pS, decided, 0.01 * width);
|
|
|
|
|
//--- The recommendation is the WIDEST pair that still resolves most of its bars inside the
|
|
|
|
|
//--- horizon. Width is the whole of the EV multiplier; the decided-rate floor is what stops it
|
|
|
|
|
//--- running away to a barrier the horizon can never deliver, which is the failure the shipped
|
|
|
|
|
//--- 128-bar clamp already caused once.
|
|
|
|
|
if(decided >= 60.0 && width > bestWidth)
|
|
|
|
|
{
|
|
|
|
|
bestWidth = width;
|
|
|
|
|
bestT = tL;
|
|
|
|
|
bestS = sL;
|
|
|
|
|
}
|
|
|
|
|
}
|
fix(shutdown): make ExitPolicy public, and stop every long loop the moment MT5 asks
Two things, one of which was a compile error.
1. ExitPolicy() was declared in the protected block but is pushed in from
Warrior_EA.mq5:770. Moved to public beside the other EA-facing setters.
2. Chart objects surviving OnDeinit. The 4,500 ms teardown budget is measured
from the STOP REQUEST, not from OnDeinit's first line, and OnDeinit cannot
begin until whatever is in flight returns - so a scan still running after
_StopFlag is raised does not delay the cleanup, it SPENDS it, and the purge
never gets its turn.
New CExpertSignalAIBase::ShutdownRequested() = IsStopped() || m_shutdownInProgress.
Deliberately NOT m_trainingStopRequested: that latches, and a latched flag
would permanently disable scans that must run again on the next Start.
Guarded, longest first:
- TuneIndicatorsByFilter - per candidate, restoring the OPERATOR's settings
on the way out (best[] is mutated in place; the tuner otherwise keeps the
last trial's parameters, which nothing chose).
- ReportBarrierGeometryScan - per pairing, breaking to ONE restore point so
m_barrierScanLiveLabels can never be left true (that makes ComputeLabelForBar
read the last candidate's multiples as the configured geometry).
- ReportFeatureLabelInformation / ReportExcursionInformation / lag profile -
nulls ABANDON rather than truncate: fewer draws is not a smaller null, it
is a wrong one, and p shifts toward significance. m_dirEvidence staying
false is the safe direction.
- SimulateExitPolicyOutcomes - zeroes its accumulators so the divergence line
is dropped instead of latching a partial expectancy as the run's only report.
- ReportGeometryExpectancyScan - per ladder rung.
- HttpGet - one choke point for up to a dozen blocking WebRequests per
first-pass Update(). An in-flight request cannot be cancelled; refusing to
start another is the whole remedy.
- PollTraining, OnChartEventHandler's study event, TuneIndicatorsAndTrain -
entry points, so a queued event cannot open an era during teardown.
TuneIndicatorsAndTrain's guard is the first statement, ahead of the
m_tuneFilterDone / g_ensembleChartTuneDone latches.
- OnTick / OnTimer / OnChartEvent.
Training's own bar loops already honoured this (pass 1 per bar, passes 2/2.5/3
yield on a 120 ms budget); the warm-up scans did not, and they are the longest
uninterruptible stretches the EA has.
StopTraining() is unchanged: the operator's Stop still finalises synchronously.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 17:03:11 -04:00
|
|
|
}
|
feat: first-passage ladder + expectancy scan - price every geometry, not just the chosen one
Corrects the premise of the previous plan. Break-even is NOT a ceiling.
If the model shifts the win probability on the bars it selects from
p0 = m/(m+k) to p0 + d, then
EV = (p0+d)*k - (1-p0-d)*m = d*(k+m)
because p0*k - (1-p0)*m is zero by construction. The stop:target RATIO
is expectancy-neutral - a punishing break-even is exactly repaid by the
payoff - and only the real edge d and the TOTAL WIDTH (k+m) move EV.
Width matters because the spread is charged once per trade however wide
the barriers are, so a narrow barrier spends much of its own range on
costs. DeriveBarrierGeometry's own comment already said the ratio buys
nothing; the objective just never followed from it.
Blocker this had to solve first: m_excUpCache/m_excDownCache hold only
MAXIMUM travel each way, and a maximum cannot say which side was
reached FIRST - so any geometry other than the walked one was
undecidable on precisely the bars where both barriers were touched,
~28% of the sample.
- BARRIER_LADDER: per bar, the first-touch AGE for 8 travel distances
in each direction, filled during the walk the labels already run.
Cursors keep it O(1) amortised per walked bar rather than 16
comparisons. Levels are travel FROM ENTRY, not barrier prices, so one
ladder serves both directions and the spread is applied analytically
when a level converts back to an SL/TP multiple - storing prices
would need four ladders and bake today's spread into the cache.
Sized, invalidated and validity-gated with the label caches.
- ReportGeometryExpectancyScan: every ladder pair priced exactly off
that cache - width in ATR and in SPREADS (cost efficiency, knowable
without knowing d), break-even, both base rates, the share of bars
resolved inside the horizon, and EV per unit of edge. Compares the
widest resolvable pair against the quantile rule's pick.
MEASUREMENT ONLY - the quantile rule still chooses. Nothing here can
measure d, and width buys nothing if the wider target is less
predictable. Base rates are printed beside each break-even because a
persistent gap is DRIFT and must not be credited to the model.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-10 12:59:18 -04:00
|
|
|
if(bestT >= 0)
|
|
|
|
|
Print(ID + StringFormat(": barrier expectancy scan - on width alone the best resolvable pair is "
|
|
|
|
|
"stop %.2f*ATR target %.2f*ATR (width %.2f*ATR, %.1f spreads), against the "
|
|
|
|
|
"quantile rule's stop %.2f target %.2f (width %.2f*ATR, %.1f spreads) - a "
|
|
|
|
|
"%.2fx difference in EV per unit of edge. MEASUREMENT ONLY: the quantile "
|
|
|
|
|
"rule still chooses, because width buys nothing if the wider target is "
|
|
|
|
|
"less predictable, and this scan cannot see that.",
|
|
|
|
|
BARRIER_LADDER[bestS] + spreadAtr, BARRIER_LADDER[bestT] - spreadAtr,
|
|
|
|
|
bestWidth, (spreadAtr > 0.0 ? bestWidth / spreadAtr : 0.0),
|
|
|
|
|
m_derivedSlMult, m_derivedTpMult, m_derivedSlMult + m_derivedTpMult,
|
|
|
|
|
(spreadAtr > 0.0 ? (m_derivedSlMult + m_derivedTpMult) / spreadAtr : 0.0),
|
|
|
|
|
(m_derivedSlMult + m_derivedTpMult > 0.0
|
|
|
|
|
? bestWidth / (m_derivedSlMult + m_derivedTpMult) : 0.0)));
|
|
|
|
|
}
|
|
|
|
|
//+------------------------------------------------------------------+
|
fix(geometry): the reachability floor measured the WRONG WINDOW - my bug from bc57aca, and it cost real width
RECONCILED: the derivation reported "target reached on 17.7% of bars" while the
label cache reported Buy on 35.9%. Nothing was broken. They measure different
windows, and both are correct:
EXCURSION window ~12 bars (the SWING MEDIAN) - what m_excUpCache accumulates
over. Deliberately short: sizing a barrier off travel
measured over a horizon that itself scales with the barrier
is circular, and it ran away to 14-31*ATR on EURUSD/USDCAD
in 2026-08-07. That guard is correct and stays.
BARRIER horizon 64 bars - what the LABEL walk and the first-passage ladder
run over, and how long the EA actually holds the trade.
So `up >= target` is a 12-bar question and `label == Buy` is a 64-bar one, and
the second can freely exceed the first. TripleBarrierLabel gates the excursion
accumulation on `idx - t <= excWindow` while the barrier walk and the ladder run
the full horizon - the split is explicit and intentional.
THE BUG IS MINE. bc57aca's scale ladder tested reachability with `up[i] >= tp`,
i.e. it asked the 12-bar question about a 64-bar trade. That understates
reachability by ~2x, which is why EVERY wide rung was rejected and the geometry
fell back to the tightest rung at 1.61/3.21. The data supported considerably
wider; the test was just asking the wrong question.
FIX: LadderWinShare() reads the answer off the first-passage ladder - target
touched strictly before the stop, over the full horizon, tie to the stop. That
is the identical question the label walk asks, so the ladder share and the Buy
rate should now agree to within rung discretisation. Both legs snap to the
SMALLEST rung at or above the requested multiple (harder target, harder stop) so
the floor stays conservative.
Expect the scale ladder to select a WIDER rung on the next relabel. On this
data the excursion test read 17.7% at q50 where the true full-horizon share is
35.9%, so rungs that scored 8.1% and 2.8% were likely well above the floor.
ALSO:
- Window reconciliation now PRINTED every derivation: excursion travel share,
ladder win share, and the label cache's Buy share side by side, with the
ladder-vs-label gap flagged if it exceeds rung discretisation. Those two must
agree; if they ever stop agreeing, one of them is wrong and the line says so.
- Renamed tpReach/slReach -> tpTravel/slTravel and relabelled the log line. They
describe the EXCURSION window and are near-tautological there (a q50 stop is
exceeded by ~50% of bars); calling them "reached within the horizon" is what
made the two quantities look like one.
- BARRIER_MIN_TP_REACH_PCT is now BARRIER_MIN_REACH_FRACTION_OF_BE (0.60) x
break-even instead of a hardcoded 20.0. Break-even for 1:RR is 100/(1+RR), so
the absolute floor silently tightened as RR rose - 0.60x at RR=2 but 0.80x at
RR=3, penalising the user for asking for a bigger target. Evaluates to exactly
20.0% at the shipped RR=2, so this is a no-op today and correct if the knob
moves.
NOT COMPILED - user compiles.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 00:51:08 -04:00
|
|
|
//+------------------------------------------------------------------+
|
|
|
|
|
//| Share of the given bars that a LONG at (sl, tp) would have WON - |
|
|
|
|
|
//| target touched strictly before the stop - read off the |
|
|
|
|
|
//| first-passage ladder, which runs the FULL barrier horizon. |
|
|
|
|
|
//| |
|
|
|
|
|
//| THIS IS THE WINDOW THE TRADE ACTUALLY LIVES IN, and using the |
|
|
|
|
|
//| excursion cache instead was a real error (introduced bc57aca, |
|
|
|
|
|
//| found the same night). m_excUpCache accumulates only over |
|
|
|
|
|
//| excWindow = the SWING MEDIAN - deliberately, because sizing a |
|
|
|
|
|
//| barrier off travel measured over a horizon that itself scales |
|
|
|
|
|
//| with the barrier is circular and ran away to 14-31*ATR on |
|
|
|
|
|
//| EURUSD/USDCAD in 2026-08-07. That guard is correct and stays. |
|
|
|
|
|
//| But it makes the excursion cache the WRONG instrument for asking |
|
|
|
|
|
//| "would this target be reached", because the trade is held for |
|
|
|
|
|
//| m_barrierHorizonBars, not for the swing median. |
|
|
|
|
|
//| |
|
|
|
|
|
//| Measured on SP500 H4: swing median ~12 bars against a 64-bar |
|
|
|
|
|
//| horizon, so the excursion-based test understated reachability by |
|
|
|
|
|
//| ~2x (17.7% vs a true 35.9%) and rejected every wide rung of the |
|
|
|
|
|
//| scale ladder - which is exactly how the geometry came out at |
|
|
|
|
|
//| 1.61/3.21 when the data supported considerably wider. |
|
|
|
|
|
//| |
|
fix(labels): overlapping-label sample correction + horizon cap on the scale ladder
Three defects, all surfaced by the 2026-08-17 SP500 H4 run that shipped
stop 4.86 / target 9.71 (width 14.57*ATR, horizon 384).
1. EVERY STANDARD ERROR ASSUMED INDEPENDENT SAMPLES. Triple-barrier labels
started one per bar overlap by the label's lifespan, so n calls are worth
~n/L independent observations (Lopez de Prado, AFML ch. 4 - sample
uniqueness). All three sqrt(p(1-p)/n) sites divided by the RAW count.
The tell: the operating point's null-of-the-maximum gate is family-wise and
should fire on ~5% of eras under the null. Measured fire rates - PAI 47/73
(64%), ConvLSTM 9/24, LSTM 8/21 (38%), CONV 4/62 (6%). CONV, the only model
whose margin distribution admits few bins, sat on the null; the rest cleared
a bar that was too low by ~sqrt(L). PAI's deployed threshold consequently
alternated between the ENDS of its own range era to era (0.10 -> 0.88 ->
0.86 -> 0.66; coverage 16% <-> 73%).
TripleBarrierLabel now records when each label became KNOWABLE - the first
winning touch, or both stops, or the timeout - and the prebuild accumulates
the mean. EffectiveSampleSize() feeds the operating point, the member deploy
gate and the ensemble vote gate. Conservative by construction (n/L is an
upper bound on the damage); gates get harder, never easier.
2. THE SCALE LADDER RAN AWAY, again. Horizon scales as swingMedian*sl*tp, and
since 4d8cb08 reachability is measured OVER that horizon - so a wider rung
buys itself the time that makes it look reachable. Same target -> horizon ->
reach -> target loop the excursion window is kept short to avoid; fixing the
window confusion reopened it through the other door. It walked 128 -> 256 ->
384 bars and stopped at q90, the widest rung there is, with every rung
reading 39-48% against a 20% floor. A floor nothing fails selects nothing.
Rungs whose required horizon exceeds BARRIER_HORIZON_MAX are now rejected -
the same rule ReportGeometryExpectancyScan already applied. It was printing
the shipped pair as CLAMPED and disqualified ('h384!') two lines under the
deriver that chose it: two subsystems, one geometry, opposite verdicts.
3. THE RUNG SNAP DESTROYED THE RATIO IT WAS COMPARING. Both legs snapped
independently to the coarse first-passage grid, re-rating each candidate:
q90 4.86/9.71 -> 5.00/10.00 (2.00), q85 4.07/8.14 -> 5.00/10.00 (IDENTICAL
measurement), q75 3.07/6.13 -> 4.00/6.50 (1.63 - a nearer target). So the
ladder compared win shares taken at ratios from 1.63 to 2.17 and read the
differences as scale. It is why the reach column came out non-monotone in
width (q75 48.5% above q90 42.9%). The stop now snaps to its nearest rung in
log space and the target follows the ratio off it; the pair actually measured
is returned and logged, so a collision reads as a collision.
Also: LadderWinShare guarded against the conditional (fractal) geometry path,
which fills n from m_fracLegCount while leaving idxList empty - a latent
out-of-bounds on a currently-dead path.
New log lines: mean label lifespan and effective n on the label-cache line, the
required-vs-available horizon per rung, and the grid pair the reconciliation
actually measured (its tolerance now scales with the grid skew instead of a flat
5pp).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 01:12:05 -04:00
|
|
|
//| THE SNAP MUST PRESERVE THE RATIO, and the first version did not. |
|
|
|
|
|
//| Both legs used to snap independently to the smallest rung at or |
|
|
|
|
|
//| above the request, which silently re-rated each candidate: |
|
|
|
|
|
//| |
|
|
|
|
|
//| q90 4.86/9.71 -> 5.00/10.00 = 2.00 (asked 2.00) |
|
|
|
|
|
//| q85 4.07/8.14 -> 5.00/10.00 = 2.00 IDENTICAL to q90 |
|
|
|
|
|
//| q75 3.07/6.13 -> 4.00/ 6.50 = 1.63 a NEARER target |
|
|
|
|
|
//| |
|
|
|
|
|
//| So the scale ladder was comparing win shares measured at ratios |
|
|
|
|
|
//| from 1.63 to 2.17 and reading the differences as scale effects. |
|
|
|
|
|
//| It is why the reported reach column came out non-monotone in |
|
|
|
|
|
//| width (q75 48.5% ABOVE q90 42.9% on SP500 H4, 2026-08-17) - q75 |
|
|
|
|
|
//| was simply being tested against an easier target. |
|
|
|
|
|
//| |
|
|
|
|
|
//| Now the STOP snaps to its nearest rung and the target is taken |
|
|
|
|
|
//| relative to THAT, so the measured ratio is as close to the asked |
|
|
|
|
|
//| ratio as the grid can express. Rungs become comparable to each |
|
|
|
|
|
//| other, which is the only thing the ladder does with them. The |
|
|
|
|
|
//| pair actually measured is returned so the caller can print it: |
|
|
|
|
|
//| a collision (two quantiles landing on one grid pair) is a real |
|
|
|
|
|
//| limit of the instrument's resolution and must be visible, not |
|
|
|
|
|
//| hidden behind two identical-looking percentages. |
|
|
|
|
|
//| |
|
|
|
|
|
//| A target past the top rung returns 0 - unreachable as far as this |
|
|
|
|
|
//| instrument can measure, which is the honest answer now that the |
|
|
|
|
|
//| rung ceiling is 20*ATR. |
|
fix(geometry): the reachability floor measured the WRONG WINDOW - my bug from bc57aca, and it cost real width
RECONCILED: the derivation reported "target reached on 17.7% of bars" while the
label cache reported Buy on 35.9%. Nothing was broken. They measure different
windows, and both are correct:
EXCURSION window ~12 bars (the SWING MEDIAN) - what m_excUpCache accumulates
over. Deliberately short: sizing a barrier off travel
measured over a horizon that itself scales with the barrier
is circular, and it ran away to 14-31*ATR on EURUSD/USDCAD
in 2026-08-07. That guard is correct and stays.
BARRIER horizon 64 bars - what the LABEL walk and the first-passage ladder
run over, and how long the EA actually holds the trade.
So `up >= target` is a 12-bar question and `label == Buy` is a 64-bar one, and
the second can freely exceed the first. TripleBarrierLabel gates the excursion
accumulation on `idx - t <= excWindow` while the barrier walk and the ladder run
the full horizon - the split is explicit and intentional.
THE BUG IS MINE. bc57aca's scale ladder tested reachability with `up[i] >= tp`,
i.e. it asked the 12-bar question about a 64-bar trade. That understates
reachability by ~2x, which is why EVERY wide rung was rejected and the geometry
fell back to the tightest rung at 1.61/3.21. The data supported considerably
wider; the test was just asking the wrong question.
FIX: LadderWinShare() reads the answer off the first-passage ladder - target
touched strictly before the stop, over the full horizon, tie to the stop. That
is the identical question the label walk asks, so the ladder share and the Buy
rate should now agree to within rung discretisation. Both legs snap to the
SMALLEST rung at or above the requested multiple (harder target, harder stop) so
the floor stays conservative.
Expect the scale ladder to select a WIDER rung on the next relabel. On this
data the excursion test read 17.7% at q50 where the true full-horizon share is
35.9%, so rungs that scored 8.1% and 2.8% were likely well above the floor.
ALSO:
- Window reconciliation now PRINTED every derivation: excursion travel share,
ladder win share, and the label cache's Buy share side by side, with the
ladder-vs-label gap flagged if it exceeds rung discretisation. Those two must
agree; if they ever stop agreeing, one of them is wrong and the line says so.
- Renamed tpReach/slReach -> tpTravel/slTravel and relabelled the log line. They
describe the EXCURSION window and are near-tautological there (a q50 stop is
exceeded by ~50% of bars); calling them "reached within the horizon" is what
made the two quantities look like one.
- BARRIER_MIN_TP_REACH_PCT is now BARRIER_MIN_REACH_FRACTION_OF_BE (0.60) x
break-even instead of a hardcoded 20.0. Break-even for 1:RR is 100/(1+RR), so
the absolute floor silently tightened as RR rose - 0.60x at RR=2 but 0.80x at
RR=3, penalising the user for asking for a bigger target. Evaluates to exactly
20.0% at the shipped RR=2, so this is a no-op today and correct if the knob
moves.
NOT COMPILED - user compiles.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 00:51:08 -04:00
|
|
|
//+------------------------------------------------------------------+
|
fix(labels): overlapping-label sample correction + horizon cap on the scale ladder
Three defects, all surfaced by the 2026-08-17 SP500 H4 run that shipped
stop 4.86 / target 9.71 (width 14.57*ATR, horizon 384).
1. EVERY STANDARD ERROR ASSUMED INDEPENDENT SAMPLES. Triple-barrier labels
started one per bar overlap by the label's lifespan, so n calls are worth
~n/L independent observations (Lopez de Prado, AFML ch. 4 - sample
uniqueness). All three sqrt(p(1-p)/n) sites divided by the RAW count.
The tell: the operating point's null-of-the-maximum gate is family-wise and
should fire on ~5% of eras under the null. Measured fire rates - PAI 47/73
(64%), ConvLSTM 9/24, LSTM 8/21 (38%), CONV 4/62 (6%). CONV, the only model
whose margin distribution admits few bins, sat on the null; the rest cleared
a bar that was too low by ~sqrt(L). PAI's deployed threshold consequently
alternated between the ENDS of its own range era to era (0.10 -> 0.88 ->
0.86 -> 0.66; coverage 16% <-> 73%).
TripleBarrierLabel now records when each label became KNOWABLE - the first
winning touch, or both stops, or the timeout - and the prebuild accumulates
the mean. EffectiveSampleSize() feeds the operating point, the member deploy
gate and the ensemble vote gate. Conservative by construction (n/L is an
upper bound on the damage); gates get harder, never easier.
2. THE SCALE LADDER RAN AWAY, again. Horizon scales as swingMedian*sl*tp, and
since 4d8cb08 reachability is measured OVER that horizon - so a wider rung
buys itself the time that makes it look reachable. Same target -> horizon ->
reach -> target loop the excursion window is kept short to avoid; fixing the
window confusion reopened it through the other door. It walked 128 -> 256 ->
384 bars and stopped at q90, the widest rung there is, with every rung
reading 39-48% against a 20% floor. A floor nothing fails selects nothing.
Rungs whose required horizon exceeds BARRIER_HORIZON_MAX are now rejected -
the same rule ReportGeometryExpectancyScan already applied. It was printing
the shipped pair as CLAMPED and disqualified ('h384!') two lines under the
deriver that chose it: two subsystems, one geometry, opposite verdicts.
3. THE RUNG SNAP DESTROYED THE RATIO IT WAS COMPARING. Both legs snapped
independently to the coarse first-passage grid, re-rating each candidate:
q90 4.86/9.71 -> 5.00/10.00 (2.00), q85 4.07/8.14 -> 5.00/10.00 (IDENTICAL
measurement), q75 3.07/6.13 -> 4.00/6.50 (1.63 - a nearer target). So the
ladder compared win shares taken at ratios from 1.63 to 2.17 and read the
differences as scale. It is why the reach column came out non-monotone in
width (q75 48.5% above q90 42.9%). The stop now snaps to its nearest rung in
log space and the target follows the ratio off it; the pair actually measured
is returned and logged, so a collision reads as a collision.
Also: LadderWinShare guarded against the conditional (fractal) geometry path,
which fills n from m_fracLegCount while leaving idxList empty - a latent
out-of-bounds on a currently-dead path.
New log lines: mean label lifespan and effective n on the label-cache line, the
required-vs-available horizon per rung, and the grid pair the reconciliation
actually measured (its tolerance now scales with the grid skew instead of a flat
5pp).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 01:12:05 -04:00
|
|
|
double CExpertSignalAIBase::LadderWinShare(const int &idxList[], int n, double sl, double tp,
|
|
|
|
|
double &effSl, double &effTp)
|
fix(geometry): the reachability floor measured the WRONG WINDOW - my bug from bc57aca, and it cost real width
RECONCILED: the derivation reported "target reached on 17.7% of bars" while the
label cache reported Buy on 35.9%. Nothing was broken. They measure different
windows, and both are correct:
EXCURSION window ~12 bars (the SWING MEDIAN) - what m_excUpCache accumulates
over. Deliberately short: sizing a barrier off travel
measured over a horizon that itself scales with the barrier
is circular, and it ran away to 14-31*ATR on EURUSD/USDCAD
in 2026-08-07. That guard is correct and stays.
BARRIER horizon 64 bars - what the LABEL walk and the first-passage ladder
run over, and how long the EA actually holds the trade.
So `up >= target` is a 12-bar question and `label == Buy` is a 64-bar one, and
the second can freely exceed the first. TripleBarrierLabel gates the excursion
accumulation on `idx - t <= excWindow` while the barrier walk and the ladder run
the full horizon - the split is explicit and intentional.
THE BUG IS MINE. bc57aca's scale ladder tested reachability with `up[i] >= tp`,
i.e. it asked the 12-bar question about a 64-bar trade. That understates
reachability by ~2x, which is why EVERY wide rung was rejected and the geometry
fell back to the tightest rung at 1.61/3.21. The data supported considerably
wider; the test was just asking the wrong question.
FIX: LadderWinShare() reads the answer off the first-passage ladder - target
touched strictly before the stop, over the full horizon, tie to the stop. That
is the identical question the label walk asks, so the ladder share and the Buy
rate should now agree to within rung discretisation. Both legs snap to the
SMALLEST rung at or above the requested multiple (harder target, harder stop) so
the floor stays conservative.
Expect the scale ladder to select a WIDER rung on the next relabel. On this
data the excursion test read 17.7% at q50 where the true full-horizon share is
35.9%, so rungs that scored 8.1% and 2.8% were likely well above the floor.
ALSO:
- Window reconciliation now PRINTED every derivation: excursion travel share,
ladder win share, and the label cache's Buy share side by side, with the
ladder-vs-label gap flagged if it exceeds rung discretisation. Those two must
agree; if they ever stop agreeing, one of them is wrong and the line says so.
- Renamed tpReach/slReach -> tpTravel/slTravel and relabelled the log line. They
describe the EXCURSION window and are near-tautological there (a q50 stop is
exceeded by ~50% of bars); calling them "reached within the horizon" is what
made the two quantities look like one.
- BARRIER_MIN_TP_REACH_PCT is now BARRIER_MIN_REACH_FRACTION_OF_BE (0.60) x
break-even instead of a hardcoded 20.0. Break-even for 1:RR is 100/(1+RR), so
the absolute floor silently tightened as RR rose - 0.60x at RR=2 but 0.80x at
RR=3, penalising the user for asking for a bigger target. Evaluates to exactly
20.0% at the shipped RR=2, so this is a no-op today and correct if the knob
moves.
NOT COMPILED - user compiles.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 00:51:08 -04:00
|
|
|
{
|
fix(labels): overlapping-label sample correction + horizon cap on the scale ladder
Three defects, all surfaced by the 2026-08-17 SP500 H4 run that shipped
stop 4.86 / target 9.71 (width 14.57*ATR, horizon 384).
1. EVERY STANDARD ERROR ASSUMED INDEPENDENT SAMPLES. Triple-barrier labels
started one per bar overlap by the label's lifespan, so n calls are worth
~n/L independent observations (Lopez de Prado, AFML ch. 4 - sample
uniqueness). All three sqrt(p(1-p)/n) sites divided by the RAW count.
The tell: the operating point's null-of-the-maximum gate is family-wise and
should fire on ~5% of eras under the null. Measured fire rates - PAI 47/73
(64%), ConvLSTM 9/24, LSTM 8/21 (38%), CONV 4/62 (6%). CONV, the only model
whose margin distribution admits few bins, sat on the null; the rest cleared
a bar that was too low by ~sqrt(L). PAI's deployed threshold consequently
alternated between the ENDS of its own range era to era (0.10 -> 0.88 ->
0.86 -> 0.66; coverage 16% <-> 73%).
TripleBarrierLabel now records when each label became KNOWABLE - the first
winning touch, or both stops, or the timeout - and the prebuild accumulates
the mean. EffectiveSampleSize() feeds the operating point, the member deploy
gate and the ensemble vote gate. Conservative by construction (n/L is an
upper bound on the damage); gates get harder, never easier.
2. THE SCALE LADDER RAN AWAY, again. Horizon scales as swingMedian*sl*tp, and
since 4d8cb08 reachability is measured OVER that horizon - so a wider rung
buys itself the time that makes it look reachable. Same target -> horizon ->
reach -> target loop the excursion window is kept short to avoid; fixing the
window confusion reopened it through the other door. It walked 128 -> 256 ->
384 bars and stopped at q90, the widest rung there is, with every rung
reading 39-48% against a 20% floor. A floor nothing fails selects nothing.
Rungs whose required horizon exceeds BARRIER_HORIZON_MAX are now rejected -
the same rule ReportGeometryExpectancyScan already applied. It was printing
the shipped pair as CLAMPED and disqualified ('h384!') two lines under the
deriver that chose it: two subsystems, one geometry, opposite verdicts.
3. THE RUNG SNAP DESTROYED THE RATIO IT WAS COMPARING. Both legs snapped
independently to the coarse first-passage grid, re-rating each candidate:
q90 4.86/9.71 -> 5.00/10.00 (2.00), q85 4.07/8.14 -> 5.00/10.00 (IDENTICAL
measurement), q75 3.07/6.13 -> 4.00/6.50 (1.63 - a nearer target). So the
ladder compared win shares taken at ratios from 1.63 to 2.17 and read the
differences as scale. It is why the reach column came out non-monotone in
width (q75 48.5% above q90 42.9%). The stop now snaps to its nearest rung in
log space and the target follows the ratio off it; the pair actually measured
is returned and logged, so a collision reads as a collision.
Also: LadderWinShare guarded against the conditional (fractal) geometry path,
which fills n from m_fracLegCount while leaving idxList empty - a latent
out-of-bounds on a currently-dead path.
New log lines: mean label lifespan and effective n on the label-cache line, the
required-vs-available horizon per rung, and the grid pair the reconciliation
actually measured (its tolerance now scales with the grid skew instead of a flat
5pp).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 01:12:05 -04:00
|
|
|
effSl = 0.0;
|
|
|
|
|
effTp = 0.0;
|
feat(measurement): fix zero-skill denominator, publish the deploy bar, measure lifespan per rung, add a MEASURE scale objective
The last run could not have demonstrated an edge either way, and nothing in the
log said so. Four changes so it does.
1. THE ZERO-SKILL LINE DIVIDED BY THE WRONG DENOMINATOR. m_oosWinLongTotal resets
every era; m_oosSamples only resets on a full model reset. So 'always-long %'
decayed as ~1/era: a run whose true rate is 37% printed 1.2% at era 33 and
0.0% at era 2219. This is the SAME bug already found and fixed for
logBuyPredPct thirty lines above ('era-15 Buy:2% that was really ~30%'), left
in the one line whose whole job is to be the reference every other number is
read against. Correct at era 1, wrong everywhere after - including the '62%
zero-skill' figure in the 2026-08-16 notes. Now per-era, and always-short is
finally readable.
2. THE DEPLOY GATE STATES ITS OWN BAR. 'edge -1pp' era after era cannot separate
'short by a hair' from 'short by an amount no strategy could cover'. The era
line now prints the required win rate, the SE, the effective n and the
lifespan it was deflated by; above 100% it says UNREACHABLE. At 4,738 OOS bars
and L=75.6 there are ~63 independent observations, putting the bar near 66% at
typical coverage.
3. LIFESPAN MEASURED PER RUNG. The first-passage cache already stores touch ages
at every ladder level, so each candidate geometry's resolution time is
readable without training on it - L-vs-width becomes a measurement across the
whole ladder in ONE run rather than a second chart. Each rung reports L,
n_eff, min provable edge and min provable EV.
4. SCALE OBJECTIVE IS PHASE-AWARE, defaulting to MEASURE. Width and detectability
are opposed: labels overlap by L, L grows like m*k = width^2 at fixed ratio,
so min provable EV ~ width^2 while the cost saving from width is only linear.
Doubling width quadruples the smallest EV you can prove. DEPLOY (widest that
clears reachability) is right once an edge is known; MEASURE (narrowest that
keeps round-trip spread under BARRIER_MAX_COST_FRACTION_PCT) is right while it
still has to be shown. The direction does not depend on the exponent, and
item 3 makes the exponent checkable.
Fixed in review: m_lastRungLifespan is cleared on every LadderWinShare entry or a
rejected rung reports the previous rung's lifespan as its own; per-rung
detectability is labelled IS-sample based (the deriver may not see the holdout),
so absolute figures are optimistic while the ranking is unaffected.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 10:37:05 -04:00
|
|
|
//--- Cleared with the out-params and for the same reason: every early return below leaves this
|
|
|
|
|
//--- untouched otherwise, and the caller loops over rungs - so a rejected rung would report the
|
|
|
|
|
//--- PREVIOUS rung's lifespan as its own. Same class of bug as the excursion publication at the top
|
|
|
|
|
//--- of TripleBarrierLabel; 0 reads as "not measured".
|
|
|
|
|
m_lastRungLifespan = 0.0;
|
fix(labels): overlapping-label sample correction + horizon cap on the scale ladder
Three defects, all surfaced by the 2026-08-17 SP500 H4 run that shipped
stop 4.86 / target 9.71 (width 14.57*ATR, horizon 384).
1. EVERY STANDARD ERROR ASSUMED INDEPENDENT SAMPLES. Triple-barrier labels
started one per bar overlap by the label's lifespan, so n calls are worth
~n/L independent observations (Lopez de Prado, AFML ch. 4 - sample
uniqueness). All three sqrt(p(1-p)/n) sites divided by the RAW count.
The tell: the operating point's null-of-the-maximum gate is family-wise and
should fire on ~5% of eras under the null. Measured fire rates - PAI 47/73
(64%), ConvLSTM 9/24, LSTM 8/21 (38%), CONV 4/62 (6%). CONV, the only model
whose margin distribution admits few bins, sat on the null; the rest cleared
a bar that was too low by ~sqrt(L). PAI's deployed threshold consequently
alternated between the ENDS of its own range era to era (0.10 -> 0.88 ->
0.86 -> 0.66; coverage 16% <-> 73%).
TripleBarrierLabel now records when each label became KNOWABLE - the first
winning touch, or both stops, or the timeout - and the prebuild accumulates
the mean. EffectiveSampleSize() feeds the operating point, the member deploy
gate and the ensemble vote gate. Conservative by construction (n/L is an
upper bound on the damage); gates get harder, never easier.
2. THE SCALE LADDER RAN AWAY, again. Horizon scales as swingMedian*sl*tp, and
since 4d8cb08 reachability is measured OVER that horizon - so a wider rung
buys itself the time that makes it look reachable. Same target -> horizon ->
reach -> target loop the excursion window is kept short to avoid; fixing the
window confusion reopened it through the other door. It walked 128 -> 256 ->
384 bars and stopped at q90, the widest rung there is, with every rung
reading 39-48% against a 20% floor. A floor nothing fails selects nothing.
Rungs whose required horizon exceeds BARRIER_HORIZON_MAX are now rejected -
the same rule ReportGeometryExpectancyScan already applied. It was printing
the shipped pair as CLAMPED and disqualified ('h384!') two lines under the
deriver that chose it: two subsystems, one geometry, opposite verdicts.
3. THE RUNG SNAP DESTROYED THE RATIO IT WAS COMPARING. Both legs snapped
independently to the coarse first-passage grid, re-rating each candidate:
q90 4.86/9.71 -> 5.00/10.00 (2.00), q85 4.07/8.14 -> 5.00/10.00 (IDENTICAL
measurement), q75 3.07/6.13 -> 4.00/6.50 (1.63 - a nearer target). So the
ladder compared win shares taken at ratios from 1.63 to 2.17 and read the
differences as scale. It is why the reach column came out non-monotone in
width (q75 48.5% above q90 42.9%). The stop now snaps to its nearest rung in
log space and the target follows the ratio off it; the pair actually measured
is returned and logged, so a collision reads as a collision.
Also: LadderWinShare guarded against the conditional (fractal) geometry path,
which fills n from m_fracLegCount while leaving idxList empty - a latent
out-of-bounds on a currently-dead path.
New log lines: mean label lifespan and effective n on the label-cache line, the
required-vs-available horizon per rung, and the grid pair the reconciliation
actually measured (its tolerance now scales with the grid skew instead of a flat
5pp).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 01:12:05 -04:00
|
|
|
if(sl <= 0.0 || tp <= 0.0 || n <= 0)
|
|
|
|
|
return 0.0;
|
|
|
|
|
//--- n is the EXCURSION sample size, which is not always idxList's size: the conditional (fractal)
|
|
|
|
|
//--- geometry path fills up[]/dn[] from m_fracLegFav with n = m_fracLegCount while leaving idxList
|
|
|
|
|
//--- empty, because leg-scoped excursions carry no bar index to look a first-passage row up by. The
|
|
|
|
|
//--- loop below would then index a zero-length array. Answer "not measurable" rather than fault:
|
|
|
|
|
//--- reachability off the ladder is undefined without the bar indices, and the caller's floor treats
|
|
|
|
|
//--- 0 as "this rung does not qualify", which is the correct handling of an unmeasurable rung.
|
|
|
|
|
if(ArraySize(idxList) < n)
|
|
|
|
|
return 0.0;
|
|
|
|
|
//--- Stop leg: NEAREST rung, in log space because the ladder is roughly geometric and a linear
|
|
|
|
|
//--- "nearest" would bias every choice toward the coarse upper end.
|
|
|
|
|
int dnLvl = -1;
|
|
|
|
|
double bestErr = DBL_MAX;
|
|
|
|
|
for(int k = 0; k < BARRIER_LADDER_COUNT; k++)
|
|
|
|
|
{
|
|
|
|
|
double err = MathAbs(MathLog(BARRIER_LADDER[k] / sl));
|
|
|
|
|
if(err < bestErr)
|
|
|
|
|
{
|
|
|
|
|
bestErr = err;
|
fix(geometry): the reachability floor measured the WRONG WINDOW - my bug from bc57aca, and it cost real width
RECONCILED: the derivation reported "target reached on 17.7% of bars" while the
label cache reported Buy on 35.9%. Nothing was broken. They measure different
windows, and both are correct:
EXCURSION window ~12 bars (the SWING MEDIAN) - what m_excUpCache accumulates
over. Deliberately short: sizing a barrier off travel
measured over a horizon that itself scales with the barrier
is circular, and it ran away to 14-31*ATR on EURUSD/USDCAD
in 2026-08-07. That guard is correct and stays.
BARRIER horizon 64 bars - what the LABEL walk and the first-passage ladder
run over, and how long the EA actually holds the trade.
So `up >= target` is a 12-bar question and `label == Buy` is a 64-bar one, and
the second can freely exceed the first. TripleBarrierLabel gates the excursion
accumulation on `idx - t <= excWindow` while the barrier walk and the ladder run
the full horizon - the split is explicit and intentional.
THE BUG IS MINE. bc57aca's scale ladder tested reachability with `up[i] >= tp`,
i.e. it asked the 12-bar question about a 64-bar trade. That understates
reachability by ~2x, which is why EVERY wide rung was rejected and the geometry
fell back to the tightest rung at 1.61/3.21. The data supported considerably
wider; the test was just asking the wrong question.
FIX: LadderWinShare() reads the answer off the first-passage ladder - target
touched strictly before the stop, over the full horizon, tie to the stop. That
is the identical question the label walk asks, so the ladder share and the Buy
rate should now agree to within rung discretisation. Both legs snap to the
SMALLEST rung at or above the requested multiple (harder target, harder stop) so
the floor stays conservative.
Expect the scale ladder to select a WIDER rung on the next relabel. On this
data the excursion test read 17.7% at q50 where the true full-horizon share is
35.9%, so rungs that scored 8.1% and 2.8% were likely well above the floor.
ALSO:
- Window reconciliation now PRINTED every derivation: excursion travel share,
ladder win share, and the label cache's Buy share side by side, with the
ladder-vs-label gap flagged if it exceeds rung discretisation. Those two must
agree; if they ever stop agreeing, one of them is wrong and the line says so.
- Renamed tpReach/slReach -> tpTravel/slTravel and relabelled the log line. They
describe the EXCURSION window and are near-tautological there (a q50 stop is
exceeded by ~50% of bars); calling them "reached within the horizon" is what
made the two quantities look like one.
- BARRIER_MIN_TP_REACH_PCT is now BARRIER_MIN_REACH_FRACTION_OF_BE (0.60) x
break-even instead of a hardcoded 20.0. Break-even for 1:RR is 100/(1+RR), so
the absolute floor silently tightened as RR rose - 0.60x at RR=2 but 0.80x at
RR=3, penalising the user for asking for a bigger target. Evaluates to exactly
20.0% at the shipped RR=2, so this is a no-op today and correct if the knob
moves.
NOT COMPILED - user compiles.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 00:51:08 -04:00
|
|
|
dnLvl = k;
|
fix(labels): overlapping-label sample correction + horizon cap on the scale ladder
Three defects, all surfaced by the 2026-08-17 SP500 H4 run that shipped
stop 4.86 / target 9.71 (width 14.57*ATR, horizon 384).
1. EVERY STANDARD ERROR ASSUMED INDEPENDENT SAMPLES. Triple-barrier labels
started one per bar overlap by the label's lifespan, so n calls are worth
~n/L independent observations (Lopez de Prado, AFML ch. 4 - sample
uniqueness). All three sqrt(p(1-p)/n) sites divided by the RAW count.
The tell: the operating point's null-of-the-maximum gate is family-wise and
should fire on ~5% of eras under the null. Measured fire rates - PAI 47/73
(64%), ConvLSTM 9/24, LSTM 8/21 (38%), CONV 4/62 (6%). CONV, the only model
whose margin distribution admits few bins, sat on the null; the rest cleared
a bar that was too low by ~sqrt(L). PAI's deployed threshold consequently
alternated between the ENDS of its own range era to era (0.10 -> 0.88 ->
0.86 -> 0.66; coverage 16% <-> 73%).
TripleBarrierLabel now records when each label became KNOWABLE - the first
winning touch, or both stops, or the timeout - and the prebuild accumulates
the mean. EffectiveSampleSize() feeds the operating point, the member deploy
gate and the ensemble vote gate. Conservative by construction (n/L is an
upper bound on the damage); gates get harder, never easier.
2. THE SCALE LADDER RAN AWAY, again. Horizon scales as swingMedian*sl*tp, and
since 4d8cb08 reachability is measured OVER that horizon - so a wider rung
buys itself the time that makes it look reachable. Same target -> horizon ->
reach -> target loop the excursion window is kept short to avoid; fixing the
window confusion reopened it through the other door. It walked 128 -> 256 ->
384 bars and stopped at q90, the widest rung there is, with every rung
reading 39-48% against a 20% floor. A floor nothing fails selects nothing.
Rungs whose required horizon exceeds BARRIER_HORIZON_MAX are now rejected -
the same rule ReportGeometryExpectancyScan already applied. It was printing
the shipped pair as CLAMPED and disqualified ('h384!') two lines under the
deriver that chose it: two subsystems, one geometry, opposite verdicts.
3. THE RUNG SNAP DESTROYED THE RATIO IT WAS COMPARING. Both legs snapped
independently to the coarse first-passage grid, re-rating each candidate:
q90 4.86/9.71 -> 5.00/10.00 (2.00), q85 4.07/8.14 -> 5.00/10.00 (IDENTICAL
measurement), q75 3.07/6.13 -> 4.00/6.50 (1.63 - a nearer target). So the
ladder compared win shares taken at ratios from 1.63 to 2.17 and read the
differences as scale. It is why the reach column came out non-monotone in
width (q75 48.5% above q90 42.9%). The stop now snaps to its nearest rung in
log space and the target follows the ratio off it; the pair actually measured
is returned and logged, so a collision reads as a collision.
Also: LadderWinShare guarded against the conditional (fractal) geometry path,
which fills n from m_fracLegCount while leaving idxList empty - a latent
out-of-bounds on a currently-dead path.
New log lines: mean label lifespan and effective n on the label-cache line, the
required-vs-available horizon per rung, and the grid pair the reconciliation
actually measured (its tolerance now scales with the grid skew instead of a flat
5pp).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 01:12:05 -04:00
|
|
|
}
|
|
|
|
|
}
|
|
|
|
|
if(dnLvl < 0)
|
|
|
|
|
return 0.0;
|
|
|
|
|
//--- Target leg: nearest rung to the ratio applied to the SNAPPED stop, so the pair that gets measured
|
|
|
|
|
//--- is a 1:RR pair on the grid rather than the requested pair re-rated by two independent roundings.
|
|
|
|
|
double wantTp = (BARRIER_LADDER[dnLvl] / sl) * tp;
|
|
|
|
|
int upLvl = -1;
|
|
|
|
|
bestErr = DBL_MAX;
|
|
|
|
|
for(int k = 0; k < BARRIER_LADDER_COUNT; k++)
|
|
|
|
|
{
|
|
|
|
|
double err = MathAbs(MathLog(BARRIER_LADDER[k] / wantTp));
|
|
|
|
|
if(err < bestErr)
|
|
|
|
|
{
|
|
|
|
|
bestErr = err;
|
|
|
|
|
upLvl = k;
|
|
|
|
|
}
|
|
|
|
|
}
|
|
|
|
|
if(upLvl < 0)
|
|
|
|
|
return 0.0;
|
|
|
|
|
//--- The grid cannot express this target at all (it sits past the top rung and the nearest rung is a
|
|
|
|
|
//--- materially different trade). Refuse rather than silently measure something else - the caller's
|
|
|
|
|
//--- floor is supposed to reject exactly this case.
|
|
|
|
|
if(wantTp > BARRIER_LADDER[BARRIER_LADDER_COUNT - 1])
|
fix(geometry): the reachability floor measured the WRONG WINDOW - my bug from bc57aca, and it cost real width
RECONCILED: the derivation reported "target reached on 17.7% of bars" while the
label cache reported Buy on 35.9%. Nothing was broken. They measure different
windows, and both are correct:
EXCURSION window ~12 bars (the SWING MEDIAN) - what m_excUpCache accumulates
over. Deliberately short: sizing a barrier off travel
measured over a horizon that itself scales with the barrier
is circular, and it ran away to 14-31*ATR on EURUSD/USDCAD
in 2026-08-07. That guard is correct and stays.
BARRIER horizon 64 bars - what the LABEL walk and the first-passage ladder
run over, and how long the EA actually holds the trade.
So `up >= target` is a 12-bar question and `label == Buy` is a 64-bar one, and
the second can freely exceed the first. TripleBarrierLabel gates the excursion
accumulation on `idx - t <= excWindow` while the barrier walk and the ladder run
the full horizon - the split is explicit and intentional.
THE BUG IS MINE. bc57aca's scale ladder tested reachability with `up[i] >= tp`,
i.e. it asked the 12-bar question about a 64-bar trade. That understates
reachability by ~2x, which is why EVERY wide rung was rejected and the geometry
fell back to the tightest rung at 1.61/3.21. The data supported considerably
wider; the test was just asking the wrong question.
FIX: LadderWinShare() reads the answer off the first-passage ladder - target
touched strictly before the stop, over the full horizon, tie to the stop. That
is the identical question the label walk asks, so the ladder share and the Buy
rate should now agree to within rung discretisation. Both legs snap to the
SMALLEST rung at or above the requested multiple (harder target, harder stop) so
the floor stays conservative.
Expect the scale ladder to select a WIDER rung on the next relabel. On this
data the excursion test read 17.7% at q50 where the true full-horizon share is
35.9%, so rungs that scored 8.1% and 2.8% were likely well above the floor.
ALSO:
- Window reconciliation now PRINTED every derivation: excursion travel share,
ladder win share, and the label cache's Buy share side by side, with the
ladder-vs-label gap flagged if it exceeds rung discretisation. Those two must
agree; if they ever stop agreeing, one of them is wrong and the line says so.
- Renamed tpReach/slReach -> tpTravel/slTravel and relabelled the log line. They
describe the EXCURSION window and are near-tautological there (a q50 stop is
exceeded by ~50% of bars); calling them "reached within the horizon" is what
made the two quantities look like one.
- BARRIER_MIN_TP_REACH_PCT is now BARRIER_MIN_REACH_FRACTION_OF_BE (0.60) x
break-even instead of a hardcoded 20.0. Break-even for 1:RR is 100/(1+RR), so
the absolute floor silently tightened as RR rose - 0.60x at RR=2 but 0.80x at
RR=3, penalising the user for asking for a bigger target. Evaluates to exactly
20.0% at the shipped RR=2, so this is a no-op today and correct if the knob
moves.
NOT COMPILED - user compiles.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 00:51:08 -04:00
|
|
|
return 0.0;
|
fix(labels): overlapping-label sample correction + horizon cap on the scale ladder
Three defects, all surfaced by the 2026-08-17 SP500 H4 run that shipped
stop 4.86 / target 9.71 (width 14.57*ATR, horizon 384).
1. EVERY STANDARD ERROR ASSUMED INDEPENDENT SAMPLES. Triple-barrier labels
started one per bar overlap by the label's lifespan, so n calls are worth
~n/L independent observations (Lopez de Prado, AFML ch. 4 - sample
uniqueness). All three sqrt(p(1-p)/n) sites divided by the RAW count.
The tell: the operating point's null-of-the-maximum gate is family-wise and
should fire on ~5% of eras under the null. Measured fire rates - PAI 47/73
(64%), ConvLSTM 9/24, LSTM 8/21 (38%), CONV 4/62 (6%). CONV, the only model
whose margin distribution admits few bins, sat on the null; the rest cleared
a bar that was too low by ~sqrt(L). PAI's deployed threshold consequently
alternated between the ENDS of its own range era to era (0.10 -> 0.88 ->
0.86 -> 0.66; coverage 16% <-> 73%).
TripleBarrierLabel now records when each label became KNOWABLE - the first
winning touch, or both stops, or the timeout - and the prebuild accumulates
the mean. EffectiveSampleSize() feeds the operating point, the member deploy
gate and the ensemble vote gate. Conservative by construction (n/L is an
upper bound on the damage); gates get harder, never easier.
2. THE SCALE LADDER RAN AWAY, again. Horizon scales as swingMedian*sl*tp, and
since 4d8cb08 reachability is measured OVER that horizon - so a wider rung
buys itself the time that makes it look reachable. Same target -> horizon ->
reach -> target loop the excursion window is kept short to avoid; fixing the
window confusion reopened it through the other door. It walked 128 -> 256 ->
384 bars and stopped at q90, the widest rung there is, with every rung
reading 39-48% against a 20% floor. A floor nothing fails selects nothing.
Rungs whose required horizon exceeds BARRIER_HORIZON_MAX are now rejected -
the same rule ReportGeometryExpectancyScan already applied. It was printing
the shipped pair as CLAMPED and disqualified ('h384!') two lines under the
deriver that chose it: two subsystems, one geometry, opposite verdicts.
3. THE RUNG SNAP DESTROYED THE RATIO IT WAS COMPARING. Both legs snapped
independently to the coarse first-passage grid, re-rating each candidate:
q90 4.86/9.71 -> 5.00/10.00 (2.00), q85 4.07/8.14 -> 5.00/10.00 (IDENTICAL
measurement), q75 3.07/6.13 -> 4.00/6.50 (1.63 - a nearer target). So the
ladder compared win shares taken at ratios from 1.63 to 2.17 and read the
differences as scale. It is why the reach column came out non-monotone in
width (q75 48.5% above q90 42.9%). The stop now snaps to its nearest rung in
log space and the target follows the ratio off it; the pair actually measured
is returned and logged, so a collision reads as a collision.
Also: LadderWinShare guarded against the conditional (fractal) geometry path,
which fills n from m_fracLegCount while leaving idxList empty - a latent
out-of-bounds on a currently-dead path.
New log lines: mean label lifespan and effective n on the label-cache line, the
required-vs-available horizon per rung, and the grid pair the reconciliation
actually measured (its tolerance now scales with the grid skew instead of a flat
5pp).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 01:12:05 -04:00
|
|
|
effSl = BARRIER_LADDER[dnLvl];
|
|
|
|
|
effTp = BARRIER_LADDER[upLvl];
|
fix(geometry): the reachability floor measured the WRONG WINDOW - my bug from bc57aca, and it cost real width
RECONCILED: the derivation reported "target reached on 17.7% of bars" while the
label cache reported Buy on 35.9%. Nothing was broken. They measure different
windows, and both are correct:
EXCURSION window ~12 bars (the SWING MEDIAN) - what m_excUpCache accumulates
over. Deliberately short: sizing a barrier off travel
measured over a horizon that itself scales with the barrier
is circular, and it ran away to 14-31*ATR on EURUSD/USDCAD
in 2026-08-07. That guard is correct and stays.
BARRIER horizon 64 bars - what the LABEL walk and the first-passage ladder
run over, and how long the EA actually holds the trade.
So `up >= target` is a 12-bar question and `label == Buy` is a 64-bar one, and
the second can freely exceed the first. TripleBarrierLabel gates the excursion
accumulation on `idx - t <= excWindow` while the barrier walk and the ladder run
the full horizon - the split is explicit and intentional.
THE BUG IS MINE. bc57aca's scale ladder tested reachability with `up[i] >= tp`,
i.e. it asked the 12-bar question about a 64-bar trade. That understates
reachability by ~2x, which is why EVERY wide rung was rejected and the geometry
fell back to the tightest rung at 1.61/3.21. The data supported considerably
wider; the test was just asking the wrong question.
FIX: LadderWinShare() reads the answer off the first-passage ladder - target
touched strictly before the stop, over the full horizon, tie to the stop. That
is the identical question the label walk asks, so the ladder share and the Buy
rate should now agree to within rung discretisation. Both legs snap to the
SMALLEST rung at or above the requested multiple (harder target, harder stop) so
the floor stays conservative.
Expect the scale ladder to select a WIDER rung on the next relabel. On this
data the excursion test read 17.7% at q50 where the true full-horizon share is
35.9%, so rungs that scored 8.1% and 2.8% were likely well above the floor.
ALSO:
- Window reconciliation now PRINTED every derivation: excursion travel share,
ladder win share, and the label cache's Buy share side by side, with the
ladder-vs-label gap flagged if it exceeds rung discretisation. Those two must
agree; if they ever stop agreeing, one of them is wrong and the line says so.
- Renamed tpReach/slReach -> tpTravel/slTravel and relabelled the log line. They
describe the EXCURSION window and are near-tautological there (a q50 stop is
exceeded by ~50% of bars); calling them "reached within the horizon" is what
made the two quantities look like one.
- BARRIER_MIN_TP_REACH_PCT is now BARRIER_MIN_REACH_FRACTION_OF_BE (0.60) x
break-even instead of a hardcoded 20.0. Break-even for 1:RR is 100/(1+RR), so
the absolute floor silently tightened as RR rose - 0.60x at RR=2 but 0.80x at
RR=3, penalising the user for asking for a bigger target. Evaluates to exactly
20.0% at the shipped RR=2, so this is a no-op today and correct if the knob
moves.
NOT COMPILED - user compiles.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 00:51:08 -04:00
|
|
|
int won = 0, seen = 0;
|
feat(measurement): fix zero-skill denominator, publish the deploy bar, measure lifespan per rung, add a MEASURE scale objective
The last run could not have demonstrated an edge either way, and nothing in the
log said so. Four changes so it does.
1. THE ZERO-SKILL LINE DIVIDED BY THE WRONG DENOMINATOR. m_oosWinLongTotal resets
every era; m_oosSamples only resets on a full model reset. So 'always-long %'
decayed as ~1/era: a run whose true rate is 37% printed 1.2% at era 33 and
0.0% at era 2219. This is the SAME bug already found and fixed for
logBuyPredPct thirty lines above ('era-15 Buy:2% that was really ~30%'), left
in the one line whose whole job is to be the reference every other number is
read against. Correct at era 1, wrong everywhere after - including the '62%
zero-skill' figure in the 2026-08-16 notes. Now per-era, and always-short is
finally readable.
2. THE DEPLOY GATE STATES ITS OWN BAR. 'edge -1pp' era after era cannot separate
'short by a hair' from 'short by an amount no strategy could cover'. The era
line now prints the required win rate, the SE, the effective n and the
lifespan it was deflated by; above 100% it says UNREACHABLE. At 4,738 OOS bars
and L=75.6 there are ~63 independent observations, putting the bar near 66% at
typical coverage.
3. LIFESPAN MEASURED PER RUNG. The first-passage cache already stores touch ages
at every ladder level, so each candidate geometry's resolution time is
readable without training on it - L-vs-width becomes a measurement across the
whole ladder in ONE run rather than a second chart. Each rung reports L,
n_eff, min provable edge and min provable EV.
4. SCALE OBJECTIVE IS PHASE-AWARE, defaulting to MEASURE. Width and detectability
are opposed: labels overlap by L, L grows like m*k = width^2 at fixed ratio,
so min provable EV ~ width^2 while the cost saving from width is only linear.
Doubling width quadruples the smallest EV you can prove. DEPLOY (widest that
clears reachability) is right once an edge is known; MEASURE (narrowest that
keeps round-trip spread under BARRIER_MAX_COST_FRACTION_PCT) is right while it
still has to be shown. The direction does not depend on the exponent, and
item 3 makes the exponent checkable.
Fixed in review: m_lastRungLifespan is cleared on every LadderWinShare entry or a
rejected rung reports the previous rung's lifespan as its own; per-rung
detectability is labelled IS-sample based (the deriver may not see the holdout),
so absolute figures are optimistic while the ranking is unaffected.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 10:37:05 -04:00
|
|
|
//--- LIFESPAN AT THIS RUNG, harvested in the same pass. The first-passage cache already stores the
|
|
|
|
|
//--- touch AGE at every ladder level, so the time this candidate geometry would take to resolve is
|
|
|
|
|
//--- readable for a rung the run is not training on - which turns "narrow barriers give more
|
|
|
|
|
//--- independent samples" from an argument into a measurement, across the whole ladder, in one run.
|
|
|
|
|
//--- Unresolved bars contribute the horizon: that is when their label becomes knowable.
|
|
|
|
|
double lifeSum = 0.0;
|
|
|
|
|
int lifeN = 0;
|
|
|
|
|
int horizon = MathMax(m_barrierHorizonBars, 1);
|
fix(geometry): the reachability floor measured the WRONG WINDOW - my bug from bc57aca, and it cost real width
RECONCILED: the derivation reported "target reached on 17.7% of bars" while the
label cache reported Buy on 35.9%. Nothing was broken. They measure different
windows, and both are correct:
EXCURSION window ~12 bars (the SWING MEDIAN) - what m_excUpCache accumulates
over. Deliberately short: sizing a barrier off travel
measured over a horizon that itself scales with the barrier
is circular, and it ran away to 14-31*ATR on EURUSD/USDCAD
in 2026-08-07. That guard is correct and stays.
BARRIER horizon 64 bars - what the LABEL walk and the first-passage ladder
run over, and how long the EA actually holds the trade.
So `up >= target` is a 12-bar question and `label == Buy` is a 64-bar one, and
the second can freely exceed the first. TripleBarrierLabel gates the excursion
accumulation on `idx - t <= excWindow` while the barrier walk and the ladder run
the full horizon - the split is explicit and intentional.
THE BUG IS MINE. bc57aca's scale ladder tested reachability with `up[i] >= tp`,
i.e. it asked the 12-bar question about a 64-bar trade. That understates
reachability by ~2x, which is why EVERY wide rung was rejected and the geometry
fell back to the tightest rung at 1.61/3.21. The data supported considerably
wider; the test was just asking the wrong question.
FIX: LadderWinShare() reads the answer off the first-passage ladder - target
touched strictly before the stop, over the full horizon, tie to the stop. That
is the identical question the label walk asks, so the ladder share and the Buy
rate should now agree to within rung discretisation. Both legs snap to the
SMALLEST rung at or above the requested multiple (harder target, harder stop) so
the floor stays conservative.
Expect the scale ladder to select a WIDER rung on the next relabel. On this
data the excursion test read 17.7% at q50 where the true full-horizon share is
35.9%, so rungs that scored 8.1% and 2.8% were likely well above the floor.
ALSO:
- Window reconciliation now PRINTED every derivation: excursion travel share,
ladder win share, and the label cache's Buy share side by side, with the
ladder-vs-label gap flagged if it exceeds rung discretisation. Those two must
agree; if they ever stop agreeing, one of them is wrong and the line says so.
- Renamed tpReach/slReach -> tpTravel/slTravel and relabelled the log line. They
describe the EXCURSION window and are near-tautological there (a q50 stop is
exceeded by ~50% of bars); calling them "reached within the horizon" is what
made the two quantities look like one.
- BARRIER_MIN_TP_REACH_PCT is now BARRIER_MIN_REACH_FRACTION_OF_BE (0.60) x
break-even instead of a hardcoded 20.0. Break-even for 1:RR is 100/(1+RR), so
the absolute floor silently tightened as RR rose - 0.60x at RR=2 but 0.80x at
RR=3, penalising the user for asking for a bigger target. Evaluates to exactly
20.0% at the shipped RR=2, so this is a no-op today and correct if the knob
moves.
NOT COMPILED - user compiles.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 00:51:08 -04:00
|
|
|
for(int i = 0; i < n; i++)
|
|
|
|
|
{
|
|
|
|
|
int b = idxList[i] * BARRIER_LADDER_COUNT;
|
|
|
|
|
if(b + BARRIER_LADDER_COUNT > ArraySize(m_ladderUpAt) ||
|
|
|
|
|
b + BARRIER_LADDER_COUNT > ArraySize(m_ladderDownAt))
|
|
|
|
|
continue;
|
|
|
|
|
//--- 0 means "never touched inside the horizon"; a SMALLER age is the EARLIER touch. Tie goes to
|
|
|
|
|
//--- the stop, the same pessimistic convention the label walk and the expectancy scan both use.
|
|
|
|
|
int tUp = m_ladderUpAt[b + upLvl];
|
|
|
|
|
int tDn = m_ladderDownAt[b + dnLvl];
|
|
|
|
|
if(tUp > 0 && (tDn == 0 || tUp < tDn))
|
|
|
|
|
won++;
|
|
|
|
|
seen++;
|
feat(measurement): fix zero-skill denominator, publish the deploy bar, measure lifespan per rung, add a MEASURE scale objective
The last run could not have demonstrated an edge either way, and nothing in the
log said so. Four changes so it does.
1. THE ZERO-SKILL LINE DIVIDED BY THE WRONG DENOMINATOR. m_oosWinLongTotal resets
every era; m_oosSamples only resets on a full model reset. So 'always-long %'
decayed as ~1/era: a run whose true rate is 37% printed 1.2% at era 33 and
0.0% at era 2219. This is the SAME bug already found and fixed for
logBuyPredPct thirty lines above ('era-15 Buy:2% that was really ~30%'), left
in the one line whose whole job is to be the reference every other number is
read against. Correct at era 1, wrong everywhere after - including the '62%
zero-skill' figure in the 2026-08-16 notes. Now per-era, and always-short is
finally readable.
2. THE DEPLOY GATE STATES ITS OWN BAR. 'edge -1pp' era after era cannot separate
'short by a hair' from 'short by an amount no strategy could cover'. The era
line now prints the required win rate, the SE, the effective n and the
lifespan it was deflated by; above 100% it says UNREACHABLE. At 4,738 OOS bars
and L=75.6 there are ~63 independent observations, putting the bar near 66% at
typical coverage.
3. LIFESPAN MEASURED PER RUNG. The first-passage cache already stores touch ages
at every ladder level, so each candidate geometry's resolution time is
readable without training on it - L-vs-width becomes a measurement across the
whole ladder in ONE run rather than a second chart. Each rung reports L,
n_eff, min provable edge and min provable EV.
4. SCALE OBJECTIVE IS PHASE-AWARE, defaulting to MEASURE. Width and detectability
are opposed: labels overlap by L, L grows like m*k = width^2 at fixed ratio,
so min provable EV ~ width^2 while the cost saving from width is only linear.
Doubling width quadruples the smallest EV you can prove. DEPLOY (widest that
clears reachability) is right once an edge is known; MEASURE (narrowest that
keeps round-trip spread under BARRIER_MAX_COST_FRACTION_PCT) is right while it
still has to be shown. The direction does not depend on the exponent, and
item 3 makes the exponent checkable.
Fixed in review: m_lastRungLifespan is cleared on every LadderWinShare entry or a
rejected rung reports the previous rung's lifespan as its own; per-rung
detectability is labelled IS-sample based (the deriver may not see the holdout),
so absolute figures are optimistic while the ranking is unaffected.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 10:37:05 -04:00
|
|
|
//--- Resolution age = the FIRST of the two touches; neither touching means it ran to the horizon.
|
|
|
|
|
int age = 0;
|
|
|
|
|
if(tUp > 0 && tDn > 0)
|
|
|
|
|
age = (int)MathMin(tUp, tDn);
|
|
|
|
|
else
|
|
|
|
|
if(tUp > 0)
|
|
|
|
|
age = tUp;
|
|
|
|
|
else
|
|
|
|
|
if(tDn > 0)
|
|
|
|
|
age = tDn;
|
|
|
|
|
else
|
|
|
|
|
age = horizon;
|
|
|
|
|
lifeSum += (double)age;
|
|
|
|
|
lifeN++;
|
fix(geometry): the reachability floor measured the WRONG WINDOW - my bug from bc57aca, and it cost real width
RECONCILED: the derivation reported "target reached on 17.7% of bars" while the
label cache reported Buy on 35.9%. Nothing was broken. They measure different
windows, and both are correct:
EXCURSION window ~12 bars (the SWING MEDIAN) - what m_excUpCache accumulates
over. Deliberately short: sizing a barrier off travel
measured over a horizon that itself scales with the barrier
is circular, and it ran away to 14-31*ATR on EURUSD/USDCAD
in 2026-08-07. That guard is correct and stays.
BARRIER horizon 64 bars - what the LABEL walk and the first-passage ladder
run over, and how long the EA actually holds the trade.
So `up >= target` is a 12-bar question and `label == Buy` is a 64-bar one, and
the second can freely exceed the first. TripleBarrierLabel gates the excursion
accumulation on `idx - t <= excWindow` while the barrier walk and the ladder run
the full horizon - the split is explicit and intentional.
THE BUG IS MINE. bc57aca's scale ladder tested reachability with `up[i] >= tp`,
i.e. it asked the 12-bar question about a 64-bar trade. That understates
reachability by ~2x, which is why EVERY wide rung was rejected and the geometry
fell back to the tightest rung at 1.61/3.21. The data supported considerably
wider; the test was just asking the wrong question.
FIX: LadderWinShare() reads the answer off the first-passage ladder - target
touched strictly before the stop, over the full horizon, tie to the stop. That
is the identical question the label walk asks, so the ladder share and the Buy
rate should now agree to within rung discretisation. Both legs snap to the
SMALLEST rung at or above the requested multiple (harder target, harder stop) so
the floor stays conservative.
Expect the scale ladder to select a WIDER rung on the next relabel. On this
data the excursion test read 17.7% at q50 where the true full-horizon share is
35.9%, so rungs that scored 8.1% and 2.8% were likely well above the floor.
ALSO:
- Window reconciliation now PRINTED every derivation: excursion travel share,
ladder win share, and the label cache's Buy share side by side, with the
ladder-vs-label gap flagged if it exceeds rung discretisation. Those two must
agree; if they ever stop agreeing, one of them is wrong and the line says so.
- Renamed tpReach/slReach -> tpTravel/slTravel and relabelled the log line. They
describe the EXCURSION window and are near-tautological there (a q50 stop is
exceeded by ~50% of bars); calling them "reached within the horizon" is what
made the two quantities look like one.
- BARRIER_MIN_TP_REACH_PCT is now BARRIER_MIN_REACH_FRACTION_OF_BE (0.60) x
break-even instead of a hardcoded 20.0. Break-even for 1:RR is 100/(1+RR), so
the absolute floor silently tightened as RR rose - 0.60x at RR=2 but 0.80x at
RR=3, penalising the user for asking for a bigger target. Evaluates to exactly
20.0% at the shipped RR=2, so this is a no-op today and correct if the knob
moves.
NOT COMPILED - user compiles.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 00:51:08 -04:00
|
|
|
}
|
feat(measurement): fix zero-skill denominator, publish the deploy bar, measure lifespan per rung, add a MEASURE scale objective
The last run could not have demonstrated an edge either way, and nothing in the
log said so. Four changes so it does.
1. THE ZERO-SKILL LINE DIVIDED BY THE WRONG DENOMINATOR. m_oosWinLongTotal resets
every era; m_oosSamples only resets on a full model reset. So 'always-long %'
decayed as ~1/era: a run whose true rate is 37% printed 1.2% at era 33 and
0.0% at era 2219. This is the SAME bug already found and fixed for
logBuyPredPct thirty lines above ('era-15 Buy:2% that was really ~30%'), left
in the one line whose whole job is to be the reference every other number is
read against. Correct at era 1, wrong everywhere after - including the '62%
zero-skill' figure in the 2026-08-16 notes. Now per-era, and always-short is
finally readable.
2. THE DEPLOY GATE STATES ITS OWN BAR. 'edge -1pp' era after era cannot separate
'short by a hair' from 'short by an amount no strategy could cover'. The era
line now prints the required win rate, the SE, the effective n and the
lifespan it was deflated by; above 100% it says UNREACHABLE. At 4,738 OOS bars
and L=75.6 there are ~63 independent observations, putting the bar near 66% at
typical coverage.
3. LIFESPAN MEASURED PER RUNG. The first-passage cache already stores touch ages
at every ladder level, so each candidate geometry's resolution time is
readable without training on it - L-vs-width becomes a measurement across the
whole ladder in ONE run rather than a second chart. Each rung reports L,
n_eff, min provable edge and min provable EV.
4. SCALE OBJECTIVE IS PHASE-AWARE, defaulting to MEASURE. Width and detectability
are opposed: labels overlap by L, L grows like m*k = width^2 at fixed ratio,
so min provable EV ~ width^2 while the cost saving from width is only linear.
Doubling width quadruples the smallest EV you can prove. DEPLOY (widest that
clears reachability) is right once an edge is known; MEASURE (narrowest that
keeps round-trip spread under BARRIER_MAX_COST_FRACTION_PCT) is right while it
still has to be shown. The direction does not depend on the exponent, and
item 3 makes the exponent checkable.
Fixed in review: m_lastRungLifespan is cleared on every LadderWinShare entry or a
rejected rung reports the previous rung's lifespan as its own; per-rung
detectability is labelled IS-sample based (the deriver may not see the holdout),
so absolute figures are optimistic while the ranking is unaffected.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 10:37:05 -04:00
|
|
|
m_lastRungLifespan = (lifeN > 0) ? lifeSum / lifeN : 0.0;
|
fix(geometry): the reachability floor measured the WRONG WINDOW - my bug from bc57aca, and it cost real width
RECONCILED: the derivation reported "target reached on 17.7% of bars" while the
label cache reported Buy on 35.9%. Nothing was broken. They measure different
windows, and both are correct:
EXCURSION window ~12 bars (the SWING MEDIAN) - what m_excUpCache accumulates
over. Deliberately short: sizing a barrier off travel
measured over a horizon that itself scales with the barrier
is circular, and it ran away to 14-31*ATR on EURUSD/USDCAD
in 2026-08-07. That guard is correct and stays.
BARRIER horizon 64 bars - what the LABEL walk and the first-passage ladder
run over, and how long the EA actually holds the trade.
So `up >= target` is a 12-bar question and `label == Buy` is a 64-bar one, and
the second can freely exceed the first. TripleBarrierLabel gates the excursion
accumulation on `idx - t <= excWindow` while the barrier walk and the ladder run
the full horizon - the split is explicit and intentional.
THE BUG IS MINE. bc57aca's scale ladder tested reachability with `up[i] >= tp`,
i.e. it asked the 12-bar question about a 64-bar trade. That understates
reachability by ~2x, which is why EVERY wide rung was rejected and the geometry
fell back to the tightest rung at 1.61/3.21. The data supported considerably
wider; the test was just asking the wrong question.
FIX: LadderWinShare() reads the answer off the first-passage ladder - target
touched strictly before the stop, over the full horizon, tie to the stop. That
is the identical question the label walk asks, so the ladder share and the Buy
rate should now agree to within rung discretisation. Both legs snap to the
SMALLEST rung at or above the requested multiple (harder target, harder stop) so
the floor stays conservative.
Expect the scale ladder to select a WIDER rung on the next relabel. On this
data the excursion test read 17.7% at q50 where the true full-horizon share is
35.9%, so rungs that scored 8.1% and 2.8% were likely well above the floor.
ALSO:
- Window reconciliation now PRINTED every derivation: excursion travel share,
ladder win share, and the label cache's Buy share side by side, with the
ladder-vs-label gap flagged if it exceeds rung discretisation. Those two must
agree; if they ever stop agreeing, one of them is wrong and the line says so.
- Renamed tpReach/slReach -> tpTravel/slTravel and relabelled the log line. They
describe the EXCURSION window and are near-tautological there (a q50 stop is
exceeded by ~50% of bars); calling them "reached within the horizon" is what
made the two quantities look like one.
- BARRIER_MIN_TP_REACH_PCT is now BARRIER_MIN_REACH_FRACTION_OF_BE (0.60) x
break-even instead of a hardcoded 20.0. Break-even for 1:RR is 100/(1+RR), so
the absolute floor silently tightened as RR rose - 0.60x at RR=2 but 0.80x at
RR=3, penalising the user for asking for a bigger target. Evaluates to exactly
20.0% at the shipped RR=2, so this is a no-op today and correct if the knob
moves.
NOT COMPILED - user compiles.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 00:51:08 -04:00
|
|
|
return (seen > 0) ? 100.0 * won / seen : 0.0;
|
|
|
|
|
}
|
|
|
|
|
//+------------------------------------------------------------------+
|
feat: derive the ATR multiples from measured excursions - no hardcoded geometry
The barrier was still two constants. SL_Mode/TP_Mode left the Inputs tab in
3482b6c, but the fallback was a hardcoded 2:6 and the geometry scan only ever
chose from a hardcoded grid {2,3} x {2,3,4,6,8,10}. Picking the least-bad of
eleven guesses is not deriving anything.
WHY THE SCAN WAS THE WRONG INSTRUMENT, now measurable rather than argued. It
ranks pairings by how predictable their OUTCOME is - a question about direction.
The excursion test (2c78f3b) ran on SP500 H1 and direction is the one thing
absent: ASYMMETRY p=0.0846, against RANGE/UP/DOWN all at p=0.0050, with RANGE
scoring 0.01345 vs a 0.00343 null - 4x, where the barrier label sits at 1.01x.
Hence the scan failing its own gate on every run, and its "winner" wandering
2:8 -> 3:8 -> 2:8 -> 2:4 across four runs of the same data. Excursion SIZE is
strongly measurable, so derive the geometry from that instead.
stop = q25 of measured ADVERSE travel (ordinary noise does not reach it)
target = q50 of measured FAVOURABLE travel (reached ~half the time, by
construction, inside the horizon)
Continuous, in ATR units, superseding the enum multiples. Reachability ("target
on X% of bars, stop on Y%") and the implied break-even are printed so the choice
is auditable rather than trusted.
FIXED-POINT ITERATION, not one-shot. ComputeBarrierHorizonBars scales the
horizon with the target (first-passage time grows with the band) and the
excursions are measured OVER the horizon, so target -> horizon -> excursions ->
target is a real loop - deriving once sizes the target from travel measured
under the PREVIOUS horizon. Re-measures until the multiples move <5%, capped at
3 passes, and says so if it does not settle.
Does NOT create expectancy, and the log says as much: chance precision equals
break-even at every geometry (m/(m+k) on both sides). It buys a target the
market reaches and a stop that survives noise. Where Min_Risk_Reward_Ratio
forces a target the market rarely reaches, it WARNS rather than overriding -
the ratio is the user's risk policy, so the honest move is to state its cost.
That is the collision that once rejected 100% of setups.
Pinned in the .cfg as doubles appended AFTER this morning's two ints, so .cfg
files written earlier today still load (their length guard finds no doubles) and
a model that carries them was trained on them and never re-derives.
Also fixes a message from e5ceed6 that claimed "this model resumed from disk"
unconditionally - it printed above a "seeding era 0" line on a brand-new model,
because the branch fires whenever the cache is not built, which is equally true
before a fresh model's first prebuild. A diagnostic that misreports its own
trigger is worse than one that says nothing: it gets quoted back as evidence.
FORCES A FULL RETRAIN (labels change).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-07 12:06:25 -04:00
|
|
|
bool CExpertSignalAIBase::DeriveBarrierGeometry(void)
|
|
|
|
|
{
|
|
|
|
|
int bars = m_labelCacheBars;
|
|
|
|
|
double up[], dn[];
|
|
|
|
|
ArrayResize(up, bars);
|
|
|
|
|
ArrayResize(dn, bars);
|
|
|
|
|
int n = 0;
|
2026-08-15 19:00:40 -04:00
|
|
|
//--- CONDITIONAL source for the fractal target: quantiles of the leg-scoped MFE/MAE recorded at
|
|
|
|
|
//--- Buy/Sell-LABELED bars (see m_fracLegFav) instead of the pooled every-bar excursions. The bars
|
|
|
|
|
//--- the model trades are the bars the label marks, and their excursion distribution is provably
|
|
|
|
|
//--- different from the pooled one whenever the label carries information - sizing the barriers on
|
|
|
|
|
//--- all bars mis-sizes them for the traded ones. No circularity: the fractal label does not
|
|
|
|
|
//--- depend on SL/TP. Falls through to the pooled source (with its own log line) when too few
|
|
|
|
|
//--- labeled legs exist, so a thin chart still gets a geometry.
|
fix(geometry): the reachability floor measured the WRONG WINDOW - my bug from bc57aca, and it cost real width
RECONCILED: the derivation reported "target reached on 17.7% of bars" while the
label cache reported Buy on 35.9%. Nothing was broken. They measure different
windows, and both are correct:
EXCURSION window ~12 bars (the SWING MEDIAN) - what m_excUpCache accumulates
over. Deliberately short: sizing a barrier off travel
measured over a horizon that itself scales with the barrier
is circular, and it ran away to 14-31*ATR on EURUSD/USDCAD
in 2026-08-07. That guard is correct and stays.
BARRIER horizon 64 bars - what the LABEL walk and the first-passage ladder
run over, and how long the EA actually holds the trade.
So `up >= target` is a 12-bar question and `label == Buy` is a 64-bar one, and
the second can freely exceed the first. TripleBarrierLabel gates the excursion
accumulation on `idx - t <= excWindow` while the barrier walk and the ladder run
the full horizon - the split is explicit and intentional.
THE BUG IS MINE. bc57aca's scale ladder tested reachability with `up[i] >= tp`,
i.e. it asked the 12-bar question about a 64-bar trade. That understates
reachability by ~2x, which is why EVERY wide rung was rejected and the geometry
fell back to the tightest rung at 1.61/3.21. The data supported considerably
wider; the test was just asking the wrong question.
FIX: LadderWinShare() reads the answer off the first-passage ladder - target
touched strictly before the stop, over the full horizon, tie to the stop. That
is the identical question the label walk asks, so the ladder share and the Buy
rate should now agree to within rung discretisation. Both legs snap to the
SMALLEST rung at or above the requested multiple (harder target, harder stop) so
the floor stays conservative.
Expect the scale ladder to select a WIDER rung on the next relabel. On this
data the excursion test read 17.7% at q50 where the true full-horizon share is
35.9%, so rungs that scored 8.1% and 2.8% were likely well above the floor.
ALSO:
- Window reconciliation now PRINTED every derivation: excursion travel share,
ladder win share, and the label cache's Buy share side by side, with the
ladder-vs-label gap flagged if it exceeds rung discretisation. Those two must
agree; if they ever stop agreeing, one of them is wrong and the line says so.
- Renamed tpReach/slReach -> tpTravel/slTravel and relabelled the log line. They
describe the EXCURSION window and are near-tautological there (a q50 stop is
exceeded by ~50% of bars); calling them "reached within the horizon" is what
made the two quantities look like one.
- BARRIER_MIN_TP_REACH_PCT is now BARRIER_MIN_REACH_FRACTION_OF_BE (0.60) x
break-even instead of a hardcoded 20.0. Break-even for 1:RR is 100/(1+RR), so
the absolute floor silently tightened as RR rose - 0.60x at RR=2 but 0.80x at
RR=3, penalising the user for asking for a bigger target. Evaluates to exactly
20.0% at the shipped RR=2, so this is a no-op today and correct if the knob
moves.
NOT COMPILED - user compiles.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 00:51:08 -04:00
|
|
|
//--- Parallel to up[]/dn[]: was THIS bar labelled Buy, and which bar was it? The label feeds the
|
|
|
|
|
//--- consistency report; the index feeds the first-passage reachability test in the scale ladder.
|
|
|
|
|
bool labUp[];
|
|
|
|
|
int idxList[];
|
|
|
|
|
ArrayResize(labUp, 0);
|
|
|
|
|
ArrayResize(idxList, 0);
|
2026-08-15 19:00:40 -04:00
|
|
|
bool conditional = false;
|
|
|
|
|
if(IsFractalTarget() && m_fracLegCount >= BARRIER_DERIVE_MIN_SAMPLES)
|
feat: derive the ATR multiples from measured excursions - no hardcoded geometry
The barrier was still two constants. SL_Mode/TP_Mode left the Inputs tab in
3482b6c, but the fallback was a hardcoded 2:6 and the geometry scan only ever
chose from a hardcoded grid {2,3} x {2,3,4,6,8,10}. Picking the least-bad of
eleven guesses is not deriving anything.
WHY THE SCAN WAS THE WRONG INSTRUMENT, now measurable rather than argued. It
ranks pairings by how predictable their OUTCOME is - a question about direction.
The excursion test (2c78f3b) ran on SP500 H1 and direction is the one thing
absent: ASYMMETRY p=0.0846, against RANGE/UP/DOWN all at p=0.0050, with RANGE
scoring 0.01345 vs a 0.00343 null - 4x, where the barrier label sits at 1.01x.
Hence the scan failing its own gate on every run, and its "winner" wandering
2:8 -> 3:8 -> 2:8 -> 2:4 across four runs of the same data. Excursion SIZE is
strongly measurable, so derive the geometry from that instead.
stop = q25 of measured ADVERSE travel (ordinary noise does not reach it)
target = q50 of measured FAVOURABLE travel (reached ~half the time, by
construction, inside the horizon)
Continuous, in ATR units, superseding the enum multiples. Reachability ("target
on X% of bars, stop on Y%") and the implied break-even are printed so the choice
is auditable rather than trusted.
FIXED-POINT ITERATION, not one-shot. ComputeBarrierHorizonBars scales the
horizon with the target (first-passage time grows with the band) and the
excursions are measured OVER the horizon, so target -> horizon -> excursions ->
target is a real loop - deriving once sizes the target from travel measured
under the PREVIOUS horizon. Re-measures until the multiples move <5%, capped at
3 passes, and says so if it does not settle.
Does NOT create expectancy, and the log says as much: chance precision equals
break-even at every geometry (m/(m+k) on both sides). It buys a target the
market reaches and a stop that survives noise. Where Min_Risk_Reward_Ratio
forces a target the market rarely reaches, it WARNS rather than overriding -
the ratio is the user's risk policy, so the honest move is to state its cost.
That is the collision that once rejected 100% of setups.
Pinned in the .cfg as doubles appended AFTER this morning's two ints, so .cfg
files written earlier today still load (their length guard finds no doubles) and
a model that carries them was trained on them and never re-derives.
Also fixes a message from e5ceed6 that claimed "this model resumed from disk"
unconditionally - it printed above a "seeding era 0" line on a brand-new model,
because the branch fires whenever the cache is not built, which is equally true
before a fresh model's first prebuild. A diagnostic that misreports its own
trigger is worse than one that says nothing: it gets quoted back as evidence.
FORCES A FULL RETRAIN (labels change).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-07 12:06:25 -04:00
|
|
|
{
|
2026-08-15 19:00:40 -04:00
|
|
|
conditional = true;
|
|
|
|
|
n = m_fracLegCount;
|
|
|
|
|
ArrayResize(up, n);
|
|
|
|
|
ArrayResize(dn, n);
|
|
|
|
|
for(int i = 0; i < n; i++)
|
|
|
|
|
{
|
|
|
|
|
up[i] = m_fracLegFav[i];
|
|
|
|
|
dn[i] = m_fracLegAdv[i];
|
|
|
|
|
}
|
|
|
|
|
}
|
|
|
|
|
else
|
|
|
|
|
{
|
|
|
|
|
if(IsFractalTarget())
|
|
|
|
|
Print(ID + StringFormat(": conditional geometry NOT available - only %d labeled fractal legs "
|
|
|
|
|
"(need %d); deriving from the pooled every-bar excursions instead.",
|
|
|
|
|
m_fracLegCount, BARRIER_DERIVE_MIN_SAMPLES));
|
|
|
|
|
//--- IS region only, matching BuildMiSample: a geometry chosen with the holdout in view has used
|
|
|
|
|
//--- the holdout for selection, and it stops being a holdout.
|
|
|
|
|
int oosCutoff = (int)(MathMax(0, MathMin(100, m_oosSplitPct)) / 100.0
|
|
|
|
|
* MathMax(bars - MathMax(m_historyBars, 0), 0));
|
|
|
|
|
for(int i = MathMax(oosCutoff, 0); i < bars; i++)
|
|
|
|
|
{
|
|
|
|
|
if(i >= ArraySize(m_labelCacheHasValue) || !m_labelCacheHasValue[i])
|
|
|
|
|
continue;
|
|
|
|
|
if(i >= ArraySize(m_excUpCache))
|
|
|
|
|
continue;
|
|
|
|
|
double u = m_excUpCache[i], d = m_excDownCache[i];
|
|
|
|
|
if(!MathIsValidNumber(u) || !MathIsValidNumber(d) || (u <= 0.0 && d <= 0.0))
|
|
|
|
|
continue; // unresolvable bar - see the same guard in BuildMiSample
|
|
|
|
|
up[n] = u;
|
|
|
|
|
dn[n] = d;
|
fix(geometry): the reachability floor measured the WRONG WINDOW - my bug from bc57aca, and it cost real width
RECONCILED: the derivation reported "target reached on 17.7% of bars" while the
label cache reported Buy on 35.9%. Nothing was broken. They measure different
windows, and both are correct:
EXCURSION window ~12 bars (the SWING MEDIAN) - what m_excUpCache accumulates
over. Deliberately short: sizing a barrier off travel
measured over a horizon that itself scales with the barrier
is circular, and it ran away to 14-31*ATR on EURUSD/USDCAD
in 2026-08-07. That guard is correct and stays.
BARRIER horizon 64 bars - what the LABEL walk and the first-passage ladder
run over, and how long the EA actually holds the trade.
So `up >= target` is a 12-bar question and `label == Buy` is a 64-bar one, and
the second can freely exceed the first. TripleBarrierLabel gates the excursion
accumulation on `idx - t <= excWindow` while the barrier walk and the ladder run
the full horizon - the split is explicit and intentional.
THE BUG IS MINE. bc57aca's scale ladder tested reachability with `up[i] >= tp`,
i.e. it asked the 12-bar question about a 64-bar trade. That understates
reachability by ~2x, which is why EVERY wide rung was rejected and the geometry
fell back to the tightest rung at 1.61/3.21. The data supported considerably
wider; the test was just asking the wrong question.
FIX: LadderWinShare() reads the answer off the first-passage ladder - target
touched strictly before the stop, over the full horizon, tie to the stop. That
is the identical question the label walk asks, so the ladder share and the Buy
rate should now agree to within rung discretisation. Both legs snap to the
SMALLEST rung at or above the requested multiple (harder target, harder stop) so
the floor stays conservative.
Expect the scale ladder to select a WIDER rung on the next relabel. On this
data the excursion test read 17.7% at q50 where the true full-horizon share is
35.9%, so rungs that scored 8.1% and 2.8% were likely well above the floor.
ALSO:
- Window reconciliation now PRINTED every derivation: excursion travel share,
ladder win share, and the label cache's Buy share side by side, with the
ladder-vs-label gap flagged if it exceeds rung discretisation. Those two must
agree; if they ever stop agreeing, one of them is wrong and the line says so.
- Renamed tpReach/slReach -> tpTravel/slTravel and relabelled the log line. They
describe the EXCURSION window and are near-tautological there (a q50 stop is
exceeded by ~50% of bars); calling them "reached within the horizon" is what
made the two quantities look like one.
- BARRIER_MIN_TP_REACH_PCT is now BARRIER_MIN_REACH_FRACTION_OF_BE (0.60) x
break-even instead of a hardcoded 20.0. Break-even for 1:RR is 100/(1+RR), so
the absolute floor silently tightened as RR rose - 0.60x at RR=2 but 0.80x at
RR=3, penalising the user for asking for a bigger target. Evaluates to exactly
20.0% at the shipped RR=2, so this is a no-op today and correct if the knob
moves.
NOT COMPILED - user compiles.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 00:51:08 -04:00
|
|
|
//--- CONSISTENCY CHECK + first-passage reachability, harvested on the SAME bar in the SAME pass.
|
|
|
|
|
ArrayResize(labUp, n + 1);
|
|
|
|
|
labUp[n] = (i < ArraySize(m_labelCacheBuy) && m_labelCacheBuy[i]);
|
|
|
|
|
ArrayResize(idxList, n + 1);
|
|
|
|
|
idxList[n] = i;
|
2026-08-15 19:00:40 -04:00
|
|
|
n++;
|
|
|
|
|
}
|
feat: derive the ATR multiples from measured excursions - no hardcoded geometry
The barrier was still two constants. SL_Mode/TP_Mode left the Inputs tab in
3482b6c, but the fallback was a hardcoded 2:6 and the geometry scan only ever
chose from a hardcoded grid {2,3} x {2,3,4,6,8,10}. Picking the least-bad of
eleven guesses is not deriving anything.
WHY THE SCAN WAS THE WRONG INSTRUMENT, now measurable rather than argued. It
ranks pairings by how predictable their OUTCOME is - a question about direction.
The excursion test (2c78f3b) ran on SP500 H1 and direction is the one thing
absent: ASYMMETRY p=0.0846, against RANGE/UP/DOWN all at p=0.0050, with RANGE
scoring 0.01345 vs a 0.00343 null - 4x, where the barrier label sits at 1.01x.
Hence the scan failing its own gate on every run, and its "winner" wandering
2:8 -> 3:8 -> 2:8 -> 2:4 across four runs of the same data. Excursion SIZE is
strongly measurable, so derive the geometry from that instead.
stop = q25 of measured ADVERSE travel (ordinary noise does not reach it)
target = q50 of measured FAVOURABLE travel (reached ~half the time, by
construction, inside the horizon)
Continuous, in ATR units, superseding the enum multiples. Reachability ("target
on X% of bars, stop on Y%") and the implied break-even are printed so the choice
is auditable rather than trusted.
FIXED-POINT ITERATION, not one-shot. ComputeBarrierHorizonBars scales the
horizon with the target (first-passage time grows with the band) and the
excursions are measured OVER the horizon, so target -> horizon -> excursions ->
target is a real loop - deriving once sizes the target from travel measured
under the PREVIOUS horizon. Re-measures until the multiples move <5%, capped at
3 passes, and says so if it does not settle.
Does NOT create expectancy, and the log says as much: chance precision equals
break-even at every geometry (m/(m+k) on both sides). It buys a target the
market reaches and a stop that survives noise. Where Min_Risk_Reward_Ratio
forces a target the market rarely reaches, it WARNS rather than overriding -
the ratio is the user's risk policy, so the honest move is to state its cost.
That is the collision that once rejected 100% of setups.
Pinned in the .cfg as doubles appended AFTER this morning's two ints, so .cfg
files written earlier today still load (their length guard finds no doubles) and
a model that carries them was trained on them and never re-derives.
Also fixes a message from e5ceed6 that claimed "this model resumed from disk"
unconditionally - it printed above a "seeding era 0" line on a brand-new model,
because the branch fires whenever the cache is not built, which is equally true
before a fresh model's first prebuild. A diagnostic that misreports its own
trigger is worse than one that says nothing: it gets quoted back as evidence.
FORCES A FULL RETRAIN (labels change).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-07 12:06:25 -04:00
|
|
|
}
|
|
|
|
|
if(n < BARRIER_DERIVE_MIN_SAMPLES)
|
|
|
|
|
{
|
|
|
|
|
Print(ID + StringFormat(": barrier geometry NOT derived - only %d usable excursion samples "
|
|
|
|
|
"(need %d). Falling back to the configured %d:%d.", n,
|
|
|
|
|
BARRIER_DERIVE_MIN_SAMPLES, m_sl_mode, m_tp_mode));
|
|
|
|
|
return false;
|
|
|
|
|
}
|
|
|
|
|
ArrayResize(up, n);
|
|
|
|
|
ArrayResize(dn, n);
|
fix(geometry): the reachability floor measured the WRONG WINDOW - my bug from bc57aca, and it cost real width
RECONCILED: the derivation reported "target reached on 17.7% of bars" while the
label cache reported Buy on 35.9%. Nothing was broken. They measure different
windows, and both are correct:
EXCURSION window ~12 bars (the SWING MEDIAN) - what m_excUpCache accumulates
over. Deliberately short: sizing a barrier off travel
measured over a horizon that itself scales with the barrier
is circular, and it ran away to 14-31*ATR on EURUSD/USDCAD
in 2026-08-07. That guard is correct and stays.
BARRIER horizon 64 bars - what the LABEL walk and the first-passage ladder
run over, and how long the EA actually holds the trade.
So `up >= target` is a 12-bar question and `label == Buy` is a 64-bar one, and
the second can freely exceed the first. TripleBarrierLabel gates the excursion
accumulation on `idx - t <= excWindow` while the barrier walk and the ladder run
the full horizon - the split is explicit and intentional.
THE BUG IS MINE. bc57aca's scale ladder tested reachability with `up[i] >= tp`,
i.e. it asked the 12-bar question about a 64-bar trade. That understates
reachability by ~2x, which is why EVERY wide rung was rejected and the geometry
fell back to the tightest rung at 1.61/3.21. The data supported considerably
wider; the test was just asking the wrong question.
FIX: LadderWinShare() reads the answer off the first-passage ladder - target
touched strictly before the stop, over the full horizon, tie to the stop. That
is the identical question the label walk asks, so the ladder share and the Buy
rate should now agree to within rung discretisation. Both legs snap to the
SMALLEST rung at or above the requested multiple (harder target, harder stop) so
the floor stays conservative.
Expect the scale ladder to select a WIDER rung on the next relabel. On this
data the excursion test read 17.7% at q50 where the true full-horizon share is
35.9%, so rungs that scored 8.1% and 2.8% were likely well above the floor.
ALSO:
- Window reconciliation now PRINTED every derivation: excursion travel share,
ladder win share, and the label cache's Buy share side by side, with the
ladder-vs-label gap flagged if it exceeds rung discretisation. Those two must
agree; if they ever stop agreeing, one of them is wrong and the line says so.
- Renamed tpReach/slReach -> tpTravel/slTravel and relabelled the log line. They
describe the EXCURSION window and are near-tautological there (a q50 stop is
exceeded by ~50% of bars); calling them "reached within the horizon" is what
made the two quantities look like one.
- BARRIER_MIN_TP_REACH_PCT is now BARRIER_MIN_REACH_FRACTION_OF_BE (0.60) x
break-even instead of a hardcoded 20.0. Break-even for 1:RR is 100/(1+RR), so
the absolute floor silently tightened as RR rose - 0.60x at RR=2 but 0.80x at
RR=3, penalising the user for asking for a bigger target. Evaluates to exactly
20.0% at the shipped RR=2, so this is a no-op today and correct if the knob
moves.
NOT COMPILED - user compiles.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 00:51:08 -04:00
|
|
|
//--- Unsorted copy: ArraySort below destroys the index correspondence with labUp[], and the
|
|
|
|
|
//--- consistency report needs to pair each bar's excursion with its own label.
|
|
|
|
|
double upUnsorted[];
|
|
|
|
|
ArrayResize(upUnsorted, n);
|
|
|
|
|
ArrayCopy(upUnsorted, up, 0, 0, n);
|
feat: derive the ATR multiples from measured excursions - no hardcoded geometry
The barrier was still two constants. SL_Mode/TP_Mode left the Inputs tab in
3482b6c, but the fallback was a hardcoded 2:6 and the geometry scan only ever
chose from a hardcoded grid {2,3} x {2,3,4,6,8,10}. Picking the least-bad of
eleven guesses is not deriving anything.
WHY THE SCAN WAS THE WRONG INSTRUMENT, now measurable rather than argued. It
ranks pairings by how predictable their OUTCOME is - a question about direction.
The excursion test (2c78f3b) ran on SP500 H1 and direction is the one thing
absent: ASYMMETRY p=0.0846, against RANGE/UP/DOWN all at p=0.0050, with RANGE
scoring 0.01345 vs a 0.00343 null - 4x, where the barrier label sits at 1.01x.
Hence the scan failing its own gate on every run, and its "winner" wandering
2:8 -> 3:8 -> 2:8 -> 2:4 across four runs of the same data. Excursion SIZE is
strongly measurable, so derive the geometry from that instead.
stop = q25 of measured ADVERSE travel (ordinary noise does not reach it)
target = q50 of measured FAVOURABLE travel (reached ~half the time, by
construction, inside the horizon)
Continuous, in ATR units, superseding the enum multiples. Reachability ("target
on X% of bars, stop on Y%") and the implied break-even are printed so the choice
is auditable rather than trusted.
FIXED-POINT ITERATION, not one-shot. ComputeBarrierHorizonBars scales the
horizon with the target (first-passage time grows with the band) and the
excursions are measured OVER the horizon, so target -> horizon -> excursions ->
target is a real loop - deriving once sizes the target from travel measured
under the PREVIOUS horizon. Re-measures until the multiples move <5%, capped at
3 passes, and says so if it does not settle.
Does NOT create expectancy, and the log says as much: chance precision equals
break-even at every geometry (m/(m+k) on both sides). It buys a target the
market reaches and a stop that survives noise. Where Min_Risk_Reward_Ratio
forces a target the market rarely reaches, it WARNS rather than overriding -
the ratio is the user's risk policy, so the honest move is to state its cost.
That is the collision that once rejected 100% of setups.
Pinned in the .cfg as doubles appended AFTER this morning's two ints, so .cfg
files written earlier today still load (their length guard finds no doubles) and
a model that carries them was trained on them and never re-derives.
Also fixes a message from e5ceed6 that claimed "this model resumed from disk"
unconditionally - it printed above a "seeding era 0" line on a brand-new model,
because the branch fires whenever the cache is not built, which is equally true
before a fresh model's first prebuild. A diagnostic that misreports its own
trigger is worse than one that says nothing: it gets quoted back as evidence.
FORCES A FULL RETRAIN (labels change).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-07 12:06:25 -04:00
|
|
|
ArraySort(up);
|
|
|
|
|
ArraySort(dn);
|
|
|
|
|
//--- STOP from the ADVERSE distribution, TARGET from the FAVOURABLE one - each leg sized by the thing
|
fix: normalise the asymmetry target - the raw one is confounded by volatility
Three symbols ran the excursion test. RANGE/UP/DOWN cleared on all three;
raw ASYMMETRY cleared on EURUSD and USDCAD at p=0.0050 and not on SP500
(p=0.1045). That looked like the first directional signal this project has
found. It probably is not, and the test as built could not tell.
(up-dn) IS NOT SCALE-FREE. If sigma is predictable - and RANGE clears at ~4x its
null on every instrument - and the directional part is symmetric noise eps, then
up-dn ~ sigma*eps, so a large sigma pushes the value into BOTH outer terciles. A
pure volatility predictor scores positive MI against a 3-bin (up-dn) while
carrying no directional information at all. Crucially that confound REPLICATES,
so reproducing on two instruments is not evidence against it - and the effect
sizes fit it: asymmetry runs 1.3-1.6x its null where RANGE runs ~4x, and carries
~0.1% of the target's entropy against RANGE's ~0.9%. That is the shape of a
leaked fraction of the volatility signal, not an independent one.
So add (up-dn)/(up+dn): bounded in [-1,+1], volatility divided out, and the only
target a directional claim may rest on. The verdict now separates the cases and
NAMES the confound when raw clears while normalised does not, instead of
reporting the raw line as a finding.
Two bugs of mine in the same block, both caught by output rather than review:
- The derived-geometry line had a MISORDERED argument list: it printed
"stop 25.00*ATR (q3 of adverse travel)" - the quantile percentage as the
multiple and the multiple as the quantile. Real values were 2.61 stop /
8.03 target. A 25*ATR stop is absurd on its face, which is why it was seen.
- THE STOP QUANTILE WAS BACKWARDS, and this one changes labels. It was 0.25
"so ordinary noise does not reach it", but q25 means 75% of bars EXCEED the
stop - hit three times in four. The printed reachability said exactly that
("stop on 75.0% of bars"). Now 0.75. A quantile is a threshold, not a rate.
This is the entire reason reachability is measured and printed rather than
assumed.
Also raises BARRIER_DERIVE_MAX_PASSES 3 -> 5: SP500 did not settle in 3 (stop
still moving ~14% per pass) while EURUSD and USDCAD converged on pass 2. And
bounds both quantile indices with MathMin(..., n-1) so q=1.0 cannot run off the
end of the sorted array.
The geometry from the previous run is NOT usable and the asymmetry result is
unresolved, not established. Both are decided by the next run.
FORCES A FULL RETRAIN (the stop quantile changes every label).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-07 13:04:13 -04:00
|
|
|
//--- it actually has to survive or reach. The stop sits at a HIGH quantile of MAE so only the minority
|
|
|
|
|
//--- of bars whose adverse travel exceeds it ever reach it; the target at the MEDIAN of MFE so it is
|
|
|
|
|
//--- reached about half the time within the horizon. See BARRIER_SL_QUANTILE for why that quantile is
|
|
|
|
|
//--- 0.75 and not 0.25 - the first version had it backwards and the printed reachability caught it.
|
fix(geometry): the target was small BY CONSTRUCTION - ratio is now policy, scale is measured, ladder ceiling removed
The derivation read the stop from q75 of ADVERSE travel and the target from q50
of FAVOURABLE travel. Over one horizon those distributions are broadly the same
shape, so q75 > q50 MECHANICALLY - the target came out smaller than the stop no
matter what the market did. SP500 H4 shipped stop 3.07 / target 1.70: a 0.55:1
payoff needing 64.3%. That was never a measurement, it was two mismatched
constants.
The reachability line printed beside it - "target on 50.0% of bars, stop on
25.0%" - is exactly 1-q50 and 1-q75. Tautological. It cannot disconfirm
anything, and it read as validation.
WIDTH AND RATIO ARE INDEPENDENT AND ONLY ONE PAYS. EV = edge x width;
ratio is EV-neutral (a driftless walk reaches +m before -k with probability
k/(k+m), which IS break-even). Width is what buys cost efficiency: the spread
is a fixed 0.047*ATR here, so the shipped 4.77*ATR width paid it 21 times per
unit of travel. So:
RATIO = policy. BARRIER_TARGET_RR = 2.0 (user's 1:2). Break-even 33.3%.
SCALE = measured. The stop quantile is chosen from a ladder, WIDEST FIRST,
taking the first rung whose implied 2x target is still reached often
enough to be a trainable class.
That last clause is the difference from the min-reward:risk raise removed in
2026-08-09, which forced target = 2 x stop with NO reachability test, landed on
6.66*ATR reachable on 3.3% of bars, and trained the model to predict something
that essentially never happened. Same ratio; the scale now retreats until the
data says the target is attainable. Every rung is logged.
LADDER CEILING REMOVED. BARRIER_LADDER stopped at 5.00 and the expectancy scan's
"best resolvable pair on width alone" came back as stop 5.05 / target 4.95 - it
pinned to the top rung. A recommendation landing exactly on the edge of its own
search space is a boundary, not a finding: it cannot tell "5 ATR is optimal"
from "5 ATR is all we allowed". Extended to 20*ATR (8 -> 14 rungs). Nothing else
needs editing - every consumer is parameterised by BARRIER_LADDER_COUNT - and
the horizon constraints (decided >= 60%, reachability floor) now bind instead of
a constant.
THE SCAN COULD NOT SEE THE SHIPPED GEOMETRY. ReportBarrierGeometryScan looked
the configured pair up in its integer grid, and DeriveBarrierGeometry produces
CONTINUOUS multiples (3.07/1.70) that can never equal a grid point - so
cfgExcess stayed at its -1.0 sentinel and the report printed "configured 3:2
scores -1.00000", which reads as a catastrophic score and actually means "never
evaluated". Worse, the grid skipped target<stop entirely because it "inverts the
trade's whole premise" - while the derivation was shipping exactly that. The
incumbent is now always scored as a peer (never crowned; it is already in force
and is not an enum pairing the scan could adopt).
BREAK-EVEN NOW INCLUDES THE SPREAD. Every report quoted the frictionless
SL/(SL+TP). On SP500 H4 that read 64.3% while the MEASURED zero-skill rate was
62.1% - a 2.2pp gap that IS the cost, and that made every model look 2.2pp
better than it was. CostAdjustedBreakEvenPct() prices a win at (TP - spread) and
a loss at (SL + spread), matching the expectancy scan's convention exactly so
the two reports cannot disagree.
It also feeds FitDirConfThreshold, which is the correctness half: the operating
point subtracts break-even from precision, so the frictionless figure made every
candidate threshold look better by the width of the spread - 2.2pp against a
measured edge of 2.3pp, i.e. very nearly all of it.
Era line now carries both: "break-even 64.3% frictionless, 66.6% AFTER SPREAD".
Forces a full relabel and retrain. Requested.
NOT COMPILED - user compiles.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 00:34:32 -04:00
|
|
|
//--- SCALE FROM THE DATA, RATIO FROM POLICY - see BARRIER_TARGET_RR. Walk the quantile ladder widest
|
|
|
|
|
//--- first and take the first rung whose implied target is still reached often enough to be a
|
|
|
|
|
//--- trainable class. Width is what pays (EV = edge x width) so wider is strictly better on cost;
|
|
|
|
|
//--- reachability is the only thing that stops it running away, and it is measured here rather than
|
|
|
|
|
//--- assumed. Every rung is reported so the choice is auditable.
|
fix(labels): overlapping-label sample correction + horizon cap on the scale ladder
Three defects, all surfaced by the 2026-08-17 SP500 H4 run that shipped
stop 4.86 / target 9.71 (width 14.57*ATR, horizon 384).
1. EVERY STANDARD ERROR ASSUMED INDEPENDENT SAMPLES. Triple-barrier labels
started one per bar overlap by the label's lifespan, so n calls are worth
~n/L independent observations (Lopez de Prado, AFML ch. 4 - sample
uniqueness). All three sqrt(p(1-p)/n) sites divided by the RAW count.
The tell: the operating point's null-of-the-maximum gate is family-wise and
should fire on ~5% of eras under the null. Measured fire rates - PAI 47/73
(64%), ConvLSTM 9/24, LSTM 8/21 (38%), CONV 4/62 (6%). CONV, the only model
whose margin distribution admits few bins, sat on the null; the rest cleared
a bar that was too low by ~sqrt(L). PAI's deployed threshold consequently
alternated between the ENDS of its own range era to era (0.10 -> 0.88 ->
0.86 -> 0.66; coverage 16% <-> 73%).
TripleBarrierLabel now records when each label became KNOWABLE - the first
winning touch, or both stops, or the timeout - and the prebuild accumulates
the mean. EffectiveSampleSize() feeds the operating point, the member deploy
gate and the ensemble vote gate. Conservative by construction (n/L is an
upper bound on the damage); gates get harder, never easier.
2. THE SCALE LADDER RAN AWAY, again. Horizon scales as swingMedian*sl*tp, and
since 4d8cb08 reachability is measured OVER that horizon - so a wider rung
buys itself the time that makes it look reachable. Same target -> horizon ->
reach -> target loop the excursion window is kept short to avoid; fixing the
window confusion reopened it through the other door. It walked 128 -> 256 ->
384 bars and stopped at q90, the widest rung there is, with every rung
reading 39-48% against a 20% floor. A floor nothing fails selects nothing.
Rungs whose required horizon exceeds BARRIER_HORIZON_MAX are now rejected -
the same rule ReportGeometryExpectancyScan already applied. It was printing
the shipped pair as CLAMPED and disqualified ('h384!') two lines under the
deriver that chose it: two subsystems, one geometry, opposite verdicts.
3. THE RUNG SNAP DESTROYED THE RATIO IT WAS COMPARING. Both legs snapped
independently to the coarse first-passage grid, re-rating each candidate:
q90 4.86/9.71 -> 5.00/10.00 (2.00), q85 4.07/8.14 -> 5.00/10.00 (IDENTICAL
measurement), q75 3.07/6.13 -> 4.00/6.50 (1.63 - a nearer target). So the
ladder compared win shares taken at ratios from 1.63 to 2.17 and read the
differences as scale. It is why the reach column came out non-monotone in
width (q75 48.5% above q90 42.9%). The stop now snaps to its nearest rung in
log space and the target follows the ratio off it; the pair actually measured
is returned and logged, so a collision reads as a collision.
Also: LadderWinShare guarded against the conditional (fractal) geometry path,
which fills n from m_fracLegCount while leaving idxList empty - a latent
out-of-bounds on a currently-dead path.
New log lines: mean label lifespan and effective n on the label-cache line, the
required-vs-available horizon per rung, and the grid pair the reconciliation
actually measured (its tolerance now scales with the grid skew instead of a flat
5pp).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 01:12:05 -04:00
|
|
|
//--- THE HORIZON IS A HARD CONSTRAINT ON WIDTH, added 2026-08-17 after this ladder ran away.
|
|
|
|
|
//--- First-passage time for the band [-m, +k] grows like m*k, so widening the barrier lengthens the
|
|
|
|
|
//--- horizon QUADRATICALLY - and because chosenReach is measured over that horizon, a wider rung buys
|
|
|
|
|
//--- itself the very time that makes it look reachable. target -> horizon -> reach -> target is a loop,
|
|
|
|
|
//--- and it is the SAME loop the excursion window was deliberately kept short to avoid (see
|
|
|
|
|
//--- ComputeBarrierHorizonBars, and the 2026-08-07 EURUSD/USDCAD runaway to 14-31*ATR that "converged"
|
|
|
|
|
//--- only because the horizon ladder caps at 384). Measuring reachability over the full horizon - the
|
|
|
|
|
//--- correct fix for the excursion-window confusion - reopened it through the other door: on SP500 H4
|
|
|
|
|
//--- the geometry walked 128 -> 256 -> 384 bars over three derive passes and stopped at q90, the
|
|
|
|
|
//--- WIDEST rung there is, with every rung reading 39-48% against a 20% floor. A floor that nothing
|
|
|
|
|
//--- fails is not selecting anything; the ladder had degenerated to "take the widest".
|
|
|
|
|
//---
|
|
|
|
|
//--- What actually stops it is the constraint the loop cannot buy its way out of: a rung whose
|
|
|
|
|
//--- required horizon exceeds BARRIER_HORIZON_MAX gets a CLAMPED label - "target before stop" quietly
|
|
|
|
|
//--- becomes "...or 384 bars, whichever comes first" - while the deployed EA holds to SL or TP with no
|
|
|
|
|
//--- bar limit. That is a train/deploy mismatch in the target itself. ReportGeometryExpectancyScan has
|
|
|
|
|
//--- always disqualified those candidates ('!' in its output); the deriver did not, and on 2026-08-17
|
|
|
|
|
//--- shipped 4.86/9.71 - needing 613 bars - while the scan printed that same pair as CLAMPED two lines
|
|
|
|
|
//--- later. Two subsystems, one geometry, opposite verdicts. Now they apply the same rule.
|
fix(geometry): the target was small BY CONSTRUCTION - ratio is now policy, scale is measured, ladder ceiling removed
The derivation read the stop from q75 of ADVERSE travel and the target from q50
of FAVOURABLE travel. Over one horizon those distributions are broadly the same
shape, so q75 > q50 MECHANICALLY - the target came out smaller than the stop no
matter what the market did. SP500 H4 shipped stop 3.07 / target 1.70: a 0.55:1
payoff needing 64.3%. That was never a measurement, it was two mismatched
constants.
The reachability line printed beside it - "target on 50.0% of bars, stop on
25.0%" - is exactly 1-q50 and 1-q75. Tautological. It cannot disconfirm
anything, and it read as validation.
WIDTH AND RATIO ARE INDEPENDENT AND ONLY ONE PAYS. EV = edge x width;
ratio is EV-neutral (a driftless walk reaches +m before -k with probability
k/(k+m), which IS break-even). Width is what buys cost efficiency: the spread
is a fixed 0.047*ATR here, so the shipped 4.77*ATR width paid it 21 times per
unit of travel. So:
RATIO = policy. BARRIER_TARGET_RR = 2.0 (user's 1:2). Break-even 33.3%.
SCALE = measured. The stop quantile is chosen from a ladder, WIDEST FIRST,
taking the first rung whose implied 2x target is still reached often
enough to be a trainable class.
That last clause is the difference from the min-reward:risk raise removed in
2026-08-09, which forced target = 2 x stop with NO reachability test, landed on
6.66*ATR reachable on 3.3% of bars, and trained the model to predict something
that essentially never happened. Same ratio; the scale now retreats until the
data says the target is attainable. Every rung is logged.
LADDER CEILING REMOVED. BARRIER_LADDER stopped at 5.00 and the expectancy scan's
"best resolvable pair on width alone" came back as stop 5.05 / target 4.95 - it
pinned to the top rung. A recommendation landing exactly on the edge of its own
search space is a boundary, not a finding: it cannot tell "5 ATR is optimal"
from "5 ATR is all we allowed". Extended to 20*ATR (8 -> 14 rungs). Nothing else
needs editing - every consumer is parameterised by BARRIER_LADDER_COUNT - and
the horizon constraints (decided >= 60%, reachability floor) now bind instead of
a constant.
THE SCAN COULD NOT SEE THE SHIPPED GEOMETRY. ReportBarrierGeometryScan looked
the configured pair up in its integer grid, and DeriveBarrierGeometry produces
CONTINUOUS multiples (3.07/1.70) that can never equal a grid point - so
cfgExcess stayed at its -1.0 sentinel and the report printed "configured 3:2
scores -1.00000", which reads as a catastrophic score and actually means "never
evaluated". Worse, the grid skipped target<stop entirely because it "inverts the
trade's whole premise" - while the derivation was shipping exactly that. The
incumbent is now always scored as a peer (never crowned; it is already in force
and is not an enum pairing the scan could adopt).
BREAK-EVEN NOW INCLUDES THE SPREAD. Every report quoted the frictionless
SL/(SL+TP). On SP500 H4 that read 64.3% while the MEASURED zero-skill rate was
62.1% - a 2.2pp gap that IS the cost, and that made every model look 2.2pp
better than it was. CostAdjustedBreakEvenPct() prices a win at (TP - spread) and
a loss at (SL + spread), matching the expectancy scan's convention exactly so
the two reports cannot disagree.
It also feeds FitDirConfThreshold, which is the correctness half: the operating
point subtracts break-even from precision, so the frictionless figure made every
candidate threshold look better by the width of the spread - 2.2pp against a
measured edge of 2.3pp, i.e. very nearly all of it.
Era line now carries both: "break-even 64.3% frictionless, 66.6% AFTER SPREAD".
Forces a full relabel and retrain. Requested.
NOT COMPILED - user compiles.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 00:34:32 -04:00
|
|
|
double slRaw = 0.0, tpRaw = 0.0;
|
|
|
|
|
double chosenQ = 0.0, chosenReach = 0.0;
|
|
|
|
|
string rungRows = "";
|
|
|
|
|
for(int r = 0; r < BARRIER_SL_QUANTILE_COUNT; r++)
|
|
|
|
|
{
|
|
|
|
|
double q = BARRIER_SL_QUANTILE_LADDER[r];
|
|
|
|
|
double sl = dn[(int)MathMin(q * n, n - 1)];
|
|
|
|
|
if(sl < MIN_SL_ATR_MULTIPLIER)
|
|
|
|
|
sl = MIN_SL_ATR_MULTIPLIER;
|
|
|
|
|
double tp = BARRIER_TARGET_RR * sl;
|
fix(labels): overlapping-label sample correction + horizon cap on the scale ladder
Three defects, all surfaced by the 2026-08-17 SP500 H4 run that shipped
stop 4.86 / target 9.71 (width 14.57*ATR, horizon 384).
1. EVERY STANDARD ERROR ASSUMED INDEPENDENT SAMPLES. Triple-barrier labels
started one per bar overlap by the label's lifespan, so n calls are worth
~n/L independent observations (Lopez de Prado, AFML ch. 4 - sample
uniqueness). All three sqrt(p(1-p)/n) sites divided by the RAW count.
The tell: the operating point's null-of-the-maximum gate is family-wise and
should fire on ~5% of eras under the null. Measured fire rates - PAI 47/73
(64%), ConvLSTM 9/24, LSTM 8/21 (38%), CONV 4/62 (6%). CONV, the only model
whose margin distribution admits few bins, sat on the null; the rest cleared
a bar that was too low by ~sqrt(L). PAI's deployed threshold consequently
alternated between the ENDS of its own range era to era (0.10 -> 0.88 ->
0.86 -> 0.66; coverage 16% <-> 73%).
TripleBarrierLabel now records when each label became KNOWABLE - the first
winning touch, or both stops, or the timeout - and the prebuild accumulates
the mean. EffectiveSampleSize() feeds the operating point, the member deploy
gate and the ensemble vote gate. Conservative by construction (n/L is an
upper bound on the damage); gates get harder, never easier.
2. THE SCALE LADDER RAN AWAY, again. Horizon scales as swingMedian*sl*tp, and
since 4d8cb08 reachability is measured OVER that horizon - so a wider rung
buys itself the time that makes it look reachable. Same target -> horizon ->
reach -> target loop the excursion window is kept short to avoid; fixing the
window confusion reopened it through the other door. It walked 128 -> 256 ->
384 bars and stopped at q90, the widest rung there is, with every rung
reading 39-48% against a 20% floor. A floor nothing fails selects nothing.
Rungs whose required horizon exceeds BARRIER_HORIZON_MAX are now rejected -
the same rule ReportGeometryExpectancyScan already applied. It was printing
the shipped pair as CLAMPED and disqualified ('h384!') two lines under the
deriver that chose it: two subsystems, one geometry, opposite verdicts.
3. THE RUNG SNAP DESTROYED THE RATIO IT WAS COMPARING. Both legs snapped
independently to the coarse first-passage grid, re-rating each candidate:
q90 4.86/9.71 -> 5.00/10.00 (2.00), q85 4.07/8.14 -> 5.00/10.00 (IDENTICAL
measurement), q75 3.07/6.13 -> 4.00/6.50 (1.63 - a nearer target). So the
ladder compared win shares taken at ratios from 1.63 to 2.17 and read the
differences as scale. It is why the reach column came out non-monotone in
width (q75 48.5% above q90 42.9%). The stop now snaps to its nearest rung in
log space and the target follows the ratio off it; the pair actually measured
is returned and logged, so a collision reads as a collision.
Also: LadderWinShare guarded against the conditional (fractal) geometry path,
which fills n from m_fracLegCount while leaving idxList empty - a latent
out-of-bounds on a currently-dead path.
New log lines: mean label lifespan and effective n on the label-cache line, the
required-vs-available horizon per rung, and the grid pair the reconciliation
actually measured (its tolerance now scales with the grid skew instead of a flat
5pp).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 01:12:05 -04:00
|
|
|
double effSl = 0.0, effTp = 0.0;
|
|
|
|
|
double reach = LadderWinShare(idxList, n, sl, tp, effSl, effTp);
|
|
|
|
|
int needH = RequiredHorizonBars(sl, tp);
|
fix(labels): correct EffectiveSampleSize clamp order, share the horizon ladder, retract a false justification
Self-review of 1540ba8 against the FULL 6,930-era log rather than the first
three minutes of it. Three corrections.
1. EffectiveSampleSize() clamped in the wrong order. MathMax(2, MathMin(eff,
rawN)) returns 2 when rawN is 1 - an effective sample LARGER than the raw
one, shrinking the SE in exactly the direction the function exists to
prevent. Floor first, cap at rawN last.
2. The horizon cap rejected on the CEILING only, and said so as though that
made the label untruncated. It does not: the horizon ladder also snaps DOWN,
so a pair needing 317 bars is granted 256 and is silently truncated without
ever being flagged CLAMPED. Added SnapHorizonToLadder() / GrantedHorizonBars()
and the scale ladder now reports "needs N gets M" per rung. Rejection stays on
the ceiling alone - matching ReportGeometryExpectancyScan's '!' exactly, which
was the point - because rejecting on the snap-down would select rungs for
landing just above a ladder point rather than for anything about the market.
ComputeBarrierHorizonBars' private copy of the ladder is gone; there is now
one copy, which is the whole reason RequiredHorizonBars was factored out.
3. RETRACTED THE JUSTIFICATION IN 1540ba8's COMMENTS. That commit claimed the
overlap correction was needed because the operating point's null-of-the-
maximum gate fired on 47/73 Perceptron eras (64%) where a family-wise test
should fire on ~5%. Those 73 fits were the first three minutes of a
six-and-a-half-hour run. Over the full run:
PAI 47/3214 = 1.5% HYB 30/1200 = 2.5%
CONV 4/63 = 6.3% LSTM 75/915 = 8.2%
All at or below the null. The gate from 7414570 is working as designed and
PAI's 47 clears were a cold-start transient never repeated in 3,141 later
fits; its threshold over the run's second half has sd 0.01. The overlap
correction is still right - sqrt(p(1-p)/n) on overlapping labels is the wrong
formula - but it fixes no observed failure, and it costs nothing today
because no model is near the deploy line.
WHAT THE FULL RUN DOES CONFIRM, unchanged: the geometry ran away exactly as
described (2.00/6.00 h128 -> 3.49/6.99 h256 -> 4.86/9.71 h384, three passes,
stopping at q90 because the quantile ladder ended), the label stayed long-skewed
at Buy 42.9% / Sell 22.6%, and no checkpoint on any of the four models ever
cleared the deployability floor. Pooled declustered win rates: PAI 31.70%,
HYB 31.02%, LSTM 32.69%, CONV 31.57% - every one 4-6pp below the 37% always-long
chance rate and 1-2.7pp below the 33.7% cost-adjusted break-even.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 07:30:19 -04:00
|
|
|
int gotH = GrantedHorizonBars(sl, tp);
|
feat(measurement): fix zero-skill denominator, publish the deploy bar, measure lifespan per rung, add a MEASURE scale objective
The last run could not have demonstrated an edge either way, and nothing in the
log said so. Four changes so it does.
1. THE ZERO-SKILL LINE DIVIDED BY THE WRONG DENOMINATOR. m_oosWinLongTotal resets
every era; m_oosSamples only resets on a full model reset. So 'always-long %'
decayed as ~1/era: a run whose true rate is 37% printed 1.2% at era 33 and
0.0% at era 2219. This is the SAME bug already found and fixed for
logBuyPredPct thirty lines above ('era-15 Buy:2% that was really ~30%'), left
in the one line whose whole job is to be the reference every other number is
read against. Correct at era 1, wrong everywhere after - including the '62%
zero-skill' figure in the 2026-08-16 notes. Now per-era, and always-short is
finally readable.
2. THE DEPLOY GATE STATES ITS OWN BAR. 'edge -1pp' era after era cannot separate
'short by a hair' from 'short by an amount no strategy could cover'. The era
line now prints the required win rate, the SE, the effective n and the
lifespan it was deflated by; above 100% it says UNREACHABLE. At 4,738 OOS bars
and L=75.6 there are ~63 independent observations, putting the bar near 66% at
typical coverage.
3. LIFESPAN MEASURED PER RUNG. The first-passage cache already stores touch ages
at every ladder level, so each candidate geometry's resolution time is
readable without training on it - L-vs-width becomes a measurement across the
whole ladder in ONE run rather than a second chart. Each rung reports L,
n_eff, min provable edge and min provable EV.
4. SCALE OBJECTIVE IS PHASE-AWARE, defaulting to MEASURE. Width and detectability
are opposed: labels overlap by L, L grows like m*k = width^2 at fixed ratio,
so min provable EV ~ width^2 while the cost saving from width is only linear.
Doubling width quadruples the smallest EV you can prove. DEPLOY (widest that
clears reachability) is right once an edge is known; MEASURE (narrowest that
keeps round-trip spread under BARRIER_MAX_COST_FRACTION_PCT) is right while it
still has to be shown. The direction does not depend on the exponent, and
item 3 makes the exponent checkable.
Fixed in review: m_lastRungLifespan is cleared on every LadderWinShare entry or a
rejected rung reports the previous rung's lifespan as its own; per-rung
detectability is labelled IS-sample based (the deriver may not see the holdout),
so absolute figures are optimistic while the ranking is unaffected.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 10:37:05 -04:00
|
|
|
//--- DETECTABILITY AT THIS RUNG - the quantity BARRIER_SCALE_OBJECTIVE trades against width.
|
|
|
|
|
//--- n_eff = sample / this rung's own measured lifespan; the smallest edge 2 sigma can separate
|
|
|
|
|
//--- from chance follows, and multiplying by the width gives the smallest EV per trade that could
|
|
|
|
|
//--- ever be PROVEN at this geometry. Chance is the 1:RR break-even rather than a measured drift
|
|
|
|
|
//--- rate - this is a property of the geometry, not of any model.
|
|
|
|
|
//--- THE DENOMINATOR IS THE IS SAMPLE (n), because that is all the deriver is allowed to see - a
|
|
|
|
|
//--- geometry chosen with the holdout in view has used the holdout. The DEPLOY gate measures on
|
|
|
|
|
//--- the OOS window, which is smaller (roughly n x m_oosSplitPct/100), so these absolute figures
|
|
|
|
|
//--- are OPTIMISTIC by sqrt(that ratio) - about 1.5x at a 30% split. The RANKING across rungs is
|
|
|
|
|
//--- unaffected, since every rung is divided by the same n, and ranking is all this loop uses it
|
|
|
|
|
//--- for. Read the era line's DEPLOY BAR for the number that actually gates a deploy.
|
|
|
|
|
double rungLife = (m_lastRungLifespan > 0.0) ? m_lastRungLifespan : 1.0;
|
|
|
|
|
double rungEffN = MathMax((double)n / rungLife, 2.0);
|
|
|
|
|
double beP = 1.0 / (1.0 + BARRIER_TARGET_RR);
|
|
|
|
|
double rungSE = 100.0 * MathSqrt(beP * (1.0 - beP) / rungEffN);
|
|
|
|
|
double minEdge = EDGE_MIN_SIGMAS * rungSE; // percentage points
|
|
|
|
|
double minEV = minEdge / 100.0 * (sl + tp); // in ATR per trade
|
|
|
|
|
//--- Round-trip spread as a share of the move. This is what stops MEASURE mode at the tight end.
|
|
|
|
|
//--- m_spreadAtr is measured by ReportGeometryExpectancyScan, which runs at the END of this
|
|
|
|
|
//--- function - so on the FIRST derive pass it is still 0 and this filter is deliberately inert
|
|
|
|
|
//--- (costOK true) rather than rejecting every rung on an unmeasured cost. The pass-2 re-derivation
|
|
|
|
|
//--- applies it for real; that is exactly what the fixed-point iteration is for.
|
|
|
|
|
double costPct = (sl + tp > 0.0 && m_spreadAtr > 0.0)
|
|
|
|
|
? 100.0 * 2.0 * m_spreadAtr / (sl + tp) : 0.0;
|
|
|
|
|
bool costOK = (m_spreadAtr <= 0.0 || costPct <= BARRIER_MAX_COST_FRACTION_PCT);
|
fix(labels): correct EffectiveSampleSize clamp order, share the horizon ladder, retract a false justification
Self-review of 1540ba8 against the FULL 6,930-era log rather than the first
three minutes of it. Three corrections.
1. EffectiveSampleSize() clamped in the wrong order. MathMax(2, MathMin(eff,
rawN)) returns 2 when rawN is 1 - an effective sample LARGER than the raw
one, shrinking the SE in exactly the direction the function exists to
prevent. Floor first, cap at rawN last.
2. The horizon cap rejected on the CEILING only, and said so as though that
made the label untruncated. It does not: the horizon ladder also snaps DOWN,
so a pair needing 317 bars is granted 256 and is silently truncated without
ever being flagged CLAMPED. Added SnapHorizonToLadder() / GrantedHorizonBars()
and the scale ladder now reports "needs N gets M" per rung. Rejection stays on
the ceiling alone - matching ReportGeometryExpectancyScan's '!' exactly, which
was the point - because rejecting on the snap-down would select rungs for
landing just above a ladder point rather than for anything about the market.
ComputeBarrierHorizonBars' private copy of the ladder is gone; there is now
one copy, which is the whole reason RequiredHorizonBars was factored out.
3. RETRACTED THE JUSTIFICATION IN 1540ba8's COMMENTS. That commit claimed the
overlap correction was needed because the operating point's null-of-the-
maximum gate fired on 47/73 Perceptron eras (64%) where a family-wise test
should fire on ~5%. Those 73 fits were the first three minutes of a
six-and-a-half-hour run. Over the full run:
PAI 47/3214 = 1.5% HYB 30/1200 = 2.5%
CONV 4/63 = 6.3% LSTM 75/915 = 8.2%
All at or below the null. The gate from 7414570 is working as designed and
PAI's 47 clears were a cold-start transient never repeated in 3,141 later
fits; its threshold over the run's second half has sd 0.01. The overlap
correction is still right - sqrt(p(1-p)/n) on overlapping labels is the wrong
formula - but it fixes no observed failure, and it costs nothing today
because no model is near the deploy line.
WHAT THE FULL RUN DOES CONFIRM, unchanged: the geometry ran away exactly as
described (2.00/6.00 h128 -> 3.49/6.99 h256 -> 4.86/9.71 h384, three passes,
stopping at q90 because the quantile ladder ended), the label stayed long-skewed
at Buy 42.9% / Sell 22.6%, and no checkpoint on any of the four models ever
cleared the deployability floor. Pooled declustered win rates: PAI 31.70%,
HYB 31.02%, LSTM 32.69%, CONV 31.57% - every one 4-6pp below the 37% always-long
chance rate and 1-2.7pp below the 33.7% cost-adjusted break-even.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 07:30:19 -04:00
|
|
|
//--- REJECT ON THE CEILING ONLY, matching ReportGeometryExpectancyScan's '!' exactly - that is the
|
|
|
|
|
//--- whole point of applying one rule in two places. gotH < needH is the SEPARATE, milder
|
|
|
|
|
//--- truncation the ladder's snap-down always imposes (317 needed -> 256 granted); it is reported,
|
|
|
|
|
//--- not rejected, because rejecting on it would select rungs for landing just above a ladder point
|
|
|
|
|
//--- rather than for anything about the market. The timeout share is what says whether it bites.
|
fix(labels): overlapping-label sample correction + horizon cap on the scale ladder
Three defects, all surfaced by the 2026-08-17 SP500 H4 run that shipped
stop 4.86 / target 9.71 (width 14.57*ATR, horizon 384).
1. EVERY STANDARD ERROR ASSUMED INDEPENDENT SAMPLES. Triple-barrier labels
started one per bar overlap by the label's lifespan, so n calls are worth
~n/L independent observations (Lopez de Prado, AFML ch. 4 - sample
uniqueness). All three sqrt(p(1-p)/n) sites divided by the RAW count.
The tell: the operating point's null-of-the-maximum gate is family-wise and
should fire on ~5% of eras under the null. Measured fire rates - PAI 47/73
(64%), ConvLSTM 9/24, LSTM 8/21 (38%), CONV 4/62 (6%). CONV, the only model
whose margin distribution admits few bins, sat on the null; the rest cleared
a bar that was too low by ~sqrt(L). PAI's deployed threshold consequently
alternated between the ENDS of its own range era to era (0.10 -> 0.88 ->
0.86 -> 0.66; coverage 16% <-> 73%).
TripleBarrierLabel now records when each label became KNOWABLE - the first
winning touch, or both stops, or the timeout - and the prebuild accumulates
the mean. EffectiveSampleSize() feeds the operating point, the member deploy
gate and the ensemble vote gate. Conservative by construction (n/L is an
upper bound on the damage); gates get harder, never easier.
2. THE SCALE LADDER RAN AWAY, again. Horizon scales as swingMedian*sl*tp, and
since 4d8cb08 reachability is measured OVER that horizon - so a wider rung
buys itself the time that makes it look reachable. Same target -> horizon ->
reach -> target loop the excursion window is kept short to avoid; fixing the
window confusion reopened it through the other door. It walked 128 -> 256 ->
384 bars and stopped at q90, the widest rung there is, with every rung
reading 39-48% against a 20% floor. A floor nothing fails selects nothing.
Rungs whose required horizon exceeds BARRIER_HORIZON_MAX are now rejected -
the same rule ReportGeometryExpectancyScan already applied. It was printing
the shipped pair as CLAMPED and disqualified ('h384!') two lines under the
deriver that chose it: two subsystems, one geometry, opposite verdicts.
3. THE RUNG SNAP DESTROYED THE RATIO IT WAS COMPARING. Both legs snapped
independently to the coarse first-passage grid, re-rating each candidate:
q90 4.86/9.71 -> 5.00/10.00 (2.00), q85 4.07/8.14 -> 5.00/10.00 (IDENTICAL
measurement), q75 3.07/6.13 -> 4.00/6.50 (1.63 - a nearer target). So the
ladder compared win shares taken at ratios from 1.63 to 2.17 and read the
differences as scale. It is why the reach column came out non-monotone in
width (q75 48.5% above q90 42.9%). The stop now snaps to its nearest rung in
log space and the target follows the ratio off it; the pair actually measured
is returned and logged, so a collision reads as a collision.
Also: LadderWinShare guarded against the conditional (fractal) geometry path,
which fills n from m_fracLegCount while leaving idxList empty - a latent
out-of-bounds on a currently-dead path.
New log lines: mean label lifespan and effective n on the label-cache line, the
required-vs-available horizon per rung, and the grid pair the reconciliation
actually measured (its tolerance now scales with the grid skew instead of a flat
5pp).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 01:12:05 -04:00
|
|
|
bool fitsH = (needH <= BARRIER_HORIZON_MAX);
|
|
|
|
|
//--- The measured pair is printed beside the requested one whenever the grid could not express the
|
|
|
|
|
//--- request, so a collision between two quantiles is visible as a collision rather than as two
|
|
|
|
|
//--- rungs that happen to score identically.
|
|
|
|
|
string effNote = "";
|
|
|
|
|
if(effSl > 0.0 && (MathAbs(effSl - sl) > 0.005 || MathAbs(effTp - tp) > 0.005))
|
|
|
|
|
effNote = StringFormat(" @grid %.2f/%.2f", effSl, effTp);
|
feat(measurement): fix zero-skill denominator, publish the deploy bar, measure lifespan per rung, add a MEASURE scale objective
The last run could not have demonstrated an edge either way, and nothing in the
log said so. Four changes so it does.
1. THE ZERO-SKILL LINE DIVIDED BY THE WRONG DENOMINATOR. m_oosWinLongTotal resets
every era; m_oosSamples only resets on a full model reset. So 'always-long %'
decayed as ~1/era: a run whose true rate is 37% printed 1.2% at era 33 and
0.0% at era 2219. This is the SAME bug already found and fixed for
logBuyPredPct thirty lines above ('era-15 Buy:2% that was really ~30%'), left
in the one line whose whole job is to be the reference every other number is
read against. Correct at era 1, wrong everywhere after - including the '62%
zero-skill' figure in the 2026-08-16 notes. Now per-era, and always-short is
finally readable.
2. THE DEPLOY GATE STATES ITS OWN BAR. 'edge -1pp' era after era cannot separate
'short by a hair' from 'short by an amount no strategy could cover'. The era
line now prints the required win rate, the SE, the effective n and the
lifespan it was deflated by; above 100% it says UNREACHABLE. At 4,738 OOS bars
and L=75.6 there are ~63 independent observations, putting the bar near 66% at
typical coverage.
3. LIFESPAN MEASURED PER RUNG. The first-passage cache already stores touch ages
at every ladder level, so each candidate geometry's resolution time is
readable without training on it - L-vs-width becomes a measurement across the
whole ladder in ONE run rather than a second chart. Each rung reports L,
n_eff, min provable edge and min provable EV.
4. SCALE OBJECTIVE IS PHASE-AWARE, defaulting to MEASURE. Width and detectability
are opposed: labels overlap by L, L grows like m*k = width^2 at fixed ratio,
so min provable EV ~ width^2 while the cost saving from width is only linear.
Doubling width quadruples the smallest EV you can prove. DEPLOY (widest that
clears reachability) is right once an edge is known; MEASURE (narrowest that
keeps round-trip spread under BARRIER_MAX_COST_FRACTION_PCT) is right while it
still has to be shown. The direction does not depend on the exponent, and
item 3 makes the exponent checkable.
Fixed in review: m_lastRungLifespan is cleared on every LadderWinShare entry or a
rejected rung reports the previous rung's lifespan as its own; per-rung
detectability is labelled IS-sample based (the deriver may not see the holdout),
so absolute figures are optimistic while the ranking is unaffected.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 10:37:05 -04:00
|
|
|
rungRows += StringFormat("%sq%.0f(stop %.2f target %.2f width %.2f reach %.1f%%%s needs %d gets %d"
|
|
|
|
|
" | L %.0f -> n_eff %.0f, min provable edge %.1fpp = %.2f ATR/trade,"
|
|
|
|
|
" cost %.1f%%%s%s)",
|
fix(labels): overlapping-label sample correction + horizon cap on the scale ladder
Three defects, all surfaced by the 2026-08-17 SP500 H4 run that shipped
stop 4.86 / target 9.71 (width 14.57*ATR, horizon 384).
1. EVERY STANDARD ERROR ASSUMED INDEPENDENT SAMPLES. Triple-barrier labels
started one per bar overlap by the label's lifespan, so n calls are worth
~n/L independent observations (Lopez de Prado, AFML ch. 4 - sample
uniqueness). All three sqrt(p(1-p)/n) sites divided by the RAW count.
The tell: the operating point's null-of-the-maximum gate is family-wise and
should fire on ~5% of eras under the null. Measured fire rates - PAI 47/73
(64%), ConvLSTM 9/24, LSTM 8/21 (38%), CONV 4/62 (6%). CONV, the only model
whose margin distribution admits few bins, sat on the null; the rest cleared
a bar that was too low by ~sqrt(L). PAI's deployed threshold consequently
alternated between the ENDS of its own range era to era (0.10 -> 0.88 ->
0.86 -> 0.66; coverage 16% <-> 73%).
TripleBarrierLabel now records when each label became KNOWABLE - the first
winning touch, or both stops, or the timeout - and the prebuild accumulates
the mean. EffectiveSampleSize() feeds the operating point, the member deploy
gate and the ensemble vote gate. Conservative by construction (n/L is an
upper bound on the damage); gates get harder, never easier.
2. THE SCALE LADDER RAN AWAY, again. Horizon scales as swingMedian*sl*tp, and
since 4d8cb08 reachability is measured OVER that horizon - so a wider rung
buys itself the time that makes it look reachable. Same target -> horizon ->
reach -> target loop the excursion window is kept short to avoid; fixing the
window confusion reopened it through the other door. It walked 128 -> 256 ->
384 bars and stopped at q90, the widest rung there is, with every rung
reading 39-48% against a 20% floor. A floor nothing fails selects nothing.
Rungs whose required horizon exceeds BARRIER_HORIZON_MAX are now rejected -
the same rule ReportGeometryExpectancyScan already applied. It was printing
the shipped pair as CLAMPED and disqualified ('h384!') two lines under the
deriver that chose it: two subsystems, one geometry, opposite verdicts.
3. THE RUNG SNAP DESTROYED THE RATIO IT WAS COMPARING. Both legs snapped
independently to the coarse first-passage grid, re-rating each candidate:
q90 4.86/9.71 -> 5.00/10.00 (2.00), q85 4.07/8.14 -> 5.00/10.00 (IDENTICAL
measurement), q75 3.07/6.13 -> 4.00/6.50 (1.63 - a nearer target). So the
ladder compared win shares taken at ratios from 1.63 to 2.17 and read the
differences as scale. It is why the reach column came out non-monotone in
width (q75 48.5% above q90 42.9%). The stop now snaps to its nearest rung in
log space and the target follows the ratio off it; the pair actually measured
is returned and logged, so a collision reads as a collision.
Also: LadderWinShare guarded against the conditional (fractal) geometry path,
which fills n from m_fracLegCount while leaving idxList empty - a latent
out-of-bounds on a currently-dead path.
New log lines: mean label lifespan and effective n on the label-cache line, the
required-vs-available horizon per rung, and the grid pair the reconciliation
actually measured (its tolerance now scales with the grid skew instead of a flat
5pp).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 01:12:05 -04:00
|
|
|
(rungRows == "" ? "" : " "), 100.0 * q, sl, tp, sl + tp, reach, effNote,
|
feat(measurement): fix zero-skill denominator, publish the deploy bar, measure lifespan per rung, add a MEASURE scale objective
The last run could not have demonstrated an edge either way, and nothing in the
log said so. Four changes so it does.
1. THE ZERO-SKILL LINE DIVIDED BY THE WRONG DENOMINATOR. m_oosWinLongTotal resets
every era; m_oosSamples only resets on a full model reset. So 'always-long %'
decayed as ~1/era: a run whose true rate is 37% printed 1.2% at era 33 and
0.0% at era 2219. This is the SAME bug already found and fixed for
logBuyPredPct thirty lines above ('era-15 Buy:2% that was really ~30%'), left
in the one line whose whole job is to be the reference every other number is
read against. Correct at era 1, wrong everywhere after - including the '62%
zero-skill' figure in the 2026-08-16 notes. Now per-era, and always-short is
finally readable.
2. THE DEPLOY GATE STATES ITS OWN BAR. 'edge -1pp' era after era cannot separate
'short by a hair' from 'short by an amount no strategy could cover'. The era
line now prints the required win rate, the SE, the effective n and the
lifespan it was deflated by; above 100% it says UNREACHABLE. At 4,738 OOS bars
and L=75.6 there are ~63 independent observations, putting the bar near 66% at
typical coverage.
3. LIFESPAN MEASURED PER RUNG. The first-passage cache already stores touch ages
at every ladder level, so each candidate geometry's resolution time is
readable without training on it - L-vs-width becomes a measurement across the
whole ladder in ONE run rather than a second chart. Each rung reports L,
n_eff, min provable edge and min provable EV.
4. SCALE OBJECTIVE IS PHASE-AWARE, defaulting to MEASURE. Width and detectability
are opposed: labels overlap by L, L grows like m*k = width^2 at fixed ratio,
so min provable EV ~ width^2 while the cost saving from width is only linear.
Doubling width quadruples the smallest EV you can prove. DEPLOY (widest that
clears reachability) is right once an edge is known; MEASURE (narrowest that
keeps round-trip spread under BARRIER_MAX_COST_FRACTION_PCT) is right while it
still has to be shown. The direction does not depend on the exponent, and
item 3 makes the exponent checkable.
Fixed in review: m_lastRungLifespan is cleared on every LadderWinShare entry or a
rejected rung reports the previous rung's lifespan as its own; per-rung
detectability is labelled IS-sample based (the deriver may not see the holdout),
so absolute figures are optimistic while the ranking is unaffected.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 10:37:05 -04:00
|
|
|
needH, gotH, rungLife, rungEffN, minEdge, minEV, costPct,
|
|
|
|
|
(costOK ? "" : " COST-REJECTED"), (fitsH ? "" : " CLAMPED-REJECTED"));
|
|
|
|
|
//--- Both objectives share the reachability floor and the horizon ceiling; they differ only in
|
|
|
|
|
//--- which end of the surviving set they take. DEPLOY keeps the first (widest, ladder is ordered
|
|
|
|
|
//--- widest-first); MEASURE keeps overwriting, so it ends on the last - the narrowest that is
|
|
|
|
|
//--- still cost-efficient.
|
|
|
|
|
bool eligible = fitsH && costOK && reach >= BARRIER_MIN_TP_REACH_PCT;
|
|
|
|
|
bool takeIt = (BARRIER_SCALE_OBJECTIVE == BARRIER_SCALE_DEPLOY) ? (slRaw <= 0.0) : true;
|
|
|
|
|
if(eligible && takeIt)
|
fix(geometry): the target was small BY CONSTRUCTION - ratio is now policy, scale is measured, ladder ceiling removed
The derivation read the stop from q75 of ADVERSE travel and the target from q50
of FAVOURABLE travel. Over one horizon those distributions are broadly the same
shape, so q75 > q50 MECHANICALLY - the target came out smaller than the stop no
matter what the market did. SP500 H4 shipped stop 3.07 / target 1.70: a 0.55:1
payoff needing 64.3%. That was never a measurement, it was two mismatched
constants.
The reachability line printed beside it - "target on 50.0% of bars, stop on
25.0%" - is exactly 1-q50 and 1-q75. Tautological. It cannot disconfirm
anything, and it read as validation.
WIDTH AND RATIO ARE INDEPENDENT AND ONLY ONE PAYS. EV = edge x width;
ratio is EV-neutral (a driftless walk reaches +m before -k with probability
k/(k+m), which IS break-even). Width is what buys cost efficiency: the spread
is a fixed 0.047*ATR here, so the shipped 4.77*ATR width paid it 21 times per
unit of travel. So:
RATIO = policy. BARRIER_TARGET_RR = 2.0 (user's 1:2). Break-even 33.3%.
SCALE = measured. The stop quantile is chosen from a ladder, WIDEST FIRST,
taking the first rung whose implied 2x target is still reached often
enough to be a trainable class.
That last clause is the difference from the min-reward:risk raise removed in
2026-08-09, which forced target = 2 x stop with NO reachability test, landed on
6.66*ATR reachable on 3.3% of bars, and trained the model to predict something
that essentially never happened. Same ratio; the scale now retreats until the
data says the target is attainable. Every rung is logged.
LADDER CEILING REMOVED. BARRIER_LADDER stopped at 5.00 and the expectancy scan's
"best resolvable pair on width alone" came back as stop 5.05 / target 4.95 - it
pinned to the top rung. A recommendation landing exactly on the edge of its own
search space is a boundary, not a finding: it cannot tell "5 ATR is optimal"
from "5 ATR is all we allowed". Extended to 20*ATR (8 -> 14 rungs). Nothing else
needs editing - every consumer is parameterised by BARRIER_LADDER_COUNT - and
the horizon constraints (decided >= 60%, reachability floor) now bind instead of
a constant.
THE SCAN COULD NOT SEE THE SHIPPED GEOMETRY. ReportBarrierGeometryScan looked
the configured pair up in its integer grid, and DeriveBarrierGeometry produces
CONTINUOUS multiples (3.07/1.70) that can never equal a grid point - so
cfgExcess stayed at its -1.0 sentinel and the report printed "configured 3:2
scores -1.00000", which reads as a catastrophic score and actually means "never
evaluated". Worse, the grid skipped target<stop entirely because it "inverts the
trade's whole premise" - while the derivation was shipping exactly that. The
incumbent is now always scored as a peer (never crowned; it is already in force
and is not an enum pairing the scan could adopt).
BREAK-EVEN NOW INCLUDES THE SPREAD. Every report quoted the frictionless
SL/(SL+TP). On SP500 H4 that read 64.3% while the MEASURED zero-skill rate was
62.1% - a 2.2pp gap that IS the cost, and that made every model look 2.2pp
better than it was. CostAdjustedBreakEvenPct() prices a win at (TP - spread) and
a loss at (SL + spread), matching the expectancy scan's convention exactly so
the two reports cannot disagree.
It also feeds FitDirConfThreshold, which is the correctness half: the operating
point subtracts break-even from precision, so the frictionless figure made every
candidate threshold look better by the width of the spread - 2.2pp against a
measured edge of 2.3pp, i.e. very nearly all of it.
Era line now carries both: "break-even 64.3% frictionless, 66.6% AFTER SPREAD".
Forces a full relabel and retrain. Requested.
NOT COMPILED - user compiles.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 00:34:32 -04:00
|
|
|
{
|
|
|
|
|
slRaw = sl;
|
|
|
|
|
tpRaw = tp;
|
|
|
|
|
chosenQ = q;
|
|
|
|
|
chosenReach = reach;
|
|
|
|
|
}
|
|
|
|
|
}
|
|
|
|
|
if(slRaw <= 0.0)
|
|
|
|
|
{
|
|
|
|
|
//--- No rung clears the floor: the horizon cannot deliver a 1:RR target at ANY survivable stop on
|
|
|
|
|
//--- this instrument. Take the tightest rung (the most reachable one there is) and say so - the
|
|
|
|
|
//--- ratio is risk policy, and the honest response is to price it, not to silently abandon it.
|
|
|
|
|
double q = BARRIER_SL_QUANTILE_LADDER[BARRIER_SL_QUANTILE_COUNT - 1];
|
|
|
|
|
slRaw = dn[(int)MathMin(q * n, n - 1)];
|
|
|
|
|
if(slRaw < MIN_SL_ATR_MULTIPLIER)
|
|
|
|
|
slRaw = MIN_SL_ATR_MULTIPLIER;
|
|
|
|
|
tpRaw = BARRIER_TARGET_RR * slRaw;
|
|
|
|
|
chosenQ = q;
|
fix(labels): overlapping-label sample correction + horizon cap on the scale ladder
Three defects, all surfaced by the 2026-08-17 SP500 H4 run that shipped
stop 4.86 / target 9.71 (width 14.57*ATR, horizon 384).
1. EVERY STANDARD ERROR ASSUMED INDEPENDENT SAMPLES. Triple-barrier labels
started one per bar overlap by the label's lifespan, so n calls are worth
~n/L independent observations (Lopez de Prado, AFML ch. 4 - sample
uniqueness). All three sqrt(p(1-p)/n) sites divided by the RAW count.
The tell: the operating point's null-of-the-maximum gate is family-wise and
should fire on ~5% of eras under the null. Measured fire rates - PAI 47/73
(64%), ConvLSTM 9/24, LSTM 8/21 (38%), CONV 4/62 (6%). CONV, the only model
whose margin distribution admits few bins, sat on the null; the rest cleared
a bar that was too low by ~sqrt(L). PAI's deployed threshold consequently
alternated between the ENDS of its own range era to era (0.10 -> 0.88 ->
0.86 -> 0.66; coverage 16% <-> 73%).
TripleBarrierLabel now records when each label became KNOWABLE - the first
winning touch, or both stops, or the timeout - and the prebuild accumulates
the mean. EffectiveSampleSize() feeds the operating point, the member deploy
gate and the ensemble vote gate. Conservative by construction (n/L is an
upper bound on the damage); gates get harder, never easier.
2. THE SCALE LADDER RAN AWAY, again. Horizon scales as swingMedian*sl*tp, and
since 4d8cb08 reachability is measured OVER that horizon - so a wider rung
buys itself the time that makes it look reachable. Same target -> horizon ->
reach -> target loop the excursion window is kept short to avoid; fixing the
window confusion reopened it through the other door. It walked 128 -> 256 ->
384 bars and stopped at q90, the widest rung there is, with every rung
reading 39-48% against a 20% floor. A floor nothing fails selects nothing.
Rungs whose required horizon exceeds BARRIER_HORIZON_MAX are now rejected -
the same rule ReportGeometryExpectancyScan already applied. It was printing
the shipped pair as CLAMPED and disqualified ('h384!') two lines under the
deriver that chose it: two subsystems, one geometry, opposite verdicts.
3. THE RUNG SNAP DESTROYED THE RATIO IT WAS COMPARING. Both legs snapped
independently to the coarse first-passage grid, re-rating each candidate:
q90 4.86/9.71 -> 5.00/10.00 (2.00), q85 4.07/8.14 -> 5.00/10.00 (IDENTICAL
measurement), q75 3.07/6.13 -> 4.00/6.50 (1.63 - a nearer target). So the
ladder compared win shares taken at ratios from 1.63 to 2.17 and read the
differences as scale. It is why the reach column came out non-monotone in
width (q75 48.5% above q90 42.9%). The stop now snaps to its nearest rung in
log space and the target follows the ratio off it; the pair actually measured
is returned and logged, so a collision reads as a collision.
Also: LadderWinShare guarded against the conditional (fractal) geometry path,
which fills n from m_fracLegCount while leaving idxList empty - a latent
out-of-bounds on a currently-dead path.
New log lines: mean label lifespan and effective n on the label-cache line, the
required-vs-available horizon per rung, and the grid pair the reconciliation
actually measured (its tolerance now scales with the grid skew instead of a flat
5pp).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 01:12:05 -04:00
|
|
|
double fbSl = 0.0, fbTp = 0.0;
|
|
|
|
|
chosenReach = LadderWinShare(idxList, n, slRaw, tpRaw, fbSl, fbTp);
|
fix(geometry): the target was small BY CONSTRUCTION - ratio is now policy, scale is measured, ladder ceiling removed
The derivation read the stop from q75 of ADVERSE travel and the target from q50
of FAVOURABLE travel. Over one horizon those distributions are broadly the same
shape, so q75 > q50 MECHANICALLY - the target came out smaller than the stop no
matter what the market did. SP500 H4 shipped stop 3.07 / target 1.70: a 0.55:1
payoff needing 64.3%. That was never a measurement, it was two mismatched
constants.
The reachability line printed beside it - "target on 50.0% of bars, stop on
25.0%" - is exactly 1-q50 and 1-q75. Tautological. It cannot disconfirm
anything, and it read as validation.
WIDTH AND RATIO ARE INDEPENDENT AND ONLY ONE PAYS. EV = edge x width;
ratio is EV-neutral (a driftless walk reaches +m before -k with probability
k/(k+m), which IS break-even). Width is what buys cost efficiency: the spread
is a fixed 0.047*ATR here, so the shipped 4.77*ATR width paid it 21 times per
unit of travel. So:
RATIO = policy. BARRIER_TARGET_RR = 2.0 (user's 1:2). Break-even 33.3%.
SCALE = measured. The stop quantile is chosen from a ladder, WIDEST FIRST,
taking the first rung whose implied 2x target is still reached often
enough to be a trainable class.
That last clause is the difference from the min-reward:risk raise removed in
2026-08-09, which forced target = 2 x stop with NO reachability test, landed on
6.66*ATR reachable on 3.3% of bars, and trained the model to predict something
that essentially never happened. Same ratio; the scale now retreats until the
data says the target is attainable. Every rung is logged.
LADDER CEILING REMOVED. BARRIER_LADDER stopped at 5.00 and the expectancy scan's
"best resolvable pair on width alone" came back as stop 5.05 / target 4.95 - it
pinned to the top rung. A recommendation landing exactly on the edge of its own
search space is a boundary, not a finding: it cannot tell "5 ATR is optimal"
from "5 ATR is all we allowed". Extended to 20*ATR (8 -> 14 rungs). Nothing else
needs editing - every consumer is parameterised by BARRIER_LADDER_COUNT - and
the horizon constraints (decided >= 60%, reachability floor) now bind instead of
a constant.
THE SCAN COULD NOT SEE THE SHIPPED GEOMETRY. ReportBarrierGeometryScan looked
the configured pair up in its integer grid, and DeriveBarrierGeometry produces
CONTINUOUS multiples (3.07/1.70) that can never equal a grid point - so
cfgExcess stayed at its -1.0 sentinel and the report printed "configured 3:2
scores -1.00000", which reads as a catastrophic score and actually means "never
evaluated". Worse, the grid skipped target<stop entirely because it "inverts the
trade's whole premise" - while the derivation was shipping exactly that. The
incumbent is now always scored as a peer (never crowned; it is already in force
and is not an enum pairing the scan could adopt).
BREAK-EVEN NOW INCLUDES THE SPREAD. Every report quoted the frictionless
SL/(SL+TP). On SP500 H4 that read 64.3% while the MEASURED zero-skill rate was
62.1% - a 2.2pp gap that IS the cost, and that made every model look 2.2pp
better than it was. CostAdjustedBreakEvenPct() prices a win at (TP - spread) and
a loss at (SL + spread), matching the expectancy scan's convention exactly so
the two reports cannot disagree.
It also feeds FitDirConfThreshold, which is the correctness half: the operating
point subtracts break-even from precision, so the frictionless figure made every
candidate threshold look better by the width of the spread - 2.2pp against a
measured edge of 2.3pp, i.e. very nearly all of it.
Era line now carries both: "break-even 64.3% frictionless, 66.6% AFTER SPREAD".
Forces a full relabel and retrain. Requested.
NOT COMPILED - user compiles.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 00:34:32 -04:00
|
|
|
Print(ID + StringFormat(": WARNING - no stop quantile produced a %.1f:1 target reached on at least "
|
fix(labels): overlapping-label sample correction + horizon cap on the scale ladder
Three defects, all surfaced by the 2026-08-17 SP500 H4 run that shipped
stop 4.86 / target 9.71 (width 14.57*ATR, horizon 384).
1. EVERY STANDARD ERROR ASSUMED INDEPENDENT SAMPLES. Triple-barrier labels
started one per bar overlap by the label's lifespan, so n calls are worth
~n/L independent observations (Lopez de Prado, AFML ch. 4 - sample
uniqueness). All three sqrt(p(1-p)/n) sites divided by the RAW count.
The tell: the operating point's null-of-the-maximum gate is family-wise and
should fire on ~5% of eras under the null. Measured fire rates - PAI 47/73
(64%), ConvLSTM 9/24, LSTM 8/21 (38%), CONV 4/62 (6%). CONV, the only model
whose margin distribution admits few bins, sat on the null; the rest cleared
a bar that was too low by ~sqrt(L). PAI's deployed threshold consequently
alternated between the ENDS of its own range era to era (0.10 -> 0.88 ->
0.86 -> 0.66; coverage 16% <-> 73%).
TripleBarrierLabel now records when each label became KNOWABLE - the first
winning touch, or both stops, or the timeout - and the prebuild accumulates
the mean. EffectiveSampleSize() feeds the operating point, the member deploy
gate and the ensemble vote gate. Conservative by construction (n/L is an
upper bound on the damage); gates get harder, never easier.
2. THE SCALE LADDER RAN AWAY, again. Horizon scales as swingMedian*sl*tp, and
since 4d8cb08 reachability is measured OVER that horizon - so a wider rung
buys itself the time that makes it look reachable. Same target -> horizon ->
reach -> target loop the excursion window is kept short to avoid; fixing the
window confusion reopened it through the other door. It walked 128 -> 256 ->
384 bars and stopped at q90, the widest rung there is, with every rung
reading 39-48% against a 20% floor. A floor nothing fails selects nothing.
Rungs whose required horizon exceeds BARRIER_HORIZON_MAX are now rejected -
the same rule ReportGeometryExpectancyScan already applied. It was printing
the shipped pair as CLAMPED and disqualified ('h384!') two lines under the
deriver that chose it: two subsystems, one geometry, opposite verdicts.
3. THE RUNG SNAP DESTROYED THE RATIO IT WAS COMPARING. Both legs snapped
independently to the coarse first-passage grid, re-rating each candidate:
q90 4.86/9.71 -> 5.00/10.00 (2.00), q85 4.07/8.14 -> 5.00/10.00 (IDENTICAL
measurement), q75 3.07/6.13 -> 4.00/6.50 (1.63 - a nearer target). So the
ladder compared win shares taken at ratios from 1.63 to 2.17 and read the
differences as scale. It is why the reach column came out non-monotone in
width (q75 48.5% above q90 42.9%). The stop now snaps to its nearest rung in
log space and the target follows the ratio off it; the pair actually measured
is returned and logged, so a collision reads as a collision.
Also: LadderWinShare guarded against the conditional (fractal) geometry path,
which fills n from m_fracLegCount while leaving idxList empty - a latent
out-of-bounds on a currently-dead path.
New log lines: mean label lifespan and effective n on the label-cache line, the
required-vs-available horizon per rung, and the grid pair the reconciliation
actually measured (its tolerance now scales with the grid skew instead of a flat
5pp).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 01:12:05 -04:00
|
|
|
"%.0f%% of bars inside the %d-bar horizon AND resolvable within the %d-bar "
|
|
|
|
|
"ceiling. Taking the tightest rung (q%.0f) at %.1f%% reachability. The "
|
|
|
|
|
"positive class will be rare and training will be correspondingly hard; "
|
|
|
|
|
"raise BARRIER_HORIZON_MAX or lower BARRIER_TARGET_RR if that proves "
|
|
|
|
|
"untrainable.",
|
fix(geometry): the target was small BY CONSTRUCTION - ratio is now policy, scale is measured, ladder ceiling removed
The derivation read the stop from q75 of ADVERSE travel and the target from q50
of FAVOURABLE travel. Over one horizon those distributions are broadly the same
shape, so q75 > q50 MECHANICALLY - the target came out smaller than the stop no
matter what the market did. SP500 H4 shipped stop 3.07 / target 1.70: a 0.55:1
payoff needing 64.3%. That was never a measurement, it was two mismatched
constants.
The reachability line printed beside it - "target on 50.0% of bars, stop on
25.0%" - is exactly 1-q50 and 1-q75. Tautological. It cannot disconfirm
anything, and it read as validation.
WIDTH AND RATIO ARE INDEPENDENT AND ONLY ONE PAYS. EV = edge x width;
ratio is EV-neutral (a driftless walk reaches +m before -k with probability
k/(k+m), which IS break-even). Width is what buys cost efficiency: the spread
is a fixed 0.047*ATR here, so the shipped 4.77*ATR width paid it 21 times per
unit of travel. So:
RATIO = policy. BARRIER_TARGET_RR = 2.0 (user's 1:2). Break-even 33.3%.
SCALE = measured. The stop quantile is chosen from a ladder, WIDEST FIRST,
taking the first rung whose implied 2x target is still reached often
enough to be a trainable class.
That last clause is the difference from the min-reward:risk raise removed in
2026-08-09, which forced target = 2 x stop with NO reachability test, landed on
6.66*ATR reachable on 3.3% of bars, and trained the model to predict something
that essentially never happened. Same ratio; the scale now retreats until the
data says the target is attainable. Every rung is logged.
LADDER CEILING REMOVED. BARRIER_LADDER stopped at 5.00 and the expectancy scan's
"best resolvable pair on width alone" came back as stop 5.05 / target 4.95 - it
pinned to the top rung. A recommendation landing exactly on the edge of its own
search space is a boundary, not a finding: it cannot tell "5 ATR is optimal"
from "5 ATR is all we allowed". Extended to 20*ATR (8 -> 14 rungs). Nothing else
needs editing - every consumer is parameterised by BARRIER_LADDER_COUNT - and
the horizon constraints (decided >= 60%, reachability floor) now bind instead of
a constant.
THE SCAN COULD NOT SEE THE SHIPPED GEOMETRY. ReportBarrierGeometryScan looked
the configured pair up in its integer grid, and DeriveBarrierGeometry produces
CONTINUOUS multiples (3.07/1.70) that can never equal a grid point - so
cfgExcess stayed at its -1.0 sentinel and the report printed "configured 3:2
scores -1.00000", which reads as a catastrophic score and actually means "never
evaluated". Worse, the grid skipped target<stop entirely because it "inverts the
trade's whole premise" - while the derivation was shipping exactly that. The
incumbent is now always scored as a peer (never crowned; it is already in force
and is not an enum pairing the scan could adopt).
BREAK-EVEN NOW INCLUDES THE SPREAD. Every report quoted the frictionless
SL/(SL+TP). On SP500 H4 that read 64.3% while the MEASURED zero-skill rate was
62.1% - a 2.2pp gap that IS the cost, and that made every model look 2.2pp
better than it was. CostAdjustedBreakEvenPct() prices a win at (TP - spread) and
a loss at (SL + spread), matching the expectancy scan's convention exactly so
the two reports cannot disagree.
It also feeds FitDirConfThreshold, which is the correctness half: the operating
point subtracts break-even from precision, so the frictionless figure made every
candidate threshold look better by the width of the spread - 2.2pp against a
measured edge of 2.3pp, i.e. very nearly all of it.
Era line now carries both: "break-even 64.3% frictionless, 66.6% AFTER SPREAD".
Forces a full relabel and retrain. Requested.
NOT COMPILED - user compiles.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 00:34:32 -04:00
|
|
|
BARRIER_TARGET_RR, BARRIER_MIN_TP_REACH_PCT, m_barrierHorizonBars,
|
fix(labels): overlapping-label sample correction + horizon cap on the scale ladder
Three defects, all surfaced by the 2026-08-17 SP500 H4 run that shipped
stop 4.86 / target 9.71 (width 14.57*ATR, horizon 384).
1. EVERY STANDARD ERROR ASSUMED INDEPENDENT SAMPLES. Triple-barrier labels
started one per bar overlap by the label's lifespan, so n calls are worth
~n/L independent observations (Lopez de Prado, AFML ch. 4 - sample
uniqueness). All three sqrt(p(1-p)/n) sites divided by the RAW count.
The tell: the operating point's null-of-the-maximum gate is family-wise and
should fire on ~5% of eras under the null. Measured fire rates - PAI 47/73
(64%), ConvLSTM 9/24, LSTM 8/21 (38%), CONV 4/62 (6%). CONV, the only model
whose margin distribution admits few bins, sat on the null; the rest cleared
a bar that was too low by ~sqrt(L). PAI's deployed threshold consequently
alternated between the ENDS of its own range era to era (0.10 -> 0.88 ->
0.86 -> 0.66; coverage 16% <-> 73%).
TripleBarrierLabel now records when each label became KNOWABLE - the first
winning touch, or both stops, or the timeout - and the prebuild accumulates
the mean. EffectiveSampleSize() feeds the operating point, the member deploy
gate and the ensemble vote gate. Conservative by construction (n/L is an
upper bound on the damage); gates get harder, never easier.
2. THE SCALE LADDER RAN AWAY, again. Horizon scales as swingMedian*sl*tp, and
since 4d8cb08 reachability is measured OVER that horizon - so a wider rung
buys itself the time that makes it look reachable. Same target -> horizon ->
reach -> target loop the excursion window is kept short to avoid; fixing the
window confusion reopened it through the other door. It walked 128 -> 256 ->
384 bars and stopped at q90, the widest rung there is, with every rung
reading 39-48% against a 20% floor. A floor nothing fails selects nothing.
Rungs whose required horizon exceeds BARRIER_HORIZON_MAX are now rejected -
the same rule ReportGeometryExpectancyScan already applied. It was printing
the shipped pair as CLAMPED and disqualified ('h384!') two lines under the
deriver that chose it: two subsystems, one geometry, opposite verdicts.
3. THE RUNG SNAP DESTROYED THE RATIO IT WAS COMPARING. Both legs snapped
independently to the coarse first-passage grid, re-rating each candidate:
q90 4.86/9.71 -> 5.00/10.00 (2.00), q85 4.07/8.14 -> 5.00/10.00 (IDENTICAL
measurement), q75 3.07/6.13 -> 4.00/6.50 (1.63 - a nearer target). So the
ladder compared win shares taken at ratios from 1.63 to 2.17 and read the
differences as scale. It is why the reach column came out non-monotone in
width (q75 48.5% above q90 42.9%). The stop now snaps to its nearest rung in
log space and the target follows the ratio off it; the pair actually measured
is returned and logged, so a collision reads as a collision.
Also: LadderWinShare guarded against the conditional (fractal) geometry path,
which fills n from m_fracLegCount while leaving idxList empty - a latent
out-of-bounds on a currently-dead path.
New log lines: mean label lifespan and effective n on the label-cache line, the
required-vs-available horizon per rung, and the grid pair the reconciliation
actually measured (its tolerance now scales with the grid skew instead of a flat
5pp).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 01:12:05 -04:00
|
|
|
BARRIER_HORIZON_MAX, 100.0 * chosenQ, chosenReach));
|
fix(geometry): the target was small BY CONSTRUCTION - ratio is now policy, scale is measured, ladder ceiling removed
The derivation read the stop from q75 of ADVERSE travel and the target from q50
of FAVOURABLE travel. Over one horizon those distributions are broadly the same
shape, so q75 > q50 MECHANICALLY - the target came out smaller than the stop no
matter what the market did. SP500 H4 shipped stop 3.07 / target 1.70: a 0.55:1
payoff needing 64.3%. That was never a measurement, it was two mismatched
constants.
The reachability line printed beside it - "target on 50.0% of bars, stop on
25.0%" - is exactly 1-q50 and 1-q75. Tautological. It cannot disconfirm
anything, and it read as validation.
WIDTH AND RATIO ARE INDEPENDENT AND ONLY ONE PAYS. EV = edge x width;
ratio is EV-neutral (a driftless walk reaches +m before -k with probability
k/(k+m), which IS break-even). Width is what buys cost efficiency: the spread
is a fixed 0.047*ATR here, so the shipped 4.77*ATR width paid it 21 times per
unit of travel. So:
RATIO = policy. BARRIER_TARGET_RR = 2.0 (user's 1:2). Break-even 33.3%.
SCALE = measured. The stop quantile is chosen from a ladder, WIDEST FIRST,
taking the first rung whose implied 2x target is still reached often
enough to be a trainable class.
That last clause is the difference from the min-reward:risk raise removed in
2026-08-09, which forced target = 2 x stop with NO reachability test, landed on
6.66*ATR reachable on 3.3% of bars, and trained the model to predict something
that essentially never happened. Same ratio; the scale now retreats until the
data says the target is attainable. Every rung is logged.
LADDER CEILING REMOVED. BARRIER_LADDER stopped at 5.00 and the expectancy scan's
"best resolvable pair on width alone" came back as stop 5.05 / target 4.95 - it
pinned to the top rung. A recommendation landing exactly on the edge of its own
search space is a boundary, not a finding: it cannot tell "5 ATR is optimal"
from "5 ATR is all we allowed". Extended to 20*ATR (8 -> 14 rungs). Nothing else
needs editing - every consumer is parameterised by BARRIER_LADDER_COUNT - and
the horizon constraints (decided >= 60%, reachability floor) now bind instead of
a constant.
THE SCAN COULD NOT SEE THE SHIPPED GEOMETRY. ReportBarrierGeometryScan looked
the configured pair up in its integer grid, and DeriveBarrierGeometry produces
CONTINUOUS multiples (3.07/1.70) that can never equal a grid point - so
cfgExcess stayed at its -1.0 sentinel and the report printed "configured 3:2
scores -1.00000", which reads as a catastrophic score and actually means "never
evaluated". Worse, the grid skipped target<stop entirely because it "inverts the
trade's whole premise" - while the derivation was shipping exactly that. The
incumbent is now always scored as a peer (never crowned; it is already in force
and is not an enum pairing the scan could adopt).
BREAK-EVEN NOW INCLUDES THE SPREAD. Every report quoted the frictionless
SL/(SL+TP). On SP500 H4 that read 64.3% while the MEASURED zero-skill rate was
62.1% - a 2.2pp gap that IS the cost, and that made every model look 2.2pp
better than it was. CostAdjustedBreakEvenPct() prices a win at (TP - spread) and
a loss at (SL + spread), matching the expectancy scan's convention exactly so
the two reports cannot disagree.
It also feeds FitDirConfThreshold, which is the correctness half: the operating
point subtracts break-even from precision, so the frictionless figure made every
candidate threshold look better by the width of the spread - 2.2pp against a
measured edge of 2.3pp, i.e. very nearly all of it.
Era line now carries both: "break-even 64.3% frictionless, 66.6% AFTER SPREAD".
Forces a full relabel and retrain. Requested.
NOT COMPILED - user compiles.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 00:34:32 -04:00
|
|
|
}
|
feat(measurement): fix zero-skill denominator, publish the deploy bar, measure lifespan per rung, add a MEASURE scale objective
The last run could not have demonstrated an edge either way, and nothing in the
log said so. Four changes so it does.
1. THE ZERO-SKILL LINE DIVIDED BY THE WRONG DENOMINATOR. m_oosWinLongTotal resets
every era; m_oosSamples only resets on a full model reset. So 'always-long %'
decayed as ~1/era: a run whose true rate is 37% printed 1.2% at era 33 and
0.0% at era 2219. This is the SAME bug already found and fixed for
logBuyPredPct thirty lines above ('era-15 Buy:2% that was really ~30%'), left
in the one line whose whole job is to be the reference every other number is
read against. Correct at era 1, wrong everywhere after - including the '62%
zero-skill' figure in the 2026-08-16 notes. Now per-era, and always-short is
finally readable.
2. THE DEPLOY GATE STATES ITS OWN BAR. 'edge -1pp' era after era cannot separate
'short by a hair' from 'short by an amount no strategy could cover'. The era
line now prints the required win rate, the SE, the effective n and the
lifespan it was deflated by; above 100% it says UNREACHABLE. At 4,738 OOS bars
and L=75.6 there are ~63 independent observations, putting the bar near 66% at
typical coverage.
3. LIFESPAN MEASURED PER RUNG. The first-passage cache already stores touch ages
at every ladder level, so each candidate geometry's resolution time is
readable without training on it - L-vs-width becomes a measurement across the
whole ladder in ONE run rather than a second chart. Each rung reports L,
n_eff, min provable edge and min provable EV.
4. SCALE OBJECTIVE IS PHASE-AWARE, defaulting to MEASURE. Width and detectability
are opposed: labels overlap by L, L grows like m*k = width^2 at fixed ratio,
so min provable EV ~ width^2 while the cost saving from width is only linear.
Doubling width quadruples the smallest EV you can prove. DEPLOY (widest that
clears reachability) is right once an edge is known; MEASURE (narrowest that
keeps round-trip spread under BARRIER_MAX_COST_FRACTION_PCT) is right while it
still has to be shown. The direction does not depend on the exponent, and
item 3 makes the exponent checkable.
Fixed in review: m_lastRungLifespan is cleared on every LadderWinShare entry or a
rejected rung reports the previous rung's lifespan as its own; per-rung
detectability is labelled IS-sample based (the deriver may not see the holdout),
so absolute figures are optimistic while the ranking is unaffected.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 10:37:05 -04:00
|
|
|
Print(ID + StringFormat(": barrier SCALE ladder - objective %s (ratio fixed at 1:%.1f by policy). Every "
|
|
|
|
|
"rung must clear %.0f%% reachability, resolve inside the %d-bar horizon "
|
|
|
|
|
"ceiling, and keep the round-trip spread under %.1f%% of its width; of those, "
|
|
|
|
|
"%s. Width is EV per trade, narrowness is EV you can PROVE - min provable EV "
|
|
|
|
|
"grows as width SQUARED because the label's lifespan does, so the two ends of "
|
|
|
|
|
"this ladder are opposed and only one is right per phase. - %s | chose q%.0f, "
|
|
|
|
|
"target reached on %.1f%% of bars; needs %d bars, granted %d (ceiling %d, then "
|
|
|
|
|
"snapped DOWN - the shortfall shows up as the timeout share on the next "
|
|
|
|
|
"label-cache line)",
|
|
|
|
|
(BARRIER_SCALE_OBJECTIVE == BARRIER_SCALE_DEPLOY
|
|
|
|
|
? "DEPLOY (maximise EV per trade)"
|
|
|
|
|
: "MEASURE (maximise detectability - the edge is not proven yet)"),
|
|
|
|
|
BARRIER_TARGET_RR, BARRIER_MIN_TP_REACH_PCT, BARRIER_HORIZON_MAX,
|
|
|
|
|
BARRIER_MAX_COST_FRACTION_PCT,
|
|
|
|
|
(BARRIER_SCALE_OBJECTIVE == BARRIER_SCALE_DEPLOY
|
|
|
|
|
? "the WIDEST wins" : "the NARROWEST wins"),
|
|
|
|
|
rungRows, 100.0 * chosenQ, chosenReach, RequiredHorizonBars(slRaw, tpRaw),
|
fix(labels): correct EffectiveSampleSize clamp order, share the horizon ladder, retract a false justification
Self-review of 1540ba8 against the FULL 6,930-era log rather than the first
three minutes of it. Three corrections.
1. EffectiveSampleSize() clamped in the wrong order. MathMax(2, MathMin(eff,
rawN)) returns 2 when rawN is 1 - an effective sample LARGER than the raw
one, shrinking the SE in exactly the direction the function exists to
prevent. Floor first, cap at rawN last.
2. The horizon cap rejected on the CEILING only, and said so as though that
made the label untruncated. It does not: the horizon ladder also snaps DOWN,
so a pair needing 317 bars is granted 256 and is silently truncated without
ever being flagged CLAMPED. Added SnapHorizonToLadder() / GrantedHorizonBars()
and the scale ladder now reports "needs N gets M" per rung. Rejection stays on
the ceiling alone - matching ReportGeometryExpectancyScan's '!' exactly, which
was the point - because rejecting on the snap-down would select rungs for
landing just above a ladder point rather than for anything about the market.
ComputeBarrierHorizonBars' private copy of the ladder is gone; there is now
one copy, which is the whole reason RequiredHorizonBars was factored out.
3. RETRACTED THE JUSTIFICATION IN 1540ba8's COMMENTS. That commit claimed the
overlap correction was needed because the operating point's null-of-the-
maximum gate fired on 47/73 Perceptron eras (64%) where a family-wise test
should fire on ~5%. Those 73 fits were the first three minutes of a
six-and-a-half-hour run. Over the full run:
PAI 47/3214 = 1.5% HYB 30/1200 = 2.5%
CONV 4/63 = 6.3% LSTM 75/915 = 8.2%
All at or below the null. The gate from 7414570 is working as designed and
PAI's 47 clears were a cold-start transient never repeated in 3,141 later
fits; its threshold over the run's second half has sd 0.01. The overlap
correction is still right - sqrt(p(1-p)/n) on overlapping labels is the wrong
formula - but it fixes no observed failure, and it costs nothing today
because no model is near the deploy line.
WHAT THE FULL RUN DOES CONFIRM, unchanged: the geometry ran away exactly as
described (2.00/6.00 h128 -> 3.49/6.99 h256 -> 4.86/9.71 h384, three passes,
stopping at q90 because the quantile ladder ended), the label stayed long-skewed
at Buy 42.9% / Sell 22.6%, and no checkpoint on any of the four models ever
cleared the deployability floor. Pooled declustered win rates: PAI 31.70%,
HYB 31.02%, LSTM 32.69%, CONV 31.57% - every one 4-6pp below the 37% always-long
chance rate and 1-2.7pp below the 33.7% cost-adjusted break-even.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 07:30:19 -04:00
|
|
|
GrantedHorizonBars(slRaw, tpRaw), BARRIER_HORIZON_MAX));
|
2026-08-09 14:51:59 -04:00
|
|
|
//--- Same floor a real order gets, so the stop used for labelling is the stop that can actually be
|
|
|
|
|
//--- placed. This is the ONLY adjustment either leg receives - both multiples are otherwise read
|
|
|
|
|
//--- straight off the measured distributions.
|
feat: derive the ATR multiples from measured excursions - no hardcoded geometry
The barrier was still two constants. SL_Mode/TP_Mode left the Inputs tab in
3482b6c, but the fallback was a hardcoded 2:6 and the geometry scan only ever
chose from a hardcoded grid {2,3} x {2,3,4,6,8,10}. Picking the least-bad of
eleven guesses is not deriving anything.
WHY THE SCAN WAS THE WRONG INSTRUMENT, now measurable rather than argued. It
ranks pairings by how predictable their OUTCOME is - a question about direction.
The excursion test (2c78f3b) ran on SP500 H1 and direction is the one thing
absent: ASYMMETRY p=0.0846, against RANGE/UP/DOWN all at p=0.0050, with RANGE
scoring 0.01345 vs a 0.00343 null - 4x, where the barrier label sits at 1.01x.
Hence the scan failing its own gate on every run, and its "winner" wandering
2:8 -> 3:8 -> 2:8 -> 2:4 across four runs of the same data. Excursion SIZE is
strongly measurable, so derive the geometry from that instead.
stop = q25 of measured ADVERSE travel (ordinary noise does not reach it)
target = q50 of measured FAVOURABLE travel (reached ~half the time, by
construction, inside the horizon)
Continuous, in ATR units, superseding the enum multiples. Reachability ("target
on X% of bars, stop on Y%") and the implied break-even are printed so the choice
is auditable rather than trusted.
FIXED-POINT ITERATION, not one-shot. ComputeBarrierHorizonBars scales the
horizon with the target (first-passage time grows with the band) and the
excursions are measured OVER the horizon, so target -> horizon -> excursions ->
target is a real loop - deriving once sizes the target from travel measured
under the PREVIOUS horizon. Re-measures until the multiples move <5%, capped at
3 passes, and says so if it does not settle.
Does NOT create expectancy, and the log says as much: chance precision equals
break-even at every geometry (m/(m+k) on both sides). It buys a target the
market reaches and a stop that survives noise. Where Min_Risk_Reward_Ratio
forces a target the market rarely reaches, it WARNS rather than overriding -
the ratio is the user's risk policy, so the honest move is to state its cost.
That is the collision that once rejected 100% of setups.
Pinned in the .cfg as doubles appended AFTER this morning's two ints, so .cfg
files written earlier today still load (their length guard finds no doubles) and
a model that carries them was trained on them and never re-derives.
Also fixes a message from e5ceed6 that claimed "this model resumed from disk"
unconditionally - it printed above a "seeding era 0" line on a brand-new model,
because the branch fires whenever the cache is not built, which is equally true
before a fresh model's first prebuild. A diagnostic that misreports its own
trigger is worse than one that says nothing: it gets quoted back as evidence.
FORCES A FULL RETRAIN (labels change).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-07 12:06:25 -04:00
|
|
|
if(slRaw < MIN_SL_ATR_MULTIPLIER)
|
|
|
|
|
slRaw = MIN_SL_ATR_MULTIPLIER;
|
2026-08-09 14:51:59 -04:00
|
|
|
//--- The minimum-reward:risk raise that used to sit here is GONE (2026-08-09). It overrode a measured
|
|
|
|
|
//--- target with an arithmetic one - twice whatever the stop happened to be - and on SP500 H1 that
|
|
|
|
|
//--- pushed the target from the q50 of favourable travel out to 6.66*ATR, reachable on 3.3% of bars.
|
|
|
|
|
//--- The model was then trained to predict an outcome that essentially never happens. A measured
|
|
|
|
|
//--- target has to stay measured; see Variables\Inputs.mqh for why the ratio bought nothing in
|
|
|
|
|
//--- exchange (a reward:risk floor moves payoff and hit rate together at a fixed break-even, it does
|
|
|
|
|
//--- not create expectancy) and cost two separate outages.
|
fix(geometry): the reachability floor measured the WRONG WINDOW - my bug from bc57aca, and it cost real width
RECONCILED: the derivation reported "target reached on 17.7% of bars" while the
label cache reported Buy on 35.9%. Nothing was broken. They measure different
windows, and both are correct:
EXCURSION window ~12 bars (the SWING MEDIAN) - what m_excUpCache accumulates
over. Deliberately short: sizing a barrier off travel
measured over a horizon that itself scales with the barrier
is circular, and it ran away to 14-31*ATR on EURUSD/USDCAD
in 2026-08-07. That guard is correct and stays.
BARRIER horizon 64 bars - what the LABEL walk and the first-passage ladder
run over, and how long the EA actually holds the trade.
So `up >= target` is a 12-bar question and `label == Buy` is a 64-bar one, and
the second can freely exceed the first. TripleBarrierLabel gates the excursion
accumulation on `idx - t <= excWindow` while the barrier walk and the ladder run
the full horizon - the split is explicit and intentional.
THE BUG IS MINE. bc57aca's scale ladder tested reachability with `up[i] >= tp`,
i.e. it asked the 12-bar question about a 64-bar trade. That understates
reachability by ~2x, which is why EVERY wide rung was rejected and the geometry
fell back to the tightest rung at 1.61/3.21. The data supported considerably
wider; the test was just asking the wrong question.
FIX: LadderWinShare() reads the answer off the first-passage ladder - target
touched strictly before the stop, over the full horizon, tie to the stop. That
is the identical question the label walk asks, so the ladder share and the Buy
rate should now agree to within rung discretisation. Both legs snap to the
SMALLEST rung at or above the requested multiple (harder target, harder stop) so
the floor stays conservative.
Expect the scale ladder to select a WIDER rung on the next relabel. On this
data the excursion test read 17.7% at q50 where the true full-horizon share is
35.9%, so rungs that scored 8.1% and 2.8% were likely well above the floor.
ALSO:
- Window reconciliation now PRINTED every derivation: excursion travel share,
ladder win share, and the label cache's Buy share side by side, with the
ladder-vs-label gap flagged if it exceeds rung discretisation. Those two must
agree; if they ever stop agreeing, one of them is wrong and the line says so.
- Renamed tpReach/slReach -> tpTravel/slTravel and relabelled the log line. They
describe the EXCURSION window and are near-tautological there (a q50 stop is
exceeded by ~50% of bars); calling them "reached within the horizon" is what
made the two quantities look like one.
- BARRIER_MIN_TP_REACH_PCT is now BARRIER_MIN_REACH_FRACTION_OF_BE (0.60) x
break-even instead of a hardcoded 20.0. Break-even for 1:RR is 100/(1+RR), so
the absolute floor silently tightened as RR rose - 0.60x at RR=2 but 0.80x at
RR=3, penalising the user for asking for a bigger target. Evaluates to exactly
20.0% at the shipped RR=2, so this is a no-op today and correct if the knob
moves.
NOT COMPILED - user compiles.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 00:51:08 -04:00
|
|
|
//--- TRAVEL SHARES OVER THE EXCURSION WINDOW - what fraction of bars moved this far within the ~swing
|
|
|
|
|
//--- median. These describe the DISTRIBUTION THE MULTIPLES WERE READ OFF, which is the only thing they
|
|
|
|
|
//--- can honestly describe, and they are near-tautological by construction (a q50 stop is exceeded by
|
|
|
|
|
//--- ~50% of bars). They are NOT reachability over the horizon the trade is held for - that is
|
|
|
|
|
//--- chosenReach, measured on the first-passage ladder, and mixing the two is what produced the
|
|
|
|
|
//--- 17.7%-vs-35.9% confusion. Kept and RENAMED rather than deleted: they are the sanity check that
|
|
|
|
|
//--- the quantile read did what it claimed.
|
|
|
|
|
int travelTp = 0, travelSl = 0;
|
feat: derive the ATR multiples from measured excursions - no hardcoded geometry
The barrier was still two constants. SL_Mode/TP_Mode left the Inputs tab in
3482b6c, but the fallback was a hardcoded 2:6 and the geometry scan only ever
chose from a hardcoded grid {2,3} x {2,3,4,6,8,10}. Picking the least-bad of
eleven guesses is not deriving anything.
WHY THE SCAN WAS THE WRONG INSTRUMENT, now measurable rather than argued. It
ranks pairings by how predictable their OUTCOME is - a question about direction.
The excursion test (2c78f3b) ran on SP500 H1 and direction is the one thing
absent: ASYMMETRY p=0.0846, against RANGE/UP/DOWN all at p=0.0050, with RANGE
scoring 0.01345 vs a 0.00343 null - 4x, where the barrier label sits at 1.01x.
Hence the scan failing its own gate on every run, and its "winner" wandering
2:8 -> 3:8 -> 2:8 -> 2:4 across four runs of the same data. Excursion SIZE is
strongly measurable, so derive the geometry from that instead.
stop = q25 of measured ADVERSE travel (ordinary noise does not reach it)
target = q50 of measured FAVOURABLE travel (reached ~half the time, by
construction, inside the horizon)
Continuous, in ATR units, superseding the enum multiples. Reachability ("target
on X% of bars, stop on Y%") and the implied break-even are printed so the choice
is auditable rather than trusted.
FIXED-POINT ITERATION, not one-shot. ComputeBarrierHorizonBars scales the
horizon with the target (first-passage time grows with the band) and the
excursions are measured OVER the horizon, so target -> horizon -> excursions ->
target is a real loop - deriving once sizes the target from travel measured
under the PREVIOUS horizon. Re-measures until the multiples move <5%, capped at
3 passes, and says so if it does not settle.
Does NOT create expectancy, and the log says as much: chance precision equals
break-even at every geometry (m/(m+k) on both sides). It buys a target the
market reaches and a stop that survives noise. Where Min_Risk_Reward_Ratio
forces a target the market rarely reaches, it WARNS rather than overriding -
the ratio is the user's risk policy, so the honest move is to state its cost.
That is the collision that once rejected 100% of setups.
Pinned in the .cfg as doubles appended AFTER this morning's two ints, so .cfg
files written earlier today still load (their length guard finds no doubles) and
a model that carries them was trained on them and never re-derives.
Also fixes a message from e5ceed6 that claimed "this model resumed from disk"
unconditionally - it printed above a "seeding era 0" line on a brand-new model,
because the branch fires whenever the cache is not built, which is equally true
before a fresh model's first prebuild. A diagnostic that misreports its own
trigger is worse than one that says nothing: it gets quoted back as evidence.
FORCES A FULL RETRAIN (labels change).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-07 12:06:25 -04:00
|
|
|
for(int i = 0; i < n; i++)
|
|
|
|
|
{
|
|
|
|
|
if(up[i] >= tpRaw)
|
fix(geometry): the reachability floor measured the WRONG WINDOW - my bug from bc57aca, and it cost real width
RECONCILED: the derivation reported "target reached on 17.7% of bars" while the
label cache reported Buy on 35.9%. Nothing was broken. They measure different
windows, and both are correct:
EXCURSION window ~12 bars (the SWING MEDIAN) - what m_excUpCache accumulates
over. Deliberately short: sizing a barrier off travel
measured over a horizon that itself scales with the barrier
is circular, and it ran away to 14-31*ATR on EURUSD/USDCAD
in 2026-08-07. That guard is correct and stays.
BARRIER horizon 64 bars - what the LABEL walk and the first-passage ladder
run over, and how long the EA actually holds the trade.
So `up >= target` is a 12-bar question and `label == Buy` is a 64-bar one, and
the second can freely exceed the first. TripleBarrierLabel gates the excursion
accumulation on `idx - t <= excWindow` while the barrier walk and the ladder run
the full horizon - the split is explicit and intentional.
THE BUG IS MINE. bc57aca's scale ladder tested reachability with `up[i] >= tp`,
i.e. it asked the 12-bar question about a 64-bar trade. That understates
reachability by ~2x, which is why EVERY wide rung was rejected and the geometry
fell back to the tightest rung at 1.61/3.21. The data supported considerably
wider; the test was just asking the wrong question.
FIX: LadderWinShare() reads the answer off the first-passage ladder - target
touched strictly before the stop, over the full horizon, tie to the stop. That
is the identical question the label walk asks, so the ladder share and the Buy
rate should now agree to within rung discretisation. Both legs snap to the
SMALLEST rung at or above the requested multiple (harder target, harder stop) so
the floor stays conservative.
Expect the scale ladder to select a WIDER rung on the next relabel. On this
data the excursion test read 17.7% at q50 where the true full-horizon share is
35.9%, so rungs that scored 8.1% and 2.8% were likely well above the floor.
ALSO:
- Window reconciliation now PRINTED every derivation: excursion travel share,
ladder win share, and the label cache's Buy share side by side, with the
ladder-vs-label gap flagged if it exceeds rung discretisation. Those two must
agree; if they ever stop agreeing, one of them is wrong and the line says so.
- Renamed tpReach/slReach -> tpTravel/slTravel and relabelled the log line. They
describe the EXCURSION window and are near-tautological there (a q50 stop is
exceeded by ~50% of bars); calling them "reached within the horizon" is what
made the two quantities look like one.
- BARRIER_MIN_TP_REACH_PCT is now BARRIER_MIN_REACH_FRACTION_OF_BE (0.60) x
break-even instead of a hardcoded 20.0. Break-even for 1:RR is 100/(1+RR), so
the absolute floor silently tightened as RR rose - 0.60x at RR=2 but 0.80x at
RR=3, penalising the user for asking for a bigger target. Evaluates to exactly
20.0% at the shipped RR=2, so this is a no-op today and correct if the knob
moves.
NOT COMPILED - user compiles.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 00:51:08 -04:00
|
|
|
travelTp++;
|
feat: derive the ATR multiples from measured excursions - no hardcoded geometry
The barrier was still two constants. SL_Mode/TP_Mode left the Inputs tab in
3482b6c, but the fallback was a hardcoded 2:6 and the geometry scan only ever
chose from a hardcoded grid {2,3} x {2,3,4,6,8,10}. Picking the least-bad of
eleven guesses is not deriving anything.
WHY THE SCAN WAS THE WRONG INSTRUMENT, now measurable rather than argued. It
ranks pairings by how predictable their OUTCOME is - a question about direction.
The excursion test (2c78f3b) ran on SP500 H1 and direction is the one thing
absent: ASYMMETRY p=0.0846, against RANGE/UP/DOWN all at p=0.0050, with RANGE
scoring 0.01345 vs a 0.00343 null - 4x, where the barrier label sits at 1.01x.
Hence the scan failing its own gate on every run, and its "winner" wandering
2:8 -> 3:8 -> 2:8 -> 2:4 across four runs of the same data. Excursion SIZE is
strongly measurable, so derive the geometry from that instead.
stop = q25 of measured ADVERSE travel (ordinary noise does not reach it)
target = q50 of measured FAVOURABLE travel (reached ~half the time, by
construction, inside the horizon)
Continuous, in ATR units, superseding the enum multiples. Reachability ("target
on X% of bars, stop on Y%") and the implied break-even are printed so the choice
is auditable rather than trusted.
FIXED-POINT ITERATION, not one-shot. ComputeBarrierHorizonBars scales the
horizon with the target (first-passage time grows with the band) and the
excursions are measured OVER the horizon, so target -> horizon -> excursions ->
target is a real loop - deriving once sizes the target from travel measured
under the PREVIOUS horizon. Re-measures until the multiples move <5%, capped at
3 passes, and says so if it does not settle.
Does NOT create expectancy, and the log says as much: chance precision equals
break-even at every geometry (m/(m+k) on both sides). It buys a target the
market reaches and a stop that survives noise. Where Min_Risk_Reward_Ratio
forces a target the market rarely reaches, it WARNS rather than overriding -
the ratio is the user's risk policy, so the honest move is to state its cost.
That is the collision that once rejected 100% of setups.
Pinned in the .cfg as doubles appended AFTER this morning's two ints, so .cfg
files written earlier today still load (their length guard finds no doubles) and
a model that carries them was trained on them and never re-derives.
Also fixes a message from e5ceed6 that claimed "this model resumed from disk"
unconditionally - it printed above a "seeding era 0" line on a brand-new model,
because the branch fires whenever the cache is not built, which is equally true
before a fresh model's first prebuild. A diagnostic that misreports its own
trigger is worse than one that says nothing: it gets quoted back as evidence.
FORCES A FULL RETRAIN (labels change).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-07 12:06:25 -04:00
|
|
|
if(dn[i] >= slRaw)
|
fix(geometry): the reachability floor measured the WRONG WINDOW - my bug from bc57aca, and it cost real width
RECONCILED: the derivation reported "target reached on 17.7% of bars" while the
label cache reported Buy on 35.9%. Nothing was broken. They measure different
windows, and both are correct:
EXCURSION window ~12 bars (the SWING MEDIAN) - what m_excUpCache accumulates
over. Deliberately short: sizing a barrier off travel
measured over a horizon that itself scales with the barrier
is circular, and it ran away to 14-31*ATR on EURUSD/USDCAD
in 2026-08-07. That guard is correct and stays.
BARRIER horizon 64 bars - what the LABEL walk and the first-passage ladder
run over, and how long the EA actually holds the trade.
So `up >= target` is a 12-bar question and `label == Buy` is a 64-bar one, and
the second can freely exceed the first. TripleBarrierLabel gates the excursion
accumulation on `idx - t <= excWindow` while the barrier walk and the ladder run
the full horizon - the split is explicit and intentional.
THE BUG IS MINE. bc57aca's scale ladder tested reachability with `up[i] >= tp`,
i.e. it asked the 12-bar question about a 64-bar trade. That understates
reachability by ~2x, which is why EVERY wide rung was rejected and the geometry
fell back to the tightest rung at 1.61/3.21. The data supported considerably
wider; the test was just asking the wrong question.
FIX: LadderWinShare() reads the answer off the first-passage ladder - target
touched strictly before the stop, over the full horizon, tie to the stop. That
is the identical question the label walk asks, so the ladder share and the Buy
rate should now agree to within rung discretisation. Both legs snap to the
SMALLEST rung at or above the requested multiple (harder target, harder stop) so
the floor stays conservative.
Expect the scale ladder to select a WIDER rung on the next relabel. On this
data the excursion test read 17.7% at q50 where the true full-horizon share is
35.9%, so rungs that scored 8.1% and 2.8% were likely well above the floor.
ALSO:
- Window reconciliation now PRINTED every derivation: excursion travel share,
ladder win share, and the label cache's Buy share side by side, with the
ladder-vs-label gap flagged if it exceeds rung discretisation. Those two must
agree; if they ever stop agreeing, one of them is wrong and the line says so.
- Renamed tpReach/slReach -> tpTravel/slTravel and relabelled the log line. They
describe the EXCURSION window and are near-tautological there (a q50 stop is
exceeded by ~50% of bars); calling them "reached within the horizon" is what
made the two quantities look like one.
- BARRIER_MIN_TP_REACH_PCT is now BARRIER_MIN_REACH_FRACTION_OF_BE (0.60) x
break-even instead of a hardcoded 20.0. Break-even for 1:RR is 100/(1+RR), so
the absolute floor silently tightened as RR rose - 0.60x at RR=2 but 0.80x at
RR=3, penalising the user for asking for a bigger target. Evaluates to exactly
20.0% at the shipped RR=2, so this is a no-op today and correct if the knob
moves.
NOT COMPILED - user compiles.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 00:51:08 -04:00
|
|
|
travelSl++;
|
feat: derive the ATR multiples from measured excursions - no hardcoded geometry
The barrier was still two constants. SL_Mode/TP_Mode left the Inputs tab in
3482b6c, but the fallback was a hardcoded 2:6 and the geometry scan only ever
chose from a hardcoded grid {2,3} x {2,3,4,6,8,10}. Picking the least-bad of
eleven guesses is not deriving anything.
WHY THE SCAN WAS THE WRONG INSTRUMENT, now measurable rather than argued. It
ranks pairings by how predictable their OUTCOME is - a question about direction.
The excursion test (2c78f3b) ran on SP500 H1 and direction is the one thing
absent: ASYMMETRY p=0.0846, against RANGE/UP/DOWN all at p=0.0050, with RANGE
scoring 0.01345 vs a 0.00343 null - 4x, where the barrier label sits at 1.01x.
Hence the scan failing its own gate on every run, and its "winner" wandering
2:8 -> 3:8 -> 2:8 -> 2:4 across four runs of the same data. Excursion SIZE is
strongly measurable, so derive the geometry from that instead.
stop = q25 of measured ADVERSE travel (ordinary noise does not reach it)
target = q50 of measured FAVOURABLE travel (reached ~half the time, by
construction, inside the horizon)
Continuous, in ATR units, superseding the enum multiples. Reachability ("target
on X% of bars, stop on Y%") and the implied break-even are printed so the choice
is auditable rather than trusted.
FIXED-POINT ITERATION, not one-shot. ComputeBarrierHorizonBars scales the
horizon with the target (first-passage time grows with the band) and the
excursions are measured OVER the horizon, so target -> horizon -> excursions ->
target is a real loop - deriving once sizes the target from travel measured
under the PREVIOUS horizon. Re-measures until the multiples move <5%, capped at
3 passes, and says so if it does not settle.
Does NOT create expectancy, and the log says as much: chance precision equals
break-even at every geometry (m/(m+k) on both sides). It buys a target the
market reaches and a stop that survives noise. Where Min_Risk_Reward_Ratio
forces a target the market rarely reaches, it WARNS rather than overriding -
the ratio is the user's risk policy, so the honest move is to state its cost.
That is the collision that once rejected 100% of setups.
Pinned in the .cfg as doubles appended AFTER this morning's two ints, so .cfg
files written earlier today still load (their length guard finds no doubles) and
a model that carries them was trained on them and never re-derives.
Also fixes a message from e5ceed6 that claimed "this model resumed from disk"
unconditionally - it printed above a "seeding era 0" line on a brand-new model,
because the branch fires whenever the cache is not built, which is equally true
before a fresh model's first prebuild. A diagnostic that misreports its own
trigger is worse than one that says nothing: it gets quoted back as evidence.
FORCES A FULL RETRAIN (labels change).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-07 12:06:25 -04:00
|
|
|
}
|
fix(geometry): the reachability floor measured the WRONG WINDOW - my bug from bc57aca, and it cost real width
RECONCILED: the derivation reported "target reached on 17.7% of bars" while the
label cache reported Buy on 35.9%. Nothing was broken. They measure different
windows, and both are correct:
EXCURSION window ~12 bars (the SWING MEDIAN) - what m_excUpCache accumulates
over. Deliberately short: sizing a barrier off travel
measured over a horizon that itself scales with the barrier
is circular, and it ran away to 14-31*ATR on EURUSD/USDCAD
in 2026-08-07. That guard is correct and stays.
BARRIER horizon 64 bars - what the LABEL walk and the first-passage ladder
run over, and how long the EA actually holds the trade.
So `up >= target` is a 12-bar question and `label == Buy` is a 64-bar one, and
the second can freely exceed the first. TripleBarrierLabel gates the excursion
accumulation on `idx - t <= excWindow` while the barrier walk and the ladder run
the full horizon - the split is explicit and intentional.
THE BUG IS MINE. bc57aca's scale ladder tested reachability with `up[i] >= tp`,
i.e. it asked the 12-bar question about a 64-bar trade. That understates
reachability by ~2x, which is why EVERY wide rung was rejected and the geometry
fell back to the tightest rung at 1.61/3.21. The data supported considerably
wider; the test was just asking the wrong question.
FIX: LadderWinShare() reads the answer off the first-passage ladder - target
touched strictly before the stop, over the full horizon, tie to the stop. That
is the identical question the label walk asks, so the ladder share and the Buy
rate should now agree to within rung discretisation. Both legs snap to the
SMALLEST rung at or above the requested multiple (harder target, harder stop) so
the floor stays conservative.
Expect the scale ladder to select a WIDER rung on the next relabel. On this
data the excursion test read 17.7% at q50 where the true full-horizon share is
35.9%, so rungs that scored 8.1% and 2.8% were likely well above the floor.
ALSO:
- Window reconciliation now PRINTED every derivation: excursion travel share,
ladder win share, and the label cache's Buy share side by side, with the
ladder-vs-label gap flagged if it exceeds rung discretisation. Those two must
agree; if they ever stop agreeing, one of them is wrong and the line says so.
- Renamed tpReach/slReach -> tpTravel/slTravel and relabelled the log line. They
describe the EXCURSION window and are near-tautological there (a q50 stop is
exceeded by ~50% of bars); calling them "reached within the horizon" is what
made the two quantities look like one.
- BARRIER_MIN_TP_REACH_PCT is now BARRIER_MIN_REACH_FRACTION_OF_BE (0.60) x
break-even instead of a hardcoded 20.0. Break-even for 1:RR is 100/(1+RR), so
the absolute floor silently tightened as RR rose - 0.60x at RR=2 but 0.80x at
RR=3, penalising the user for asking for a bigger target. Evaluates to exactly
20.0% at the shipped RR=2, so this is a no-op today and correct if the knob
moves.
NOT COMPILED - user compiles.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 00:51:08 -04:00
|
|
|
double tpTravel = 100.0 * travelTp / n;
|
|
|
|
|
double slTravel = 100.0 * travelSl / n;
|
feat: derive the ATR multiples from measured excursions - no hardcoded geometry
The barrier was still two constants. SL_Mode/TP_Mode left the Inputs tab in
3482b6c, but the fallback was a hardcoded 2:6 and the geometry scan only ever
chose from a hardcoded grid {2,3} x {2,3,4,6,8,10}. Picking the least-bad of
eleven guesses is not deriving anything.
WHY THE SCAN WAS THE WRONG INSTRUMENT, now measurable rather than argued. It
ranks pairings by how predictable their OUTCOME is - a question about direction.
The excursion test (2c78f3b) ran on SP500 H1 and direction is the one thing
absent: ASYMMETRY p=0.0846, against RANGE/UP/DOWN all at p=0.0050, with RANGE
scoring 0.01345 vs a 0.00343 null - 4x, where the barrier label sits at 1.01x.
Hence the scan failing its own gate on every run, and its "winner" wandering
2:8 -> 3:8 -> 2:8 -> 2:4 across four runs of the same data. Excursion SIZE is
strongly measurable, so derive the geometry from that instead.
stop = q25 of measured ADVERSE travel (ordinary noise does not reach it)
target = q50 of measured FAVOURABLE travel (reached ~half the time, by
construction, inside the horizon)
Continuous, in ATR units, superseding the enum multiples. Reachability ("target
on X% of bars, stop on Y%") and the implied break-even are printed so the choice
is auditable rather than trusted.
FIXED-POINT ITERATION, not one-shot. ComputeBarrierHorizonBars scales the
horizon with the target (first-passage time grows with the band) and the
excursions are measured OVER the horizon, so target -> horizon -> excursions ->
target is a real loop - deriving once sizes the target from travel measured
under the PREVIOUS horizon. Re-measures until the multiples move <5%, capped at
3 passes, and says so if it does not settle.
Does NOT create expectancy, and the log says as much: chance precision equals
break-even at every geometry (m/(m+k) on both sides). It buys a target the
market reaches and a stop that survives noise. Where Min_Risk_Reward_Ratio
forces a target the market rarely reaches, it WARNS rather than overriding -
the ratio is the user's risk policy, so the honest move is to state its cost.
That is the collision that once rejected 100% of setups.
Pinned in the .cfg as doubles appended AFTER this morning's two ints, so .cfg
files written earlier today still load (their length guard finds no doubles) and
a model that carries them was trained on them and never re-derives.
Also fixes a message from e5ceed6 that claimed "this model resumed from disk"
unconditionally - it printed above a "seeding era 0" line on a brand-new model,
because the branch fires whenever the cache is not built, which is equally true
before a fresh model's first prebuild. A diagnostic that misreports its own
trigger is worse than one that says nothing: it gets quoted back as evidence.
FORCES A FULL RETRAIN (labels change).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-07 12:06:25 -04:00
|
|
|
double breakeven = 100.0 * slRaw / (slRaw + tpRaw);
|
|
|
|
|
m_derivedSlMult = slRaw;
|
|
|
|
|
m_derivedTpMult = tpRaw;
|
|
|
|
|
m_geometryDerived = true;
|
fix: live trades now use the geometry the gate certifies; perf: BN kernels
Three changes, one theme: the trade placed, the trade graded, and the trade
computed are now the same trade.
1) GEOMETRY WIRE (correctness, the ranked #1 open issue). The measured barrier
pair reached the LABELS only - OpenParams still placed orders at the enum
geometry (2*ATR/6*ATR), so the deploy gate certified "reaches 1.62*ATR before
3.33*ATR above break-even" about trades the EA never placed. Published via
g_DerivedSlAtrMult/g_DerivedTpAtrMult (ConfidenceBridge, same same-tick
contract as the confidence globals, because OpenParams runs on the root signal
which has no pointer to the AI filter). Two writers: DeriveBarrierGeometry at
era 0, and the .cfg adoption a deployed model takes. Overrides both legs and
both Intelligent modes - the certificate is exact or it is nothing. TP is
ATR-anchored like the label, NOT risk-relative, so a floor-widened stop cannot
reshape the certified target.
2) BATCH NORM RUNS DEVICE-SIDE ON OPENCL. Four kernels in Network.cl -
forward, hidden gradient, gamma/beta accumulate, gamma/beta apply - each a
line-for-line transcription of the host implementation (NormalizeHost /
HiddenGradHost / StepGammaBeta) including every NaN guard, clamp, and the
exact moment-write ordering. The host copies remain the runtime for the DLL
and pure-MQL5 tiers and the reference the kernels must match.
Because this box has no OpenCL platform, the safety story is layered:
- shim validation: kernels compiled as C and driven against a fp64 host
transcription over NaN-poisoned stats, NaN gamma, over-clamp inputs, the
frozen path, both optimizers, 3 batches - ALL PASS, worst normalized diff
0.132 vs tolerance 1.0
- in-situ self-check: each kernel is compared against its host twin ON FIRST
USE on the real device (SelfCheckBn*), covering what the shim cannot - arg
indices and buffer bindings. Any disagreement resyncs from the good copy,
latches all BN kernels off process-wide, and training continues host-side.
A transcription bug costs a warning and some speed, never a poisoned .nnw.
- sync discipline: BatchOptions is now a CBufferDouble with explicit
authority tracking (m_bnDeviceAuthoritative). Checkpoints/saves pull
read-only; restores/loads/resets push; a mid-batch handover drains the
device gamma/beta accumulator into the host arrays so no sample is lost.
3) SMALL FIXES. Apply-kernel build failure now latches the dispatch path at
init (one warning instead of warning + failed Execute). Build tag bumped to
win-scoring-gpu-v1 - first tag change since expectancy-stop-v1 despite five
binary-changing commits.
Both build variants compile 0 errors, 0 warnings.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 17:51:40 -04:00
|
|
|
//--- Publish to the LIVE order path (ConfidenceBridge.mqh). Until 2026-08-09 the derived pair
|
|
|
|
|
//--- reached the labels only, so the gate certified trades at this geometry while OpenParams()
|
|
|
|
|
//--- placed them at the enum geometry - graded on one game, paid on another.
|
|
|
|
|
g_DerivedSlAtrMult = m_derivedSlMult;
|
|
|
|
|
g_DerivedTpAtrMult = m_derivedTpMult;
|
2026-08-15 19:00:40 -04:00
|
|
|
if(conditional)
|
|
|
|
|
Print(ID + StringFormat(": geometry source - CONDITIONAL on the fractal label: MFE/MAE measured "
|
|
|
|
|
"from each labeled bar's close over the leg to its NEXT fractal extreme "
|
|
|
|
|
"(%d IS legs, entry-anchored), not over every bar. The stop/target below "
|
|
|
|
|
"are sized for the bars the model actually trades.", n));
|
fix: live trades now use the geometry the gate certifies; perf: BN kernels
Three changes, one theme: the trade placed, the trade graded, and the trade
computed are now the same trade.
1) GEOMETRY WIRE (correctness, the ranked #1 open issue). The measured barrier
pair reached the LABELS only - OpenParams still placed orders at the enum
geometry (2*ATR/6*ATR), so the deploy gate certified "reaches 1.62*ATR before
3.33*ATR above break-even" about trades the EA never placed. Published via
g_DerivedSlAtrMult/g_DerivedTpAtrMult (ConfidenceBridge, same same-tick
contract as the confidence globals, because OpenParams runs on the root signal
which has no pointer to the AI filter). Two writers: DeriveBarrierGeometry at
era 0, and the .cfg adoption a deployed model takes. Overrides both legs and
both Intelligent modes - the certificate is exact or it is nothing. TP is
ATR-anchored like the label, NOT risk-relative, so a floor-widened stop cannot
reshape the certified target.
2) BATCH NORM RUNS DEVICE-SIDE ON OPENCL. Four kernels in Network.cl -
forward, hidden gradient, gamma/beta accumulate, gamma/beta apply - each a
line-for-line transcription of the host implementation (NormalizeHost /
HiddenGradHost / StepGammaBeta) including every NaN guard, clamp, and the
exact moment-write ordering. The host copies remain the runtime for the DLL
and pure-MQL5 tiers and the reference the kernels must match.
Because this box has no OpenCL platform, the safety story is layered:
- shim validation: kernels compiled as C and driven against a fp64 host
transcription over NaN-poisoned stats, NaN gamma, over-clamp inputs, the
frozen path, both optimizers, 3 batches - ALL PASS, worst normalized diff
0.132 vs tolerance 1.0
- in-situ self-check: each kernel is compared against its host twin ON FIRST
USE on the real device (SelfCheckBn*), covering what the shim cannot - arg
indices and buffer bindings. Any disagreement resyncs from the good copy,
latches all BN kernels off process-wide, and training continues host-side.
A transcription bug costs a warning and some speed, never a poisoned .nnw.
- sync discipline: BatchOptions is now a CBufferDouble with explicit
authority tracking (m_bnDeviceAuthoritative). Checkpoints/saves pull
read-only; restores/loads/resets push; a mid-batch handover drains the
device gamma/beta accumulator into the host arrays so no sample is lost.
3) SMALL FIXES. Apply-kernel build failure now latches the dispatch path at
init (one warning instead of warning + failed Execute). Build tag bumped to
win-scoring-gpu-v1 - first tag change since expectancy-stop-v1 despite five
binary-changing commits.
Both build variants compile 0 errors, 0 warnings.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 17:51:40 -04:00
|
|
|
Print(ID + StringFormat(": live orders now use the MEASURED geometry - stop %.2f*ATR, target "
|
|
|
|
|
"%.2f*ATR - overriding the SL_Mode/TP_Mode enums (and the Intelligent "
|
|
|
|
|
"modes' confidence scaling), so the trade placed is the trade the deploy "
|
|
|
|
|
"gate certified.", m_derivedSlMult, m_derivedTpMult));
|
feat: derive the ATR multiples from measured excursions - no hardcoded geometry
The barrier was still two constants. SL_Mode/TP_Mode left the Inputs tab in
3482b6c, but the fallback was a hardcoded 2:6 and the geometry scan only ever
chose from a hardcoded grid {2,3} x {2,3,4,6,8,10}. Picking the least-bad of
eleven guesses is not deriving anything.
WHY THE SCAN WAS THE WRONG INSTRUMENT, now measurable rather than argued. It
ranks pairings by how predictable their OUTCOME is - a question about direction.
The excursion test (2c78f3b) ran on SP500 H1 and direction is the one thing
absent: ASYMMETRY p=0.0846, against RANGE/UP/DOWN all at p=0.0050, with RANGE
scoring 0.01345 vs a 0.00343 null - 4x, where the barrier label sits at 1.01x.
Hence the scan failing its own gate on every run, and its "winner" wandering
2:8 -> 3:8 -> 2:8 -> 2:4 across four runs of the same data. Excursion SIZE is
strongly measurable, so derive the geometry from that instead.
stop = q25 of measured ADVERSE travel (ordinary noise does not reach it)
target = q50 of measured FAVOURABLE travel (reached ~half the time, by
construction, inside the horizon)
Continuous, in ATR units, superseding the enum multiples. Reachability ("target
on X% of bars, stop on Y%") and the implied break-even are printed so the choice
is auditable rather than trusted.
FIXED-POINT ITERATION, not one-shot. ComputeBarrierHorizonBars scales the
horizon with the target (first-passage time grows with the band) and the
excursions are measured OVER the horizon, so target -> horizon -> excursions ->
target is a real loop - deriving once sizes the target from travel measured
under the PREVIOUS horizon. Re-measures until the multiples move <5%, capped at
3 passes, and says so if it does not settle.
Does NOT create expectancy, and the log says as much: chance precision equals
break-even at every geometry (m/(m+k) on both sides). It buys a target the
market reaches and a stop that survives noise. Where Min_Risk_Reward_Ratio
forces a target the market rarely reaches, it WARNS rather than overriding -
the ratio is the user's risk policy, so the honest move is to state its cost.
That is the collision that once rejected 100% of setups.
Pinned in the .cfg as doubles appended AFTER this morning's two ints, so .cfg
files written earlier today still load (their length guard finds no doubles) and
a model that carries them was trained on them and never re-derives.
Also fixes a message from e5ceed6 that claimed "this model resumed from disk"
unconditionally - it printed above a "seeding era 0" line on a brand-new model,
because the branch fires whenever the cache is not built, which is equally true
before a fresh model's first prebuild. A diagnostic that misreports its own
trigger is worse than one that says nothing: it gets quoted back as evidence.
FORCES A FULL RETRAIN (labels change).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-07 12:06:25 -04:00
|
|
|
Print(ID + StringFormat(": barrier geometry DERIVED from %d measured excursions - stop %.2f*ATR "
|
fix(geometry): the target was small BY CONSTRUCTION - ratio is now policy, scale is measured, ladder ceiling removed
The derivation read the stop from q75 of ADVERSE travel and the target from q50
of FAVOURABLE travel. Over one horizon those distributions are broadly the same
shape, so q75 > q50 MECHANICALLY - the target came out smaller than the stop no
matter what the market did. SP500 H4 shipped stop 3.07 / target 1.70: a 0.55:1
payoff needing 64.3%. That was never a measurement, it was two mismatched
constants.
The reachability line printed beside it - "target on 50.0% of bars, stop on
25.0%" - is exactly 1-q50 and 1-q75. Tautological. It cannot disconfirm
anything, and it read as validation.
WIDTH AND RATIO ARE INDEPENDENT AND ONLY ONE PAYS. EV = edge x width;
ratio is EV-neutral (a driftless walk reaches +m before -k with probability
k/(k+m), which IS break-even). Width is what buys cost efficiency: the spread
is a fixed 0.047*ATR here, so the shipped 4.77*ATR width paid it 21 times per
unit of travel. So:
RATIO = policy. BARRIER_TARGET_RR = 2.0 (user's 1:2). Break-even 33.3%.
SCALE = measured. The stop quantile is chosen from a ladder, WIDEST FIRST,
taking the first rung whose implied 2x target is still reached often
enough to be a trainable class.
That last clause is the difference from the min-reward:risk raise removed in
2026-08-09, which forced target = 2 x stop with NO reachability test, landed on
6.66*ATR reachable on 3.3% of bars, and trained the model to predict something
that essentially never happened. Same ratio; the scale now retreats until the
data says the target is attainable. Every rung is logged.
LADDER CEILING REMOVED. BARRIER_LADDER stopped at 5.00 and the expectancy scan's
"best resolvable pair on width alone" came back as stop 5.05 / target 4.95 - it
pinned to the top rung. A recommendation landing exactly on the edge of its own
search space is a boundary, not a finding: it cannot tell "5 ATR is optimal"
from "5 ATR is all we allowed". Extended to 20*ATR (8 -> 14 rungs). Nothing else
needs editing - every consumer is parameterised by BARRIER_LADDER_COUNT - and
the horizon constraints (decided >= 60%, reachability floor) now bind instead of
a constant.
THE SCAN COULD NOT SEE THE SHIPPED GEOMETRY. ReportBarrierGeometryScan looked
the configured pair up in its integer grid, and DeriveBarrierGeometry produces
CONTINUOUS multiples (3.07/1.70) that can never equal a grid point - so
cfgExcess stayed at its -1.0 sentinel and the report printed "configured 3:2
scores -1.00000", which reads as a catastrophic score and actually means "never
evaluated". Worse, the grid skipped target<stop entirely because it "inverts the
trade's whole premise" - while the derivation was shipping exactly that. The
incumbent is now always scored as a peer (never crowned; it is already in force
and is not an enum pairing the scan could adopt).
BREAK-EVEN NOW INCLUDES THE SPREAD. Every report quoted the frictionless
SL/(SL+TP). On SP500 H4 that read 64.3% while the MEASURED zero-skill rate was
62.1% - a 2.2pp gap that IS the cost, and that made every model look 2.2pp
better than it was. CostAdjustedBreakEvenPct() prices a win at (TP - spread) and
a loss at (SL + spread), matching the expectancy scan's convention exactly so
the two reports cannot disagree.
It also feeds FitDirConfThreshold, which is the correctness half: the operating
point subtracts break-even from precision, so the frictionless figure made every
candidate threshold look better by the width of the spread - 2.2pp against a
measured edge of 2.3pp, i.e. very nearly all of it.
Era line now carries both: "break-even 64.3% frictionless, 66.6% AFTER SPREAD".
Forces a full relabel and retrain. Requested.
NOT COMPILED - user compiles.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 00:34:32 -04:00
|
|
|
"(q%.0f of adverse travel, chosen by the SCALE ladder above), target %.2f*ATR "
|
|
|
|
|
"(= %.1f x the stop, RATIO BY POLICY not by measurement) | width %.2f*ATR | "
|
fix(geometry): the reachability floor measured the WRONG WINDOW - my bug from bc57aca, and it cost real width
RECONCILED: the derivation reported "target reached on 17.7% of bars" while the
label cache reported Buy on 35.9%. Nothing was broken. They measure different
windows, and both are correct:
EXCURSION window ~12 bars (the SWING MEDIAN) - what m_excUpCache accumulates
over. Deliberately short: sizing a barrier off travel
measured over a horizon that itself scales with the barrier
is circular, and it ran away to 14-31*ATR on EURUSD/USDCAD
in 2026-08-07. That guard is correct and stays.
BARRIER horizon 64 bars - what the LABEL walk and the first-passage ladder
run over, and how long the EA actually holds the trade.
So `up >= target` is a 12-bar question and `label == Buy` is a 64-bar one, and
the second can freely exceed the first. TripleBarrierLabel gates the excursion
accumulation on `idx - t <= excWindow` while the barrier walk and the ladder run
the full horizon - the split is explicit and intentional.
THE BUG IS MINE. bc57aca's scale ladder tested reachability with `up[i] >= tp`,
i.e. it asked the 12-bar question about a 64-bar trade. That understates
reachability by ~2x, which is why EVERY wide rung was rejected and the geometry
fell back to the tightest rung at 1.61/3.21. The data supported considerably
wider; the test was just asking the wrong question.
FIX: LadderWinShare() reads the answer off the first-passage ladder - target
touched strictly before the stop, over the full horizon, tie to the stop. That
is the identical question the label walk asks, so the ladder share and the Buy
rate should now agree to within rung discretisation. Both legs snap to the
SMALLEST rung at or above the requested multiple (harder target, harder stop) so
the floor stays conservative.
Expect the scale ladder to select a WIDER rung on the next relabel. On this
data the excursion test read 17.7% at q50 where the true full-horizon share is
35.9%, so rungs that scored 8.1% and 2.8% were likely well above the floor.
ALSO:
- Window reconciliation now PRINTED every derivation: excursion travel share,
ladder win share, and the label cache's Buy share side by side, with the
ladder-vs-label gap flagged if it exceeds rung discretisation. Those two must
agree; if they ever stop agreeing, one of them is wrong and the line says so.
- Renamed tpReach/slReach -> tpTravel/slTravel and relabelled the log line. They
describe the EXCURSION window and are near-tautological there (a q50 stop is
exceeded by ~50% of bars); calling them "reached within the horizon" is what
made the two quantities look like one.
- BARRIER_MIN_TP_REACH_PCT is now BARRIER_MIN_REACH_FRACTION_OF_BE (0.60) x
break-even instead of a hardcoded 20.0. Break-even for 1:RR is 100/(1+RR), so
the absolute floor silently tightened as RR rose - 0.60x at RR=2 but 0.80x at
RR=3, penalising the user for asking for a bigger target. Evaluates to exactly
20.0% at the shipped RR=2, so this is a no-op today and correct if the knob
moves.
NOT COMPILED - user compiles.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 00:51:08 -04:00
|
|
|
"travelled within the %d-bar EXCURSION window: target on %.1f%% of bars, stop "
|
|
|
|
|
"on %.1f%% (near-tautological - that is where the quantiles were read) | "
|
fix(geometry): the target was small BY CONSTRUCTION - ratio is now policy, scale is measured, ladder ceiling removed
The derivation read the stop from q75 of ADVERSE travel and the target from q50
of FAVOURABLE travel. Over one horizon those distributions are broadly the same
shape, so q75 > q50 MECHANICALLY - the target came out smaller than the stop no
matter what the market did. SP500 H4 shipped stop 3.07 / target 1.70: a 0.55:1
payoff needing 64.3%. That was never a measurement, it was two mismatched
constants.
The reachability line printed beside it - "target on 50.0% of bars, stop on
25.0%" - is exactly 1-q50 and 1-q75. Tautological. It cannot disconfirm
anything, and it read as validation.
WIDTH AND RATIO ARE INDEPENDENT AND ONLY ONE PAYS. EV = edge x width;
ratio is EV-neutral (a driftless walk reaches +m before -k with probability
k/(k+m), which IS break-even). Width is what buys cost efficiency: the spread
is a fixed 0.047*ATR here, so the shipped 4.77*ATR width paid it 21 times per
unit of travel. So:
RATIO = policy. BARRIER_TARGET_RR = 2.0 (user's 1:2). Break-even 33.3%.
SCALE = measured. The stop quantile is chosen from a ladder, WIDEST FIRST,
taking the first rung whose implied 2x target is still reached often
enough to be a trainable class.
That last clause is the difference from the min-reward:risk raise removed in
2026-08-09, which forced target = 2 x stop with NO reachability test, landed on
6.66*ATR reachable on 3.3% of bars, and trained the model to predict something
that essentially never happened. Same ratio; the scale now retreats until the
data says the target is attainable. Every rung is logged.
LADDER CEILING REMOVED. BARRIER_LADDER stopped at 5.00 and the expectancy scan's
"best resolvable pair on width alone" came back as stop 5.05 / target 4.95 - it
pinned to the top rung. A recommendation landing exactly on the edge of its own
search space is a boundary, not a finding: it cannot tell "5 ATR is optimal"
from "5 ATR is all we allowed". Extended to 20*ATR (8 -> 14 rungs). Nothing else
needs editing - every consumer is parameterised by BARRIER_LADDER_COUNT - and
the horizon constraints (decided >= 60%, reachability floor) now bind instead of
a constant.
THE SCAN COULD NOT SEE THE SHIPPED GEOMETRY. ReportBarrierGeometryScan looked
the configured pair up in its integer grid, and DeriveBarrierGeometry produces
CONTINUOUS multiples (3.07/1.70) that can never equal a grid point - so
cfgExcess stayed at its -1.0 sentinel and the report printed "configured 3:2
scores -1.00000", which reads as a catastrophic score and actually means "never
evaluated". Worse, the grid skipped target<stop entirely because it "inverts the
trade's whole premise" - while the derivation was shipping exactly that. The
incumbent is now always scored as a peer (never crowned; it is already in force
and is not an enum pairing the scan could adopt).
BREAK-EVEN NOW INCLUDES THE SPREAD. Every report quoted the frictionless
SL/(SL+TP). On SP500 H4 that read 64.3% while the MEASURED zero-skill rate was
62.1% - a 2.2pp gap that IS the cost, and that made every model look 2.2pp
better than it was. CostAdjustedBreakEvenPct() prices a win at (TP - spread) and
a loss at (SL + spread), matching the expectancy scan's convention exactly so
the two reports cannot disagree.
It also feeds FitDirConfThreshold, which is the correctness half: the operating
point subtracts break-even from precision, so the frictionless figure made every
candidate threshold look better by the width of the spread - 2.2pp against a
measured edge of 2.3pp, i.e. very nearly all of it.
Era line now carries both: "break-even 64.3% frictionless, 66.6% AFTER SPREAD".
Forces a full relabel and retrain. Requested.
NOT COMPILED - user compiles.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 00:34:32 -04:00
|
|
|
"implied break-even %.1f%%. This does NOT create expectancy - chance precision "
|
|
|
|
|
"equals break-even at every RATIO - what the geometry buys is WIDTH, and width "
|
|
|
|
|
"is the EV multiplier because the spread is a fixed cost per trade.",
|
fix: normalise the asymmetry target - the raw one is confounded by volatility
Three symbols ran the excursion test. RANGE/UP/DOWN cleared on all three;
raw ASYMMETRY cleared on EURUSD and USDCAD at p=0.0050 and not on SP500
(p=0.1045). That looked like the first directional signal this project has
found. It probably is not, and the test as built could not tell.
(up-dn) IS NOT SCALE-FREE. If sigma is predictable - and RANGE clears at ~4x its
null on every instrument - and the directional part is symmetric noise eps, then
up-dn ~ sigma*eps, so a large sigma pushes the value into BOTH outer terciles. A
pure volatility predictor scores positive MI against a 3-bin (up-dn) while
carrying no directional information at all. Crucially that confound REPLICATES,
so reproducing on two instruments is not evidence against it - and the effect
sizes fit it: asymmetry runs 1.3-1.6x its null where RANGE runs ~4x, and carries
~0.1% of the target's entropy against RANGE's ~0.9%. That is the shape of a
leaked fraction of the volatility signal, not an independent one.
So add (up-dn)/(up+dn): bounded in [-1,+1], volatility divided out, and the only
target a directional claim may rest on. The verdict now separates the cases and
NAMES the confound when raw clears while normalised does not, instead of
reporting the raw line as a finding.
Two bugs of mine in the same block, both caught by output rather than review:
- The derived-geometry line had a MISORDERED argument list: it printed
"stop 25.00*ATR (q3 of adverse travel)" - the quantile percentage as the
multiple and the multiple as the quantile. Real values were 2.61 stop /
8.03 target. A 25*ATR stop is absurd on its face, which is why it was seen.
- THE STOP QUANTILE WAS BACKWARDS, and this one changes labels. It was 0.25
"so ordinary noise does not reach it", but q25 means 75% of bars EXCEED the
stop - hit three times in four. The printed reachability said exactly that
("stop on 75.0% of bars"). Now 0.75. A quantile is a threshold, not a rate.
This is the entire reason reachability is measured and printed rather than
assumed.
Also raises BARRIER_DERIVE_MAX_PASSES 3 -> 5: SP500 did not settle in 3 (stop
still moving ~14% per pass) while EURUSD and USDCAD converged on pass 2. And
bounds both quantile indices with MathMin(..., n-1) so q=1.0 cannot run off the
end of the sorted array.
The geometry from the previous run is NOT usable and the asymmetry result is
unresolved, not established. Both are decided by the next run.
FORCES A FULL RETRAIN (the stop quantile changes every label).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-07 13:04:13 -04:00
|
|
|
//--- ORDER MATTERS AND WAS WRONG ONCE: the multiples and the quantile labels
|
|
|
|
|
//--- were swapped, so the log read "stop 25.00*ATR (q3 ...)" - printing the
|
|
|
|
|
//--- quantile percentage as the multiple and the multiple as the quantile.
|
|
|
|
|
//--- 25*ATR is absurd on its face, which is the only reason it was caught.
|
fix(geometry): the target was small BY CONSTRUCTION - ratio is now policy, scale is measured, ladder ceiling removed
The derivation read the stop from q75 of ADVERSE travel and the target from q50
of FAVOURABLE travel. Over one horizon those distributions are broadly the same
shape, so q75 > q50 MECHANICALLY - the target came out smaller than the stop no
matter what the market did. SP500 H4 shipped stop 3.07 / target 1.70: a 0.55:1
payoff needing 64.3%. That was never a measurement, it was two mismatched
constants.
The reachability line printed beside it - "target on 50.0% of bars, stop on
25.0%" - is exactly 1-q50 and 1-q75. Tautological. It cannot disconfirm
anything, and it read as validation.
WIDTH AND RATIO ARE INDEPENDENT AND ONLY ONE PAYS. EV = edge x width;
ratio is EV-neutral (a driftless walk reaches +m before -k with probability
k/(k+m), which IS break-even). Width is what buys cost efficiency: the spread
is a fixed 0.047*ATR here, so the shipped 4.77*ATR width paid it 21 times per
unit of travel. So:
RATIO = policy. BARRIER_TARGET_RR = 2.0 (user's 1:2). Break-even 33.3%.
SCALE = measured. The stop quantile is chosen from a ladder, WIDEST FIRST,
taking the first rung whose implied 2x target is still reached often
enough to be a trainable class.
That last clause is the difference from the min-reward:risk raise removed in
2026-08-09, which forced target = 2 x stop with NO reachability test, landed on
6.66*ATR reachable on 3.3% of bars, and trained the model to predict something
that essentially never happened. Same ratio; the scale now retreats until the
data says the target is attainable. Every rung is logged.
LADDER CEILING REMOVED. BARRIER_LADDER stopped at 5.00 and the expectancy scan's
"best resolvable pair on width alone" came back as stop 5.05 / target 4.95 - it
pinned to the top rung. A recommendation landing exactly on the edge of its own
search space is a boundary, not a finding: it cannot tell "5 ATR is optimal"
from "5 ATR is all we allowed". Extended to 20*ATR (8 -> 14 rungs). Nothing else
needs editing - every consumer is parameterised by BARRIER_LADDER_COUNT - and
the horizon constraints (decided >= 60%, reachability floor) now bind instead of
a constant.
THE SCAN COULD NOT SEE THE SHIPPED GEOMETRY. ReportBarrierGeometryScan looked
the configured pair up in its integer grid, and DeriveBarrierGeometry produces
CONTINUOUS multiples (3.07/1.70) that can never equal a grid point - so
cfgExcess stayed at its -1.0 sentinel and the report printed "configured 3:2
scores -1.00000", which reads as a catastrophic score and actually means "never
evaluated". Worse, the grid skipped target<stop entirely because it "inverts the
trade's whole premise" - while the derivation was shipping exactly that. The
incumbent is now always scored as a peer (never crowned; it is already in force
and is not an enum pairing the scan could adopt).
BREAK-EVEN NOW INCLUDES THE SPREAD. Every report quoted the frictionless
SL/(SL+TP). On SP500 H4 that read 64.3% while the MEASURED zero-skill rate was
62.1% - a 2.2pp gap that IS the cost, and that made every model look 2.2pp
better than it was. CostAdjustedBreakEvenPct() prices a win at (TP - spread) and
a loss at (SL + spread), matching the expectancy scan's convention exactly so
the two reports cannot disagree.
It also feeds FitDirConfThreshold, which is the correctness half: the operating
point subtracts break-even from precision, so the frictionless figure made every
candidate threshold look better by the width of the spread - 2.2pp against a
measured edge of 2.3pp, i.e. very nearly all of it.
Era line now carries both: "break-even 64.3% frictionless, 66.6% AFTER SPREAD".
Forces a full relabel and retrain. Requested.
NOT COMPILED - user compiles.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 00:34:32 -04:00
|
|
|
n, m_derivedSlMult, 100.0 * chosenQ, m_derivedTpMult, BARRIER_TARGET_RR,
|
fix(geometry): the reachability floor measured the WRONG WINDOW - my bug from bc57aca, and it cost real width
RECONCILED: the derivation reported "target reached on 17.7% of bars" while the
label cache reported Buy on 35.9%. Nothing was broken. They measure different
windows, and both are correct:
EXCURSION window ~12 bars (the SWING MEDIAN) - what m_excUpCache accumulates
over. Deliberately short: sizing a barrier off travel
measured over a horizon that itself scales with the barrier
is circular, and it ran away to 14-31*ATR on EURUSD/USDCAD
in 2026-08-07. That guard is correct and stays.
BARRIER horizon 64 bars - what the LABEL walk and the first-passage ladder
run over, and how long the EA actually holds the trade.
So `up >= target` is a 12-bar question and `label == Buy` is a 64-bar one, and
the second can freely exceed the first. TripleBarrierLabel gates the excursion
accumulation on `idx - t <= excWindow` while the barrier walk and the ladder run
the full horizon - the split is explicit and intentional.
THE BUG IS MINE. bc57aca's scale ladder tested reachability with `up[i] >= tp`,
i.e. it asked the 12-bar question about a 64-bar trade. That understates
reachability by ~2x, which is why EVERY wide rung was rejected and the geometry
fell back to the tightest rung at 1.61/3.21. The data supported considerably
wider; the test was just asking the wrong question.
FIX: LadderWinShare() reads the answer off the first-passage ladder - target
touched strictly before the stop, over the full horizon, tie to the stop. That
is the identical question the label walk asks, so the ladder share and the Buy
rate should now agree to within rung discretisation. Both legs snap to the
SMALLEST rung at or above the requested multiple (harder target, harder stop) so
the floor stays conservative.
Expect the scale ladder to select a WIDER rung on the next relabel. On this
data the excursion test read 17.7% at q50 where the true full-horizon share is
35.9%, so rungs that scored 8.1% and 2.8% were likely well above the floor.
ALSO:
- Window reconciliation now PRINTED every derivation: excursion travel share,
ladder win share, and the label cache's Buy share side by side, with the
ladder-vs-label gap flagged if it exceeds rung discretisation. Those two must
agree; if they ever stop agreeing, one of them is wrong and the line says so.
- Renamed tpReach/slReach -> tpTravel/slTravel and relabelled the log line. They
describe the EXCURSION window and are near-tautological there (a q50 stop is
exceeded by ~50% of bars); calling them "reached within the horizon" is what
made the two quantities look like one.
- BARRIER_MIN_TP_REACH_PCT is now BARRIER_MIN_REACH_FRACTION_OF_BE (0.60) x
break-even instead of a hardcoded 20.0. Break-even for 1:RR is 100/(1+RR), so
the absolute floor silently tightened as RR rose - 0.60x at RR=2 but 0.80x at
RR=3, penalising the user for asking for a bigger target. Evaluates to exactly
20.0% at the shipped RR=2, so this is a no-op today and correct if the knob
moves.
NOT COMPILED - user compiles.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 00:51:08 -04:00
|
|
|
m_derivedSlMult + m_derivedTpMult, m_swingMedianBars,
|
|
|
|
|
tpTravel, slTravel, breakeven));
|
|
|
|
|
//--- THE THREE WINDOWS, printed together because two of them look like the same quantity and are not.
|
|
|
|
|
//--- This is the reconciliation for a 2026-08-17 discrepancy that cost an hour: the derivation
|
|
|
|
|
//--- reported "target reached on 17.7% of bars" while the label cache reported Buy on 35.9% - twice as
|
|
|
|
|
//--- many wins as there were bars that ever reached the target. Nothing was broken. They measure
|
|
|
|
|
//--- different windows:
|
|
|
|
|
//---
|
|
|
|
|
//--- excursion window ~12 bars (the SWING MEDIAN) - what m_excUpCache accumulates over, kept short
|
|
|
|
|
//--- ON PURPOSE so the barrier is not sized off travel measured over a horizon
|
|
|
|
|
//--- that scales with the barrier (the 2026-08-07 runaway to 14-31*ATR)
|
|
|
|
|
//--- barrier horizon 64 bars - what the LABEL walk and the first-passage ladder run over, and how
|
|
|
|
|
//--- long the EA actually holds the trade
|
|
|
|
|
//---
|
|
|
|
|
//--- So `up >= target` is a 12-bar question and `label == Buy` is a 64-bar one, and the second can
|
|
|
|
|
//--- freely exceed the first. The ladder share below is the 64-bar question asked the same way the
|
|
|
|
|
//--- label asks it, so THAT is the one that should match the Buy rate - and any gap between those two
|
|
|
|
|
//--- is real: it can only come from the ladder's snap-to-rung discretisation.
|
|
|
|
|
if(!conditional && ArraySize(labUp) >= n && n > 0)
|
|
|
|
|
{
|
|
|
|
|
int excReach = 0, buyCount = 0;
|
|
|
|
|
for(int i = 0; i < n; i++)
|
|
|
|
|
{
|
|
|
|
|
if(upUnsorted[i] >= tpRaw)
|
|
|
|
|
excReach++;
|
|
|
|
|
if(labUp[i])
|
|
|
|
|
buyCount++;
|
|
|
|
|
}
|
fix(labels): overlapping-label sample correction + horizon cap on the scale ladder
Three defects, all surfaced by the 2026-08-17 SP500 H4 run that shipped
stop 4.86 / target 9.71 (width 14.57*ATR, horizon 384).
1. EVERY STANDARD ERROR ASSUMED INDEPENDENT SAMPLES. Triple-barrier labels
started one per bar overlap by the label's lifespan, so n calls are worth
~n/L independent observations (Lopez de Prado, AFML ch. 4 - sample
uniqueness). All three sqrt(p(1-p)/n) sites divided by the RAW count.
The tell: the operating point's null-of-the-maximum gate is family-wise and
should fire on ~5% of eras under the null. Measured fire rates - PAI 47/73
(64%), ConvLSTM 9/24, LSTM 8/21 (38%), CONV 4/62 (6%). CONV, the only model
whose margin distribution admits few bins, sat on the null; the rest cleared
a bar that was too low by ~sqrt(L). PAI's deployed threshold consequently
alternated between the ENDS of its own range era to era (0.10 -> 0.88 ->
0.86 -> 0.66; coverage 16% <-> 73%).
TripleBarrierLabel now records when each label became KNOWABLE - the first
winning touch, or both stops, or the timeout - and the prebuild accumulates
the mean. EffectiveSampleSize() feeds the operating point, the member deploy
gate and the ensemble vote gate. Conservative by construction (n/L is an
upper bound on the damage); gates get harder, never easier.
2. THE SCALE LADDER RAN AWAY, again. Horizon scales as swingMedian*sl*tp, and
since 4d8cb08 reachability is measured OVER that horizon - so a wider rung
buys itself the time that makes it look reachable. Same target -> horizon ->
reach -> target loop the excursion window is kept short to avoid; fixing the
window confusion reopened it through the other door. It walked 128 -> 256 ->
384 bars and stopped at q90, the widest rung there is, with every rung
reading 39-48% against a 20% floor. A floor nothing fails selects nothing.
Rungs whose required horizon exceeds BARRIER_HORIZON_MAX are now rejected -
the same rule ReportGeometryExpectancyScan already applied. It was printing
the shipped pair as CLAMPED and disqualified ('h384!') two lines under the
deriver that chose it: two subsystems, one geometry, opposite verdicts.
3. THE RUNG SNAP DESTROYED THE RATIO IT WAS COMPARING. Both legs snapped
independently to the coarse first-passage grid, re-rating each candidate:
q90 4.86/9.71 -> 5.00/10.00 (2.00), q85 4.07/8.14 -> 5.00/10.00 (IDENTICAL
measurement), q75 3.07/6.13 -> 4.00/6.50 (1.63 - a nearer target). So the
ladder compared win shares taken at ratios from 1.63 to 2.17 and read the
differences as scale. It is why the reach column came out non-monotone in
width (q75 48.5% above q90 42.9%). The stop now snaps to its nearest rung in
log space and the target follows the ratio off it; the pair actually measured
is returned and logged, so a collision reads as a collision.
Also: LadderWinShare guarded against the conditional (fractal) geometry path,
which fills n from m_fracLegCount while leaving idxList empty - a latent
out-of-bounds on a currently-dead path.
New log lines: mean label lifespan and effective n on the label-cache line, the
required-vs-available horizon per rung, and the grid pair the reconciliation
actually measured (its tolerance now scales with the grid skew instead of a flat
5pp).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 01:12:05 -04:00
|
|
|
double recSl = 0.0, recTp = 0.0;
|
|
|
|
|
double ladderShare = LadderWinShare(idxList, n, slRaw, tpRaw, recSl, recTp);
|
fix(geometry): the reachability floor measured the WRONG WINDOW - my bug from bc57aca, and it cost real width
RECONCILED: the derivation reported "target reached on 17.7% of bars" while the
label cache reported Buy on 35.9%. Nothing was broken. They measure different
windows, and both are correct:
EXCURSION window ~12 bars (the SWING MEDIAN) - what m_excUpCache accumulates
over. Deliberately short: sizing a barrier off travel
measured over a horizon that itself scales with the barrier
is circular, and it ran away to 14-31*ATR on EURUSD/USDCAD
in 2026-08-07. That guard is correct and stays.
BARRIER horizon 64 bars - what the LABEL walk and the first-passage ladder
run over, and how long the EA actually holds the trade.
So `up >= target` is a 12-bar question and `label == Buy` is a 64-bar one, and
the second can freely exceed the first. TripleBarrierLabel gates the excursion
accumulation on `idx - t <= excWindow` while the barrier walk and the ladder run
the full horizon - the split is explicit and intentional.
THE BUG IS MINE. bc57aca's scale ladder tested reachability with `up[i] >= tp`,
i.e. it asked the 12-bar question about a 64-bar trade. That understates
reachability by ~2x, which is why EVERY wide rung was rejected and the geometry
fell back to the tightest rung at 1.61/3.21. The data supported considerably
wider; the test was just asking the wrong question.
FIX: LadderWinShare() reads the answer off the first-passage ladder - target
touched strictly before the stop, over the full horizon, tie to the stop. That
is the identical question the label walk asks, so the ladder share and the Buy
rate should now agree to within rung discretisation. Both legs snap to the
SMALLEST rung at or above the requested multiple (harder target, harder stop) so
the floor stays conservative.
Expect the scale ladder to select a WIDER rung on the next relabel. On this
data the excursion test read 17.7% at q50 where the true full-horizon share is
35.9%, so rungs that scored 8.1% and 2.8% were likely well above the floor.
ALSO:
- Window reconciliation now PRINTED every derivation: excursion travel share,
ladder win share, and the label cache's Buy share side by side, with the
ladder-vs-label gap flagged if it exceeds rung discretisation. Those two must
agree; if they ever stop agreeing, one of them is wrong and the line says so.
- Renamed tpReach/slReach -> tpTravel/slTravel and relabelled the log line. They
describe the EXCURSION window and are near-tautological there (a q50 stop is
exceeded by ~50% of bars); calling them "reached within the horizon" is what
made the two quantities look like one.
- BARRIER_MIN_TP_REACH_PCT is now BARRIER_MIN_REACH_FRACTION_OF_BE (0.60) x
break-even instead of a hardcoded 20.0. Break-even for 1:RR is 100/(1+RR), so
the absolute floor silently tightened as RR rose - 0.60x at RR=2 but 0.80x at
RR=3, penalising the user for asking for a bigger target. Evaluates to exactly
20.0% at the shipped RR=2, so this is a no-op today and correct if the knob
moves.
NOT COMPILED - user compiles.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 00:51:08 -04:00
|
|
|
double buyPct = 100.0 * buyCount / n;
|
|
|
|
|
double gap = MathAbs(ladderShare - buyPct);
|
fix(labels): overlapping-label sample correction + horizon cap on the scale ladder
Three defects, all surfaced by the 2026-08-17 SP500 H4 run that shipped
stop 4.86 / target 9.71 (width 14.57*ATR, horizon 384).
1. EVERY STANDARD ERROR ASSUMED INDEPENDENT SAMPLES. Triple-barrier labels
started one per bar overlap by the label's lifespan, so n calls are worth
~n/L independent observations (Lopez de Prado, AFML ch. 4 - sample
uniqueness). All three sqrt(p(1-p)/n) sites divided by the RAW count.
The tell: the operating point's null-of-the-maximum gate is family-wise and
should fire on ~5% of eras under the null. Measured fire rates - PAI 47/73
(64%), ConvLSTM 9/24, LSTM 8/21 (38%), CONV 4/62 (6%). CONV, the only model
whose margin distribution admits few bins, sat on the null; the rest cleared
a bar that was too low by ~sqrt(L). PAI's deployed threshold consequently
alternated between the ENDS of its own range era to era (0.10 -> 0.88 ->
0.86 -> 0.66; coverage 16% <-> 73%).
TripleBarrierLabel now records when each label became KNOWABLE - the first
winning touch, or both stops, or the timeout - and the prebuild accumulates
the mean. EffectiveSampleSize() feeds the operating point, the member deploy
gate and the ensemble vote gate. Conservative by construction (n/L is an
upper bound on the damage); gates get harder, never easier.
2. THE SCALE LADDER RAN AWAY, again. Horizon scales as swingMedian*sl*tp, and
since 4d8cb08 reachability is measured OVER that horizon - so a wider rung
buys itself the time that makes it look reachable. Same target -> horizon ->
reach -> target loop the excursion window is kept short to avoid; fixing the
window confusion reopened it through the other door. It walked 128 -> 256 ->
384 bars and stopped at q90, the widest rung there is, with every rung
reading 39-48% against a 20% floor. A floor nothing fails selects nothing.
Rungs whose required horizon exceeds BARRIER_HORIZON_MAX are now rejected -
the same rule ReportGeometryExpectancyScan already applied. It was printing
the shipped pair as CLAMPED and disqualified ('h384!') two lines under the
deriver that chose it: two subsystems, one geometry, opposite verdicts.
3. THE RUNG SNAP DESTROYED THE RATIO IT WAS COMPARING. Both legs snapped
independently to the coarse first-passage grid, re-rating each candidate:
q90 4.86/9.71 -> 5.00/10.00 (2.00), q85 4.07/8.14 -> 5.00/10.00 (IDENTICAL
measurement), q75 3.07/6.13 -> 4.00/6.50 (1.63 - a nearer target). So the
ladder compared win shares taken at ratios from 1.63 to 2.17 and read the
differences as scale. It is why the reach column came out non-monotone in
width (q75 48.5% above q90 42.9%). The stop now snaps to its nearest rung in
log space and the target follows the ratio off it; the pair actually measured
is returned and logged, so a collision reads as a collision.
Also: LadderWinShare guarded against the conditional (fractal) geometry path,
which fills n from m_fracLegCount while leaving idxList empty - a latent
out-of-bounds on a currently-dead path.
New log lines: mean label lifespan and effective n on the label-cache line, the
required-vs-available horizon per rung, and the grid pair the reconciliation
actually measured (its tolerance now scales with the grid skew instead of a flat
5pp).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 01:12:05 -04:00
|
|
|
//--- The gap tolerance has to scale with how badly the grid mis-states the pair, not sit at a flat
|
|
|
|
|
//--- 5pp: the ladder measures recSl/recTp, the labels measure slRaw/tpRaw, and when those differ the
|
|
|
|
|
//--- two are answering NEARLY the same question rather than exactly it. A fixed tolerance either
|
|
|
|
|
//--- cries wolf on a coarse rung or says nothing on a fine one.
|
|
|
|
|
double gridSkew = (slRaw > 0.0 && tpRaw > 0.0 && recSl > 0.0)
|
|
|
|
|
? MathAbs((recTp / recSl) - (tpRaw / slRaw)) / (tpRaw / slRaw) : 0.0;
|
|
|
|
|
double gapTol = 5.0 + 100.0 * gridSkew;
|
fix(geometry): the reachability floor measured the WRONG WINDOW - my bug from bc57aca, and it cost real width
RECONCILED: the derivation reported "target reached on 17.7% of bars" while the
label cache reported Buy on 35.9%. Nothing was broken. They measure different
windows, and both are correct:
EXCURSION window ~12 bars (the SWING MEDIAN) - what m_excUpCache accumulates
over. Deliberately short: sizing a barrier off travel
measured over a horizon that itself scales with the barrier
is circular, and it ran away to 14-31*ATR on EURUSD/USDCAD
in 2026-08-07. That guard is correct and stays.
BARRIER horizon 64 bars - what the LABEL walk and the first-passage ladder
run over, and how long the EA actually holds the trade.
So `up >= target` is a 12-bar question and `label == Buy` is a 64-bar one, and
the second can freely exceed the first. TripleBarrierLabel gates the excursion
accumulation on `idx - t <= excWindow` while the barrier walk and the ladder run
the full horizon - the split is explicit and intentional.
THE BUG IS MINE. bc57aca's scale ladder tested reachability with `up[i] >= tp`,
i.e. it asked the 12-bar question about a 64-bar trade. That understates
reachability by ~2x, which is why EVERY wide rung was rejected and the geometry
fell back to the tightest rung at 1.61/3.21. The data supported considerably
wider; the test was just asking the wrong question.
FIX: LadderWinShare() reads the answer off the first-passage ladder - target
touched strictly before the stop, over the full horizon, tie to the stop. That
is the identical question the label walk asks, so the ladder share and the Buy
rate should now agree to within rung discretisation. Both legs snap to the
SMALLEST rung at or above the requested multiple (harder target, harder stop) so
the floor stays conservative.
Expect the scale ladder to select a WIDER rung on the next relabel. On this
data the excursion test read 17.7% at q50 where the true full-horizon share is
35.9%, so rungs that scored 8.1% and 2.8% were likely well above the floor.
ALSO:
- Window reconciliation now PRINTED every derivation: excursion travel share,
ladder win share, and the label cache's Buy share side by side, with the
ladder-vs-label gap flagged if it exceeds rung discretisation. Those two must
agree; if they ever stop agreeing, one of them is wrong and the line says so.
- Renamed tpReach/slReach -> tpTravel/slTravel and relabelled the log line. They
describe the EXCURSION window and are near-tautological there (a q50 stop is
exceeded by ~50% of bars); calling them "reached within the horizon" is what
made the two quantities look like one.
- BARRIER_MIN_TP_REACH_PCT is now BARRIER_MIN_REACH_FRACTION_OF_BE (0.60) x
break-even instead of a hardcoded 20.0. Break-even for 1:RR is 100/(1+RR), so
the absolute floor silently tightened as RR rose - 0.60x at RR=2 but 0.80x at
RR=3, penalising the user for asking for a bigger target. Evaluates to exactly
20.0% at the shipped RR=2, so this is a no-op today and correct if the knob
moves.
NOT COMPILED - user compiles.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 00:51:08 -04:00
|
|
|
Print(ID + StringFormat(": window reconciliation at stop %.2f target %.2f - EXCURSION window "
|
|
|
|
|
"(%d bars, sizes the barrier): target travelled on %.1f%% of bars | "
|
|
|
|
|
"BARRIER horizon (%d bars, what the trade lives in): ladder says a long "
|
fix(labels): overlapping-label sample correction + horizon cap on the scale ladder
Three defects, all surfaced by the 2026-08-17 SP500 H4 run that shipped
stop 4.86 / target 9.71 (width 14.57*ATR, horizon 384).
1. EVERY STANDARD ERROR ASSUMED INDEPENDENT SAMPLES. Triple-barrier labels
started one per bar overlap by the label's lifespan, so n calls are worth
~n/L independent observations (Lopez de Prado, AFML ch. 4 - sample
uniqueness). All three sqrt(p(1-p)/n) sites divided by the RAW count.
The tell: the operating point's null-of-the-maximum gate is family-wise and
should fire on ~5% of eras under the null. Measured fire rates - PAI 47/73
(64%), ConvLSTM 9/24, LSTM 8/21 (38%), CONV 4/62 (6%). CONV, the only model
whose margin distribution admits few bins, sat on the null; the rest cleared
a bar that was too low by ~sqrt(L). PAI's deployed threshold consequently
alternated between the ENDS of its own range era to era (0.10 -> 0.88 ->
0.86 -> 0.66; coverage 16% <-> 73%).
TripleBarrierLabel now records when each label became KNOWABLE - the first
winning touch, or both stops, or the timeout - and the prebuild accumulates
the mean. EffectiveSampleSize() feeds the operating point, the member deploy
gate and the ensemble vote gate. Conservative by construction (n/L is an
upper bound on the damage); gates get harder, never easier.
2. THE SCALE LADDER RAN AWAY, again. Horizon scales as swingMedian*sl*tp, and
since 4d8cb08 reachability is measured OVER that horizon - so a wider rung
buys itself the time that makes it look reachable. Same target -> horizon ->
reach -> target loop the excursion window is kept short to avoid; fixing the
window confusion reopened it through the other door. It walked 128 -> 256 ->
384 bars and stopped at q90, the widest rung there is, with every rung
reading 39-48% against a 20% floor. A floor nothing fails selects nothing.
Rungs whose required horizon exceeds BARRIER_HORIZON_MAX are now rejected -
the same rule ReportGeometryExpectancyScan already applied. It was printing
the shipped pair as CLAMPED and disqualified ('h384!') two lines under the
deriver that chose it: two subsystems, one geometry, opposite verdicts.
3. THE RUNG SNAP DESTROYED THE RATIO IT WAS COMPARING. Both legs snapped
independently to the coarse first-passage grid, re-rating each candidate:
q90 4.86/9.71 -> 5.00/10.00 (2.00), q85 4.07/8.14 -> 5.00/10.00 (IDENTICAL
measurement), q75 3.07/6.13 -> 4.00/6.50 (1.63 - a nearer target). So the
ladder compared win shares taken at ratios from 1.63 to 2.17 and read the
differences as scale. It is why the reach column came out non-monotone in
width (q75 48.5% above q90 42.9%). The stop now snaps to its nearest rung in
log space and the target follows the ratio off it; the pair actually measured
is returned and logged, so a collision reads as a collision.
Also: LadderWinShare guarded against the conditional (fractal) geometry path,
which fills n from m_fracLegCount while leaving idxList empty - a latent
out-of-bounds on a currently-dead path.
New log lines: mean label lifespan and effective n on the label-cache line, the
required-vs-available horizon per rung, and the grid pair the reconciliation
actually measured (its tolerance now scales with the grid skew instead of a flat
5pp).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 01:12:05 -04:00
|
|
|
"wins %.1f%% (measured at the grid pair %.2f/%.2f, ratio %.2f vs the asked "
|
|
|
|
|
"%.2f), the label cache says Buy %.1f%% | ladder-vs-label gap %.1fpp vs a "
|
|
|
|
|
"%.1fpp tolerance %s. The first number is EXPECTED to be the smallest - it "
|
|
|
|
|
"asks a %d-bar question where the other two ask a %d-bar one.",
|
fix(geometry): the reachability floor measured the WRONG WINDOW - my bug from bc57aca, and it cost real width
RECONCILED: the derivation reported "target reached on 17.7% of bars" while the
label cache reported Buy on 35.9%. Nothing was broken. They measure different
windows, and both are correct:
EXCURSION window ~12 bars (the SWING MEDIAN) - what m_excUpCache accumulates
over. Deliberately short: sizing a barrier off travel
measured over a horizon that itself scales with the barrier
is circular, and it ran away to 14-31*ATR on EURUSD/USDCAD
in 2026-08-07. That guard is correct and stays.
BARRIER horizon 64 bars - what the LABEL walk and the first-passage ladder
run over, and how long the EA actually holds the trade.
So `up >= target` is a 12-bar question and `label == Buy` is a 64-bar one, and
the second can freely exceed the first. TripleBarrierLabel gates the excursion
accumulation on `idx - t <= excWindow` while the barrier walk and the ladder run
the full horizon - the split is explicit and intentional.
THE BUG IS MINE. bc57aca's scale ladder tested reachability with `up[i] >= tp`,
i.e. it asked the 12-bar question about a 64-bar trade. That understates
reachability by ~2x, which is why EVERY wide rung was rejected and the geometry
fell back to the tightest rung at 1.61/3.21. The data supported considerably
wider; the test was just asking the wrong question.
FIX: LadderWinShare() reads the answer off the first-passage ladder - target
touched strictly before the stop, over the full horizon, tie to the stop. That
is the identical question the label walk asks, so the ladder share and the Buy
rate should now agree to within rung discretisation. Both legs snap to the
SMALLEST rung at or above the requested multiple (harder target, harder stop) so
the floor stays conservative.
Expect the scale ladder to select a WIDER rung on the next relabel. On this
data the excursion test read 17.7% at q50 where the true full-horizon share is
35.9%, so rungs that scored 8.1% and 2.8% were likely well above the floor.
ALSO:
- Window reconciliation now PRINTED every derivation: excursion travel share,
ladder win share, and the label cache's Buy share side by side, with the
ladder-vs-label gap flagged if it exceeds rung discretisation. Those two must
agree; if they ever stop agreeing, one of them is wrong and the line says so.
- Renamed tpReach/slReach -> tpTravel/slTravel and relabelled the log line. They
describe the EXCURSION window and are near-tautological there (a q50 stop is
exceeded by ~50% of bars); calling them "reached within the horizon" is what
made the two quantities look like one.
- BARRIER_MIN_TP_REACH_PCT is now BARRIER_MIN_REACH_FRACTION_OF_BE (0.60) x
break-even instead of a hardcoded 20.0. Break-even for 1:RR is 100/(1+RR), so
the absolute floor silently tightened as RR rose - 0.60x at RR=2 but 0.80x at
RR=3, penalising the user for asking for a bigger target. Evaluates to exactly
20.0% at the shipped RR=2, so this is a no-op today and correct if the knob
moves.
NOT COMPILED - user compiles.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 00:51:08 -04:00
|
|
|
slRaw, tpRaw, m_swingMedianBars, 100.0 * excReach / n,
|
fix(labels): overlapping-label sample correction + horizon cap on the scale ladder
Three defects, all surfaced by the 2026-08-17 SP500 H4 run that shipped
stop 4.86 / target 9.71 (width 14.57*ATR, horizon 384).
1. EVERY STANDARD ERROR ASSUMED INDEPENDENT SAMPLES. Triple-barrier labels
started one per bar overlap by the label's lifespan, so n calls are worth
~n/L independent observations (Lopez de Prado, AFML ch. 4 - sample
uniqueness). All three sqrt(p(1-p)/n) sites divided by the RAW count.
The tell: the operating point's null-of-the-maximum gate is family-wise and
should fire on ~5% of eras under the null. Measured fire rates - PAI 47/73
(64%), ConvLSTM 9/24, LSTM 8/21 (38%), CONV 4/62 (6%). CONV, the only model
whose margin distribution admits few bins, sat on the null; the rest cleared
a bar that was too low by ~sqrt(L). PAI's deployed threshold consequently
alternated between the ENDS of its own range era to era (0.10 -> 0.88 ->
0.86 -> 0.66; coverage 16% <-> 73%).
TripleBarrierLabel now records when each label became KNOWABLE - the first
winning touch, or both stops, or the timeout - and the prebuild accumulates
the mean. EffectiveSampleSize() feeds the operating point, the member deploy
gate and the ensemble vote gate. Conservative by construction (n/L is an
upper bound on the damage); gates get harder, never easier.
2. THE SCALE LADDER RAN AWAY, again. Horizon scales as swingMedian*sl*tp, and
since 4d8cb08 reachability is measured OVER that horizon - so a wider rung
buys itself the time that makes it look reachable. Same target -> horizon ->
reach -> target loop the excursion window is kept short to avoid; fixing the
window confusion reopened it through the other door. It walked 128 -> 256 ->
384 bars and stopped at q90, the widest rung there is, with every rung
reading 39-48% against a 20% floor. A floor nothing fails selects nothing.
Rungs whose required horizon exceeds BARRIER_HORIZON_MAX are now rejected -
the same rule ReportGeometryExpectancyScan already applied. It was printing
the shipped pair as CLAMPED and disqualified ('h384!') two lines under the
deriver that chose it: two subsystems, one geometry, opposite verdicts.
3. THE RUNG SNAP DESTROYED THE RATIO IT WAS COMPARING. Both legs snapped
independently to the coarse first-passage grid, re-rating each candidate:
q90 4.86/9.71 -> 5.00/10.00 (2.00), q85 4.07/8.14 -> 5.00/10.00 (IDENTICAL
measurement), q75 3.07/6.13 -> 4.00/6.50 (1.63 - a nearer target). So the
ladder compared win shares taken at ratios from 1.63 to 2.17 and read the
differences as scale. It is why the reach column came out non-monotone in
width (q75 48.5% above q90 42.9%). The stop now snaps to its nearest rung in
log space and the target follows the ratio off it; the pair actually measured
is returned and logged, so a collision reads as a collision.
Also: LadderWinShare guarded against the conditional (fractal) geometry path,
which fills n from m_fracLegCount while leaving idxList empty - a latent
out-of-bounds on a currently-dead path.
New log lines: mean label lifespan and effective n on the label-cache line, the
required-vs-available horizon per rung, and the grid pair the reconciliation
actually measured (its tolerance now scales with the grid skew instead of a flat
5pp).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 01:12:05 -04:00
|
|
|
m_barrierHorizonBars, ladderShare, recSl, recTp,
|
|
|
|
|
(recSl > 0.0 ? recTp / recSl : 0.0), tpRaw / slRaw, buyPct, gap, gapTol,
|
|
|
|
|
(gap <= gapTol ? "(rung discretisation, expected)"
|
fix(geometry): the reachability floor measured the WRONG WINDOW - my bug from bc57aca, and it cost real width
RECONCILED: the derivation reported "target reached on 17.7% of bars" while the
label cache reported Buy on 35.9%. Nothing was broken. They measure different
windows, and both are correct:
EXCURSION window ~12 bars (the SWING MEDIAN) - what m_excUpCache accumulates
over. Deliberately short: sizing a barrier off travel
measured over a horizon that itself scales with the barrier
is circular, and it ran away to 14-31*ATR on EURUSD/USDCAD
in 2026-08-07. That guard is correct and stays.
BARRIER horizon 64 bars - what the LABEL walk and the first-passage ladder
run over, and how long the EA actually holds the trade.
So `up >= target` is a 12-bar question and `label == Buy` is a 64-bar one, and
the second can freely exceed the first. TripleBarrierLabel gates the excursion
accumulation on `idx - t <= excWindow` while the barrier walk and the ladder run
the full horizon - the split is explicit and intentional.
THE BUG IS MINE. bc57aca's scale ladder tested reachability with `up[i] >= tp`,
i.e. it asked the 12-bar question about a 64-bar trade. That understates
reachability by ~2x, which is why EVERY wide rung was rejected and the geometry
fell back to the tightest rung at 1.61/3.21. The data supported considerably
wider; the test was just asking the wrong question.
FIX: LadderWinShare() reads the answer off the first-passage ladder - target
touched strictly before the stop, over the full horizon, tie to the stop. That
is the identical question the label walk asks, so the ladder share and the Buy
rate should now agree to within rung discretisation. Both legs snap to the
SMALLEST rung at or above the requested multiple (harder target, harder stop) so
the floor stays conservative.
Expect the scale ladder to select a WIDER rung on the next relabel. On this
data the excursion test read 17.7% at q50 where the true full-horizon share is
35.9%, so rungs that scored 8.1% and 2.8% were likely well above the floor.
ALSO:
- Window reconciliation now PRINTED every derivation: excursion travel share,
ladder win share, and the label cache's Buy share side by side, with the
ladder-vs-label gap flagged if it exceeds rung discretisation. Those two must
agree; if they ever stop agreeing, one of them is wrong and the line says so.
- Renamed tpReach/slReach -> tpTravel/slTravel and relabelled the log line. They
describe the EXCURSION window and are near-tautological there (a q50 stop is
exceeded by ~50% of bars); calling them "reached within the horizon" is what
made the two quantities look like one.
- BARRIER_MIN_TP_REACH_PCT is now BARRIER_MIN_REACH_FRACTION_OF_BE (0.60) x
break-even instead of a hardcoded 20.0. Break-even for 1:RR is 100/(1+RR), so
the absolute floor silently tightened as RR rose - 0.60x at RR=2 but 0.80x at
RR=3, penalising the user for asking for a bigger target. Evaluates to exactly
20.0% at the shipped RR=2, so this is a no-op today and correct if the knob
moves.
NOT COMPILED - user compiles.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 00:51:08 -04:00
|
|
|
: "<-- TOO LARGE to be discretisation; the ladder and the label walk should"
|
|
|
|
|
" be answering the identical question, so one of them is wrong"),
|
|
|
|
|
m_swingMedianBars, m_barrierHorizonBars));
|
|
|
|
|
}
|
feat: first-passage ladder + expectancy scan - price every geometry, not just the chosen one
Corrects the premise of the previous plan. Break-even is NOT a ceiling.
If the model shifts the win probability on the bars it selects from
p0 = m/(m+k) to p0 + d, then
EV = (p0+d)*k - (1-p0-d)*m = d*(k+m)
because p0*k - (1-p0)*m is zero by construction. The stop:target RATIO
is expectancy-neutral - a punishing break-even is exactly repaid by the
payoff - and only the real edge d and the TOTAL WIDTH (k+m) move EV.
Width matters because the spread is charged once per trade however wide
the barriers are, so a narrow barrier spends much of its own range on
costs. DeriveBarrierGeometry's own comment already said the ratio buys
nothing; the objective just never followed from it.
Blocker this had to solve first: m_excUpCache/m_excDownCache hold only
MAXIMUM travel each way, and a maximum cannot say which side was
reached FIRST - so any geometry other than the walked one was
undecidable on precisely the bars where both barriers were touched,
~28% of the sample.
- BARRIER_LADDER: per bar, the first-touch AGE for 8 travel distances
in each direction, filled during the walk the labels already run.
Cursors keep it O(1) amortised per walked bar rather than 16
comparisons. Levels are travel FROM ENTRY, not barrier prices, so one
ladder serves both directions and the spread is applied analytically
when a level converts back to an SL/TP multiple - storing prices
would need four ladders and bake today's spread into the cache.
Sized, invalidated and validity-gated with the label caches.
- ReportGeometryExpectancyScan: every ladder pair priced exactly off
that cache - width in ATR and in SPREADS (cost efficiency, knowable
without knowing d), break-even, both base rates, the share of bars
resolved inside the horizon, and EV per unit of edge. Compares the
widest resolvable pair against the quantile rule's pick.
MEASUREMENT ONLY - the quantile rule still chooses. Nothing here can
measure d, and width buys nothing if the wider target is less
predictable. Base rates are printed beside each break-even because a
persistent gap is DRIFT and must not be credited to the model.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-10 12:59:18 -04:00
|
|
|
//--- Prices every alternative geometry against the one just chosen. Runs AFTER the pick so the report
|
|
|
|
|
//--- can compare the two, and changes nothing - see its definition for why width, not ratio, is the
|
|
|
|
|
//--- quantity that moves expectancy.
|
|
|
|
|
ReportGeometryExpectancyScan();
|
fix(geometry): the target was small BY CONSTRUCTION - ratio is now policy, scale is measured, ladder ceiling removed
The derivation read the stop from q75 of ADVERSE travel and the target from q50
of FAVOURABLE travel. Over one horizon those distributions are broadly the same
shape, so q75 > q50 MECHANICALLY - the target came out smaller than the stop no
matter what the market did. SP500 H4 shipped stop 3.07 / target 1.70: a 0.55:1
payoff needing 64.3%. That was never a measurement, it was two mismatched
constants.
The reachability line printed beside it - "target on 50.0% of bars, stop on
25.0%" - is exactly 1-q50 and 1-q75. Tautological. It cannot disconfirm
anything, and it read as validation.
WIDTH AND RATIO ARE INDEPENDENT AND ONLY ONE PAYS. EV = edge x width;
ratio is EV-neutral (a driftless walk reaches +m before -k with probability
k/(k+m), which IS break-even). Width is what buys cost efficiency: the spread
is a fixed 0.047*ATR here, so the shipped 4.77*ATR width paid it 21 times per
unit of travel. So:
RATIO = policy. BARRIER_TARGET_RR = 2.0 (user's 1:2). Break-even 33.3%.
SCALE = measured. The stop quantile is chosen from a ladder, WIDEST FIRST,
taking the first rung whose implied 2x target is still reached often
enough to be a trainable class.
That last clause is the difference from the min-reward:risk raise removed in
2026-08-09, which forced target = 2 x stop with NO reachability test, landed on
6.66*ATR reachable on 3.3% of bars, and trained the model to predict something
that essentially never happened. Same ratio; the scale now retreats until the
data says the target is attainable. Every rung is logged.
LADDER CEILING REMOVED. BARRIER_LADDER stopped at 5.00 and the expectancy scan's
"best resolvable pair on width alone" came back as stop 5.05 / target 4.95 - it
pinned to the top rung. A recommendation landing exactly on the edge of its own
search space is a boundary, not a finding: it cannot tell "5 ATR is optimal"
from "5 ATR is all we allowed". Extended to 20*ATR (8 -> 14 rungs). Nothing else
needs editing - every consumer is parameterised by BARRIER_LADDER_COUNT - and
the horizon constraints (decided >= 60%, reachability floor) now bind instead of
a constant.
THE SCAN COULD NOT SEE THE SHIPPED GEOMETRY. ReportBarrierGeometryScan looked
the configured pair up in its integer grid, and DeriveBarrierGeometry produces
CONTINUOUS multiples (3.07/1.70) that can never equal a grid point - so
cfgExcess stayed at its -1.0 sentinel and the report printed "configured 3:2
scores -1.00000", which reads as a catastrophic score and actually means "never
evaluated". Worse, the grid skipped target<stop entirely because it "inverts the
trade's whole premise" - while the derivation was shipping exactly that. The
incumbent is now always scored as a peer (never crowned; it is already in force
and is not an enum pairing the scan could adopt).
BREAK-EVEN NOW INCLUDES THE SPREAD. Every report quoted the frictionless
SL/(SL+TP). On SP500 H4 that read 64.3% while the MEASURED zero-skill rate was
62.1% - a 2.2pp gap that IS the cost, and that made every model look 2.2pp
better than it was. CostAdjustedBreakEvenPct() prices a win at (TP - spread) and
a loss at (SL + spread), matching the expectancy scan's convention exactly so
the two reports cannot disagree.
It also feeds FitDirConfThreshold, which is the correctness half: the operating
point subtracts break-even from precision, so the frictionless figure made every
candidate threshold look better by the width of the spread - 2.2pp against a
measured edge of 2.3pp, i.e. very nearly all of it.
Era line now carries both: "break-even 64.3% frictionless, 66.6% AFTER SPREAD".
Forces a full relabel and retrain. Requested.
NOT COMPILED - user compiles.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 00:34:32 -04:00
|
|
|
//--- The scale ladder already retreats until this floor is met, so reaching here means even its
|
|
|
|
|
//--- tightest rung could not - which is a HORIZON problem, not a ratio problem. Same failure the
|
|
|
|
|
//--- clamped-horizon incident produced, and the ladder report above shows every rung it tried.
|
fix(geometry): the reachability floor measured the WRONG WINDOW - my bug from bc57aca, and it cost real width
RECONCILED: the derivation reported "target reached on 17.7% of bars" while the
label cache reported Buy on 35.9%. Nothing was broken. They measure different
windows, and both are correct:
EXCURSION window ~12 bars (the SWING MEDIAN) - what m_excUpCache accumulates
over. Deliberately short: sizing a barrier off travel
measured over a horizon that itself scales with the barrier
is circular, and it ran away to 14-31*ATR on EURUSD/USDCAD
in 2026-08-07. That guard is correct and stays.
BARRIER horizon 64 bars - what the LABEL walk and the first-passage ladder
run over, and how long the EA actually holds the trade.
So `up >= target` is a 12-bar question and `label == Buy` is a 64-bar one, and
the second can freely exceed the first. TripleBarrierLabel gates the excursion
accumulation on `idx - t <= excWindow` while the barrier walk and the ladder run
the full horizon - the split is explicit and intentional.
THE BUG IS MINE. bc57aca's scale ladder tested reachability with `up[i] >= tp`,
i.e. it asked the 12-bar question about a 64-bar trade. That understates
reachability by ~2x, which is why EVERY wide rung was rejected and the geometry
fell back to the tightest rung at 1.61/3.21. The data supported considerably
wider; the test was just asking the wrong question.
FIX: LadderWinShare() reads the answer off the first-passage ladder - target
touched strictly before the stop, over the full horizon, tie to the stop. That
is the identical question the label walk asks, so the ladder share and the Buy
rate should now agree to within rung discretisation. Both legs snap to the
SMALLEST rung at or above the requested multiple (harder target, harder stop) so
the floor stays conservative.
Expect the scale ladder to select a WIDER rung on the next relabel. On this
data the excursion test read 17.7% at q50 where the true full-horizon share is
35.9%, so rungs that scored 8.1% and 2.8% were likely well above the floor.
ALSO:
- Window reconciliation now PRINTED every derivation: excursion travel share,
ladder win share, and the label cache's Buy share side by side, with the
ladder-vs-label gap flagged if it exceeds rung discretisation. Those two must
agree; if they ever stop agreeing, one of them is wrong and the line says so.
- Renamed tpReach/slReach -> tpTravel/slTravel and relabelled the log line. They
describe the EXCURSION window and are near-tautological there (a q50 stop is
exceeded by ~50% of bars); calling them "reached within the horizon" is what
made the two quantities look like one.
- BARRIER_MIN_TP_REACH_PCT is now BARRIER_MIN_REACH_FRACTION_OF_BE (0.60) x
break-even instead of a hardcoded 20.0. Break-even for 1:RR is 100/(1+RR), so
the absolute floor silently tightened as RR rose - 0.60x at RR=2 but 0.80x at
RR=3, penalising the user for asking for a bigger target. Evaluates to exactly
20.0% at the shipped RR=2, so this is a no-op today and correct if the knob
moves.
NOT COMPILED - user compiles.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 00:51:08 -04:00
|
|
|
if(chosenReach < BARRIER_MIN_TP_REACH_PCT)
|
fix(geometry): the target was small BY CONSTRUCTION - ratio is now policy, scale is measured, ladder ceiling removed
The derivation read the stop from q75 of ADVERSE travel and the target from q50
of FAVOURABLE travel. Over one horizon those distributions are broadly the same
shape, so q75 > q50 MECHANICALLY - the target came out smaller than the stop no
matter what the market did. SP500 H4 shipped stop 3.07 / target 1.70: a 0.55:1
payoff needing 64.3%. That was never a measurement, it was two mismatched
constants.
The reachability line printed beside it - "target on 50.0% of bars, stop on
25.0%" - is exactly 1-q50 and 1-q75. Tautological. It cannot disconfirm
anything, and it read as validation.
WIDTH AND RATIO ARE INDEPENDENT AND ONLY ONE PAYS. EV = edge x width;
ratio is EV-neutral (a driftless walk reaches +m before -k with probability
k/(k+m), which IS break-even). Width is what buys cost efficiency: the spread
is a fixed 0.047*ATR here, so the shipped 4.77*ATR width paid it 21 times per
unit of travel. So:
RATIO = policy. BARRIER_TARGET_RR = 2.0 (user's 1:2). Break-even 33.3%.
SCALE = measured. The stop quantile is chosen from a ladder, WIDEST FIRST,
taking the first rung whose implied 2x target is still reached often
enough to be a trainable class.
That last clause is the difference from the min-reward:risk raise removed in
2026-08-09, which forced target = 2 x stop with NO reachability test, landed on
6.66*ATR reachable on 3.3% of bars, and trained the model to predict something
that essentially never happened. Same ratio; the scale now retreats until the
data says the target is attainable. Every rung is logged.
LADDER CEILING REMOVED. BARRIER_LADDER stopped at 5.00 and the expectancy scan's
"best resolvable pair on width alone" came back as stop 5.05 / target 4.95 - it
pinned to the top rung. A recommendation landing exactly on the edge of its own
search space is a boundary, not a finding: it cannot tell "5 ATR is optimal"
from "5 ATR is all we allowed". Extended to 20*ATR (8 -> 14 rungs). Nothing else
needs editing - every consumer is parameterised by BARRIER_LADDER_COUNT - and
the horizon constraints (decided >= 60%, reachability floor) now bind instead of
a constant.
THE SCAN COULD NOT SEE THE SHIPPED GEOMETRY. ReportBarrierGeometryScan looked
the configured pair up in its integer grid, and DeriveBarrierGeometry produces
CONTINUOUS multiples (3.07/1.70) that can never equal a grid point - so
cfgExcess stayed at its -1.0 sentinel and the report printed "configured 3:2
scores -1.00000", which reads as a catastrophic score and actually means "never
evaluated". Worse, the grid skipped target<stop entirely because it "inverts the
trade's whole premise" - while the derivation was shipping exactly that. The
incumbent is now always scored as a peer (never crowned; it is already in force
and is not an enum pairing the scan could adopt).
BREAK-EVEN NOW INCLUDES THE SPREAD. Every report quoted the frictionless
SL/(SL+TP). On SP500 H4 that read 64.3% while the MEASURED zero-skill rate was
62.1% - a 2.2pp gap that IS the cost, and that made every model look 2.2pp
better than it was. CostAdjustedBreakEvenPct() prices a win at (TP - spread) and
a loss at (SL + spread), matching the expectancy scan's convention exactly so
the two reports cannot disagree.
It also feeds FitDirConfThreshold, which is the correctness half: the operating
point subtracts break-even from precision, so the frictionless figure made every
candidate threshold look better by the width of the spread - 2.2pp against a
measured edge of 2.3pp, i.e. very nearly all of it.
Era line now carries both: "break-even 64.3% frictionless, 66.6% AFTER SPREAD".
Forces a full relabel and retrain. Requested.
NOT COMPILED - user compiles.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 00:34:32 -04:00
|
|
|
Print(ID + StringFormat(": WARNING - the target of %.2f*ATR (%.1f x the %.2f stop) is reached on "
|
|
|
|
|
"only %.1f%% of bars inside the %d-bar horizon, and no rung of the scale "
|
|
|
|
|
"ladder did better. The positive class will be that rare, so expect the "
|
|
|
|
|
"recall floor to bite. Lengthen the horizon or lower BARRIER_TARGET_RR.",
|
fix(geometry): the reachability floor measured the WRONG WINDOW - my bug from bc57aca, and it cost real width
RECONCILED: the derivation reported "target reached on 17.7% of bars" while the
label cache reported Buy on 35.9%. Nothing was broken. They measure different
windows, and both are correct:
EXCURSION window ~12 bars (the SWING MEDIAN) - what m_excUpCache accumulates
over. Deliberately short: sizing a barrier off travel
measured over a horizon that itself scales with the barrier
is circular, and it ran away to 14-31*ATR on EURUSD/USDCAD
in 2026-08-07. That guard is correct and stays.
BARRIER horizon 64 bars - what the LABEL walk and the first-passage ladder
run over, and how long the EA actually holds the trade.
So `up >= target` is a 12-bar question and `label == Buy` is a 64-bar one, and
the second can freely exceed the first. TripleBarrierLabel gates the excursion
accumulation on `idx - t <= excWindow` while the barrier walk and the ladder run
the full horizon - the split is explicit and intentional.
THE BUG IS MINE. bc57aca's scale ladder tested reachability with `up[i] >= tp`,
i.e. it asked the 12-bar question about a 64-bar trade. That understates
reachability by ~2x, which is why EVERY wide rung was rejected and the geometry
fell back to the tightest rung at 1.61/3.21. The data supported considerably
wider; the test was just asking the wrong question.
FIX: LadderWinShare() reads the answer off the first-passage ladder - target
touched strictly before the stop, over the full horizon, tie to the stop. That
is the identical question the label walk asks, so the ladder share and the Buy
rate should now agree to within rung discretisation. Both legs snap to the
SMALLEST rung at or above the requested multiple (harder target, harder stop) so
the floor stays conservative.
Expect the scale ladder to select a WIDER rung on the next relabel. On this
data the excursion test read 17.7% at q50 where the true full-horizon share is
35.9%, so rungs that scored 8.1% and 2.8% were likely well above the floor.
ALSO:
- Window reconciliation now PRINTED every derivation: excursion travel share,
ladder win share, and the label cache's Buy share side by side, with the
ladder-vs-label gap flagged if it exceeds rung discretisation. Those two must
agree; if they ever stop agreeing, one of them is wrong and the line says so.
- Renamed tpReach/slReach -> tpTravel/slTravel and relabelled the log line. They
describe the EXCURSION window and are near-tautological there (a q50 stop is
exceeded by ~50% of bars); calling them "reached within the horizon" is what
made the two quantities look like one.
- BARRIER_MIN_TP_REACH_PCT is now BARRIER_MIN_REACH_FRACTION_OF_BE (0.60) x
break-even instead of a hardcoded 20.0. Break-even for 1:RR is 100/(1+RR), so
the absolute floor silently tightened as RR rose - 0.60x at RR=2 but 0.80x at
RR=3, penalising the user for asking for a bigger target. Evaluates to exactly
20.0% at the shipped RR=2, so this is a no-op today and correct if the knob
moves.
NOT COMPILED - user compiles.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 00:51:08 -04:00
|
|
|
m_derivedTpMult, BARRIER_TARGET_RR, m_derivedSlMult, chosenReach,
|
fix(geometry): the target was small BY CONSTRUCTION - ratio is now policy, scale is measured, ladder ceiling removed
The derivation read the stop from q75 of ADVERSE travel and the target from q50
of FAVOURABLE travel. Over one horizon those distributions are broadly the same
shape, so q75 > q50 MECHANICALLY - the target came out smaller than the stop no
matter what the market did. SP500 H4 shipped stop 3.07 / target 1.70: a 0.55:1
payoff needing 64.3%. That was never a measurement, it was two mismatched
constants.
The reachability line printed beside it - "target on 50.0% of bars, stop on
25.0%" - is exactly 1-q50 and 1-q75. Tautological. It cannot disconfirm
anything, and it read as validation.
WIDTH AND RATIO ARE INDEPENDENT AND ONLY ONE PAYS. EV = edge x width;
ratio is EV-neutral (a driftless walk reaches +m before -k with probability
k/(k+m), which IS break-even). Width is what buys cost efficiency: the spread
is a fixed 0.047*ATR here, so the shipped 4.77*ATR width paid it 21 times per
unit of travel. So:
RATIO = policy. BARRIER_TARGET_RR = 2.0 (user's 1:2). Break-even 33.3%.
SCALE = measured. The stop quantile is chosen from a ladder, WIDEST FIRST,
taking the first rung whose implied 2x target is still reached often
enough to be a trainable class.
That last clause is the difference from the min-reward:risk raise removed in
2026-08-09, which forced target = 2 x stop with NO reachability test, landed on
6.66*ATR reachable on 3.3% of bars, and trained the model to predict something
that essentially never happened. Same ratio; the scale now retreats until the
data says the target is attainable. Every rung is logged.
LADDER CEILING REMOVED. BARRIER_LADDER stopped at 5.00 and the expectancy scan's
"best resolvable pair on width alone" came back as stop 5.05 / target 4.95 - it
pinned to the top rung. A recommendation landing exactly on the edge of its own
search space is a boundary, not a finding: it cannot tell "5 ATR is optimal"
from "5 ATR is all we allowed". Extended to 20*ATR (8 -> 14 rungs). Nothing else
needs editing - every consumer is parameterised by BARRIER_LADDER_COUNT - and
the horizon constraints (decided >= 60%, reachability floor) now bind instead of
a constant.
THE SCAN COULD NOT SEE THE SHIPPED GEOMETRY. ReportBarrierGeometryScan looked
the configured pair up in its integer grid, and DeriveBarrierGeometry produces
CONTINUOUS multiples (3.07/1.70) that can never equal a grid point - so
cfgExcess stayed at its -1.0 sentinel and the report printed "configured 3:2
scores -1.00000", which reads as a catastrophic score and actually means "never
evaluated". Worse, the grid skipped target<stop entirely because it "inverts the
trade's whole premise" - while the derivation was shipping exactly that. The
incumbent is now always scored as a peer (never crowned; it is already in force
and is not an enum pairing the scan could adopt).
BREAK-EVEN NOW INCLUDES THE SPREAD. Every report quoted the frictionless
SL/(SL+TP). On SP500 H4 that read 64.3% while the MEASURED zero-skill rate was
62.1% - a 2.2pp gap that IS the cost, and that made every model look 2.2pp
better than it was. CostAdjustedBreakEvenPct() prices a win at (TP - spread) and
a loss at (SL + spread), matching the expectancy scan's convention exactly so
the two reports cannot disagree.
It also feeds FitDirConfThreshold, which is the correctness half: the operating
point subtracts break-even from precision, so the frictionless figure made every
candidate threshold look better by the width of the spread - 2.2pp against a
measured edge of 2.3pp, i.e. very nearly all of it.
Era line now carries both: "break-even 64.3% frictionless, 66.6% AFTER SPREAD".
Forces a full relabel and retrain. Requested.
NOT COMPILED - user compiles.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 00:34:32 -04:00
|
|
|
m_barrierHorizonBars));
|
feat: derive the ATR multiples from measured excursions - no hardcoded geometry
The barrier was still two constants. SL_Mode/TP_Mode left the Inputs tab in
3482b6c, but the fallback was a hardcoded 2:6 and the geometry scan only ever
chose from a hardcoded grid {2,3} x {2,3,4,6,8,10}. Picking the least-bad of
eleven guesses is not deriving anything.
WHY THE SCAN WAS THE WRONG INSTRUMENT, now measurable rather than argued. It
ranks pairings by how predictable their OUTCOME is - a question about direction.
The excursion test (2c78f3b) ran on SP500 H1 and direction is the one thing
absent: ASYMMETRY p=0.0846, against RANGE/UP/DOWN all at p=0.0050, with RANGE
scoring 0.01345 vs a 0.00343 null - 4x, where the barrier label sits at 1.01x.
Hence the scan failing its own gate on every run, and its "winner" wandering
2:8 -> 3:8 -> 2:8 -> 2:4 across four runs of the same data. Excursion SIZE is
strongly measurable, so derive the geometry from that instead.
stop = q25 of measured ADVERSE travel (ordinary noise does not reach it)
target = q50 of measured FAVOURABLE travel (reached ~half the time, by
construction, inside the horizon)
Continuous, in ATR units, superseding the enum multiples. Reachability ("target
on X% of bars, stop on Y%") and the implied break-even are printed so the choice
is auditable rather than trusted.
FIXED-POINT ITERATION, not one-shot. ComputeBarrierHorizonBars scales the
horizon with the target (first-passage time grows with the band) and the
excursions are measured OVER the horizon, so target -> horizon -> excursions ->
target is a real loop - deriving once sizes the target from travel measured
under the PREVIOUS horizon. Re-measures until the multiples move <5%, capped at
3 passes, and says so if it does not settle.
Does NOT create expectancy, and the log says as much: chance precision equals
break-even at every geometry (m/(m+k) on both sides). It buys a target the
market reaches and a stop that survives noise. Where Min_Risk_Reward_Ratio
forces a target the market rarely reaches, it WARNS rather than overriding -
the ratio is the user's risk policy, so the honest move is to state its cost.
That is the collision that once rejected 100% of setups.
Pinned in the .cfg as doubles appended AFTER this morning's two ints, so .cfg
files written earlier today still load (their length guard finds no doubles) and
a model that carries them was trained on them and never re-derives.
Also fixes a message from e5ceed6 that claimed "this model resumed from disk"
unconditionally - it printed above a "seeding era 0" line on a brand-new model,
because the branch fires whenever the cache is not built, which is equally true
before a fresh model's first prebuild. A diagnostic that misreports its own
trigger is worse than one that says nothing: it gets quoted back as evidence.
FORCES A FULL RETRAIN (labels change).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-07 12:06:25 -04:00
|
|
|
return true;
|
|
|
|
|
}
|
|
|
|
|
//+------------------------------------------------------------------+
|
fix(labels): overlapping-label sample correction + horizon cap on the scale ladder
Three defects, all surfaced by the 2026-08-17 SP500 H4 run that shipped
stop 4.86 / target 9.71 (width 14.57*ATR, horizon 384).
1. EVERY STANDARD ERROR ASSUMED INDEPENDENT SAMPLES. Triple-barrier labels
started one per bar overlap by the label's lifespan, so n calls are worth
~n/L independent observations (Lopez de Prado, AFML ch. 4 - sample
uniqueness). All three sqrt(p(1-p)/n) sites divided by the RAW count.
The tell: the operating point's null-of-the-maximum gate is family-wise and
should fire on ~5% of eras under the null. Measured fire rates - PAI 47/73
(64%), ConvLSTM 9/24, LSTM 8/21 (38%), CONV 4/62 (6%). CONV, the only model
whose margin distribution admits few bins, sat on the null; the rest cleared
a bar that was too low by ~sqrt(L). PAI's deployed threshold consequently
alternated between the ENDS of its own range era to era (0.10 -> 0.88 ->
0.86 -> 0.66; coverage 16% <-> 73%).
TripleBarrierLabel now records when each label became KNOWABLE - the first
winning touch, or both stops, or the timeout - and the prebuild accumulates
the mean. EffectiveSampleSize() feeds the operating point, the member deploy
gate and the ensemble vote gate. Conservative by construction (n/L is an
upper bound on the damage); gates get harder, never easier.
2. THE SCALE LADDER RAN AWAY, again. Horizon scales as swingMedian*sl*tp, and
since 4d8cb08 reachability is measured OVER that horizon - so a wider rung
buys itself the time that makes it look reachable. Same target -> horizon ->
reach -> target loop the excursion window is kept short to avoid; fixing the
window confusion reopened it through the other door. It walked 128 -> 256 ->
384 bars and stopped at q90, the widest rung there is, with every rung
reading 39-48% against a 20% floor. A floor nothing fails selects nothing.
Rungs whose required horizon exceeds BARRIER_HORIZON_MAX are now rejected -
the same rule ReportGeometryExpectancyScan already applied. It was printing
the shipped pair as CLAMPED and disqualified ('h384!') two lines under the
deriver that chose it: two subsystems, one geometry, opposite verdicts.
3. THE RUNG SNAP DESTROYED THE RATIO IT WAS COMPARING. Both legs snapped
independently to the coarse first-passage grid, re-rating each candidate:
q90 4.86/9.71 -> 5.00/10.00 (2.00), q85 4.07/8.14 -> 5.00/10.00 (IDENTICAL
measurement), q75 3.07/6.13 -> 4.00/6.50 (1.63 - a nearer target). So the
ladder compared win shares taken at ratios from 1.63 to 2.17 and read the
differences as scale. It is why the reach column came out non-monotone in
width (q75 48.5% above q90 42.9%). The stop now snaps to its nearest rung in
log space and the target follows the ratio off it; the pair actually measured
is returned and logged, so a collision reads as a collision.
Also: LadderWinShare guarded against the conditional (fractal) geometry path,
which fills n from m_fracLegCount while leaving idxList empty - a latent
out-of-bounds on a currently-dead path.
New log lines: mean label lifespan and effective n on the label-cache line, the
required-vs-available horizon per rung, and the grid pair the reconciliation
actually measured (its tolerance now scales with the grid skew instead of a flat
5pp).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 01:12:05 -04:00
|
|
|
//| Mean bars-to-resolution over the label cache. 1.0 until something |
|
|
|
|
|
//| has been measured, which makes EffectiveSampleSize() the identity |
|
|
|
|
|
//| - the pre-2026-08-17 behaviour. That default is deliberate: an |
|
|
|
|
|
//| UNMEASURED overlap must not silently shrink anyone's sample, so |
|
|
|
|
|
//| the correction switches itself on only once it has evidence. |
|
|
|
|
|
//+------------------------------------------------------------------+
|
|
|
|
|
double CExpertSignalAIBase::MeanLabelLifespan(void)
|
|
|
|
|
{
|
|
|
|
|
if(m_labelLifespanCount <= 0 || m_labelLifespanSum <= 0.0)
|
|
|
|
|
return 1.0;
|
|
|
|
|
double mean = m_labelLifespanSum / (double)m_labelLifespanCount;
|
|
|
|
|
//--- Cannot exceed the window it was measured in, and cannot be shorter than one bar.
|
|
|
|
|
double cap = (double)MathMax(m_barrierHorizonBars, 1);
|
|
|
|
|
return MathMax(1.0, MathMin(mean, cap));
|
|
|
|
|
}
|
|
|
|
|
//+------------------------------------------------------------------+
|
|
|
|
|
//| Independent observations behind `rawN` overlapping labels. |
|
|
|
|
|
//| |
|
|
|
|
|
//| Triple-barrier labels started one per bar with mean lifespan L |
|
|
|
|
|
//| have average concurrency ~L, hence average uniqueness ~1/L, hence |
|
|
|
|
|
//| n_eff = n/L (Lopez de Prado, AFML ch. 4 - sample uniqueness and |
|
|
|
|
|
//| the sequential bootstrap). Every sqrt(p(1-p)/n) in this codebase |
|
|
|
|
|
//| assumed L = 1, which on SP500 H4 at a 384-bar horizon understated |
|
|
|
|
|
//| every standard error by roughly sqrt(L). See m_lastLabelLifespan |
|
|
|
|
|
//| for the run that exposed it. |
|
|
|
|
|
//| |
|
fix(labels): correct EffectiveSampleSize clamp order, share the horizon ladder, retract a false justification
Self-review of 1540ba8 against the FULL 6,930-era log rather than the first
three minutes of it. Three corrections.
1. EffectiveSampleSize() clamped in the wrong order. MathMax(2, MathMin(eff,
rawN)) returns 2 when rawN is 1 - an effective sample LARGER than the raw
one, shrinking the SE in exactly the direction the function exists to
prevent. Floor first, cap at rawN last.
2. The horizon cap rejected on the CEILING only, and said so as though that
made the label untruncated. It does not: the horizon ladder also snaps DOWN,
so a pair needing 317 bars is granted 256 and is silently truncated without
ever being flagged CLAMPED. Added SnapHorizonToLadder() / GrantedHorizonBars()
and the scale ladder now reports "needs N gets M" per rung. Rejection stays on
the ceiling alone - matching ReportGeometryExpectancyScan's '!' exactly, which
was the point - because rejecting on the snap-down would select rungs for
landing just above a ladder point rather than for anything about the market.
ComputeBarrierHorizonBars' private copy of the ladder is gone; there is now
one copy, which is the whole reason RequiredHorizonBars was factored out.
3. RETRACTED THE JUSTIFICATION IN 1540ba8's COMMENTS. That commit claimed the
overlap correction was needed because the operating point's null-of-the-
maximum gate fired on 47/73 Perceptron eras (64%) where a family-wise test
should fire on ~5%. Those 73 fits were the first three minutes of a
six-and-a-half-hour run. Over the full run:
PAI 47/3214 = 1.5% HYB 30/1200 = 2.5%
CONV 4/63 = 6.3% LSTM 75/915 = 8.2%
All at or below the null. The gate from 7414570 is working as designed and
PAI's 47 clears were a cold-start transient never repeated in 3,141 later
fits; its threshold over the run's second half has sd 0.01. The overlap
correction is still right - sqrt(p(1-p)/n) on overlapping labels is the wrong
formula - but it fixes no observed failure, and it costs nothing today
because no model is near the deploy line.
WHAT THE FULL RUN DOES CONFIRM, unchanged: the geometry ran away exactly as
described (2.00/6.00 h128 -> 3.49/6.99 h256 -> 4.86/9.71 h384, three passes,
stopping at q90 because the quantile ladder ended), the label stayed long-skewed
at Buy 42.9% / Sell 22.6%, and no checkpoint on any of the four models ever
cleared the deployability floor. Pooled declustered win rates: PAI 31.70%,
HYB 31.02%, LSTM 32.69%, CONV 31.57% - every one 4-6pp below the 37% always-long
chance rate and 1-2.7pp below the 33.7% cost-adjusted break-even.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 07:30:19 -04:00
|
|
|
//| ORDER OF THE CLAMPS MATTERS, and the first version had it wrong: |
|
|
|
|
|
//| MathMax(2, MathMin(eff, rawN)) returns 2 when rawN is 1, i.e. an |
|
|
|
|
|
//| effective sample LARGER than the raw one, which shrinks the SE in |
|
|
|
|
|
//| exactly the direction this function exists to prevent. The cap at |
|
|
|
|
|
//| rawN has to be applied LAST, so a floor can never manufacture |
|
|
|
|
|
//| observations that were not there. |
|
fix(labels): overlapping-label sample correction + horizon cap on the scale ladder
Three defects, all surfaced by the 2026-08-17 SP500 H4 run that shipped
stop 4.86 / target 9.71 (width 14.57*ATR, horizon 384).
1. EVERY STANDARD ERROR ASSUMED INDEPENDENT SAMPLES. Triple-barrier labels
started one per bar overlap by the label's lifespan, so n calls are worth
~n/L independent observations (Lopez de Prado, AFML ch. 4 - sample
uniqueness). All three sqrt(p(1-p)/n) sites divided by the RAW count.
The tell: the operating point's null-of-the-maximum gate is family-wise and
should fire on ~5% of eras under the null. Measured fire rates - PAI 47/73
(64%), ConvLSTM 9/24, LSTM 8/21 (38%), CONV 4/62 (6%). CONV, the only model
whose margin distribution admits few bins, sat on the null; the rest cleared
a bar that was too low by ~sqrt(L). PAI's deployed threshold consequently
alternated between the ENDS of its own range era to era (0.10 -> 0.88 ->
0.86 -> 0.66; coverage 16% <-> 73%).
TripleBarrierLabel now records when each label became KNOWABLE - the first
winning touch, or both stops, or the timeout - and the prebuild accumulates
the mean. EffectiveSampleSize() feeds the operating point, the member deploy
gate and the ensemble vote gate. Conservative by construction (n/L is an
upper bound on the damage); gates get harder, never easier.
2. THE SCALE LADDER RAN AWAY, again. Horizon scales as swingMedian*sl*tp, and
since 4d8cb08 reachability is measured OVER that horizon - so a wider rung
buys itself the time that makes it look reachable. Same target -> horizon ->
reach -> target loop the excursion window is kept short to avoid; fixing the
window confusion reopened it through the other door. It walked 128 -> 256 ->
384 bars and stopped at q90, the widest rung there is, with every rung
reading 39-48% against a 20% floor. A floor nothing fails selects nothing.
Rungs whose required horizon exceeds BARRIER_HORIZON_MAX are now rejected -
the same rule ReportGeometryExpectancyScan already applied. It was printing
the shipped pair as CLAMPED and disqualified ('h384!') two lines under the
deriver that chose it: two subsystems, one geometry, opposite verdicts.
3. THE RUNG SNAP DESTROYED THE RATIO IT WAS COMPARING. Both legs snapped
independently to the coarse first-passage grid, re-rating each candidate:
q90 4.86/9.71 -> 5.00/10.00 (2.00), q85 4.07/8.14 -> 5.00/10.00 (IDENTICAL
measurement), q75 3.07/6.13 -> 4.00/6.50 (1.63 - a nearer target). So the
ladder compared win shares taken at ratios from 1.63 to 2.17 and read the
differences as scale. It is why the reach column came out non-monotone in
width (q75 48.5% above q90 42.9%). The stop now snaps to its nearest rung in
log space and the target follows the ratio off it; the pair actually measured
is returned and logged, so a collision reads as a collision.
Also: LadderWinShare guarded against the conditional (fractal) geometry path,
which fills n from m_fracLegCount while leaving idxList empty - a latent
out-of-bounds on a currently-dead path.
New log lines: mean label lifespan and effective n on the label-cache line, the
required-vs-available horizon per rung, and the grid pair the reconciliation
actually measured (its tolerance now scales with the grid skew instead of a flat
5pp).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 01:12:05 -04:00
|
|
|
//| |
|
|
|
|
|
//| THIS IS CONSERVATIVE, and knowingly so. n/L is the sample-size |
|
|
|
|
|
//| treatment AFML prescribes, but it is an upper bound on the damage:|
|
|
|
|
|
//| two labels sharing 99% of their outcome window are highly |
|
|
|
|
|
//| correlated, not identical - they enter at different prices, so |
|
|
|
|
|
//| one can win where the other loses. The true effective sample sits |
|
|
|
|
|
//| somewhere between n/L and n, and nothing here measures where. |
|
|
|
|
|
//| Erring toward n/L means gates get HARDER to clear, never easier, |
|
|
|
|
|
//| which is the safe direction on a funded account and the opposite |
|
|
|
|
|
//| of the error this replaces. Expect the operating point to sit at |
|
|
|
|
|
//| its deterministic fallback far more often and the deploy gate to |
|
|
|
|
|
//| reject eras it used to pass; that is the correction working, not |
|
|
|
|
|
//| a regression. If it proves too strict, the honest refinement is |
|
|
|
|
|
//| to MEASURE average uniqueness per label (AFML 4.2) rather than to |
|
|
|
|
|
//| soften the divisor by taste. |
|
|
|
|
|
//+------------------------------------------------------------------+
|
|
|
|
|
double CExpertSignalAIBase::EffectiveSampleSize(double rawN)
|
|
|
|
|
{
|
|
|
|
|
if(rawN <= 0.0)
|
|
|
|
|
return 0.0;
|
|
|
|
|
double eff = rawN / MeanLabelLifespan();
|
fix(labels): correct EffectiveSampleSize clamp order, share the horizon ladder, retract a false justification
Self-review of 1540ba8 against the FULL 6,930-era log rather than the first
three minutes of it. Three corrections.
1. EffectiveSampleSize() clamped in the wrong order. MathMax(2, MathMin(eff,
rawN)) returns 2 when rawN is 1 - an effective sample LARGER than the raw
one, shrinking the SE in exactly the direction the function exists to
prevent. Floor first, cap at rawN last.
2. The horizon cap rejected on the CEILING only, and said so as though that
made the label untruncated. It does not: the horizon ladder also snaps DOWN,
so a pair needing 317 bars is granted 256 and is silently truncated without
ever being flagged CLAMPED. Added SnapHorizonToLadder() / GrantedHorizonBars()
and the scale ladder now reports "needs N gets M" per rung. Rejection stays on
the ceiling alone - matching ReportGeometryExpectancyScan's '!' exactly, which
was the point - because rejecting on the snap-down would select rungs for
landing just above a ladder point rather than for anything about the market.
ComputeBarrierHorizonBars' private copy of the ladder is gone; there is now
one copy, which is the whole reason RequiredHorizonBars was factored out.
3. RETRACTED THE JUSTIFICATION IN 1540ba8's COMMENTS. That commit claimed the
overlap correction was needed because the operating point's null-of-the-
maximum gate fired on 47/73 Perceptron eras (64%) where a family-wise test
should fire on ~5%. Those 73 fits were the first three minutes of a
six-and-a-half-hour run. Over the full run:
PAI 47/3214 = 1.5% HYB 30/1200 = 2.5%
CONV 4/63 = 6.3% LSTM 75/915 = 8.2%
All at or below the null. The gate from 7414570 is working as designed and
PAI's 47 clears were a cold-start transient never repeated in 3,141 later
fits; its threshold over the run's second half has sd 0.01. The overlap
correction is still right - sqrt(p(1-p)/n) on overlapping labels is the wrong
formula - but it fixes no observed failure, and it costs nothing today
because no model is near the deploy line.
WHAT THE FULL RUN DOES CONFIRM, unchanged: the geometry ran away exactly as
described (2.00/6.00 h128 -> 3.49/6.99 h256 -> 4.86/9.71 h384, three passes,
stopping at q90 because the quantile ladder ended), the label stayed long-skewed
at Buy 42.9% / Sell 22.6%, and no checkpoint on any of the four models ever
cleared the deployability floor. Pooled declustered win rates: PAI 31.70%,
HYB 31.02%, LSTM 32.69%, CONV 31.57% - every one 4-6pp below the 37% always-long
chance rate and 1-2.7pp below the 33.7% cost-adjusted break-even.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 07:30:19 -04:00
|
|
|
//--- Floor, so a caller dividing by it cannot hit zero...
|
|
|
|
|
if(eff < 2.0)
|
|
|
|
|
eff = 2.0;
|
|
|
|
|
//--- ...then the cap, LAST, so the floor can never exceed the observations that actually exist.
|
|
|
|
|
return MathMin(eff, rawN);
|
fix(labels): overlapping-label sample correction + horizon cap on the scale ladder
Three defects, all surfaced by the 2026-08-17 SP500 H4 run that shipped
stop 4.86 / target 9.71 (width 14.57*ATR, horizon 384).
1. EVERY STANDARD ERROR ASSUMED INDEPENDENT SAMPLES. Triple-barrier labels
started one per bar overlap by the label's lifespan, so n calls are worth
~n/L independent observations (Lopez de Prado, AFML ch. 4 - sample
uniqueness). All three sqrt(p(1-p)/n) sites divided by the RAW count.
The tell: the operating point's null-of-the-maximum gate is family-wise and
should fire on ~5% of eras under the null. Measured fire rates - PAI 47/73
(64%), ConvLSTM 9/24, LSTM 8/21 (38%), CONV 4/62 (6%). CONV, the only model
whose margin distribution admits few bins, sat on the null; the rest cleared
a bar that was too low by ~sqrt(L). PAI's deployed threshold consequently
alternated between the ENDS of its own range era to era (0.10 -> 0.88 ->
0.86 -> 0.66; coverage 16% <-> 73%).
TripleBarrierLabel now records when each label became KNOWABLE - the first
winning touch, or both stops, or the timeout - and the prebuild accumulates
the mean. EffectiveSampleSize() feeds the operating point, the member deploy
gate and the ensemble vote gate. Conservative by construction (n/L is an
upper bound on the damage); gates get harder, never easier.
2. THE SCALE LADDER RAN AWAY, again. Horizon scales as swingMedian*sl*tp, and
since 4d8cb08 reachability is measured OVER that horizon - so a wider rung
buys itself the time that makes it look reachable. Same target -> horizon ->
reach -> target loop the excursion window is kept short to avoid; fixing the
window confusion reopened it through the other door. It walked 128 -> 256 ->
384 bars and stopped at q90, the widest rung there is, with every rung
reading 39-48% against a 20% floor. A floor nothing fails selects nothing.
Rungs whose required horizon exceeds BARRIER_HORIZON_MAX are now rejected -
the same rule ReportGeometryExpectancyScan already applied. It was printing
the shipped pair as CLAMPED and disqualified ('h384!') two lines under the
deriver that chose it: two subsystems, one geometry, opposite verdicts.
3. THE RUNG SNAP DESTROYED THE RATIO IT WAS COMPARING. Both legs snapped
independently to the coarse first-passage grid, re-rating each candidate:
q90 4.86/9.71 -> 5.00/10.00 (2.00), q85 4.07/8.14 -> 5.00/10.00 (IDENTICAL
measurement), q75 3.07/6.13 -> 4.00/6.50 (1.63 - a nearer target). So the
ladder compared win shares taken at ratios from 1.63 to 2.17 and read the
differences as scale. It is why the reach column came out non-monotone in
width (q75 48.5% above q90 42.9%). The stop now snaps to its nearest rung in
log space and the target follows the ratio off it; the pair actually measured
is returned and logged, so a collision reads as a collision.
Also: LadderWinShare guarded against the conditional (fractal) geometry path,
which fills n from m_fracLegCount while leaving idxList empty - a latent
out-of-bounds on a currently-dead path.
New log lines: mean label lifespan and effective n on the label-cache line, the
required-vs-available horizon per rung, and the grid pair the reconciliation
actually measured (its tolerance now scales with the grid skew instead of a flat
5pp).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 01:12:05 -04:00
|
|
|
}
|
|
|
|
|
//+------------------------------------------------------------------+
|
|
|
|
|
//| Bars this geometry needs before its label stops being truncated. |
|
|
|
|
|
//| IDENTICAL arithmetic to ComputeBarrierHorizonBars() - see the |
|
|
|
|
|
//| declaration for the 2026-08-17 divergence that made factoring it |
|
|
|
|
|
//| out necessary rather than tidy. |
|
|
|
|
|
//+------------------------------------------------------------------+
|
|
|
|
|
int CExpertSignalAIBase::RequiredHorizonBars(double slMult, double tpMult)
|
|
|
|
|
{
|
|
|
|
|
double swing = (double)MathMax(m_swingMedianBars, 1);
|
|
|
|
|
if(slMult <= 0.0 || tpMult <= 0.0)
|
|
|
|
|
return BARRIER_HORIZON_MIN;
|
|
|
|
|
return (int)MathRound(swing * slMult * tpMult);
|
|
|
|
|
}
|
|
|
|
|
//+------------------------------------------------------------------+
|
fix(labels): correct EffectiveSampleSize clamp order, share the horizon ladder, retract a false justification
Self-review of 1540ba8 against the FULL 6,930-era log rather than the first
three minutes of it. Three corrections.
1. EffectiveSampleSize() clamped in the wrong order. MathMax(2, MathMin(eff,
rawN)) returns 2 when rawN is 1 - an effective sample LARGER than the raw
one, shrinking the SE in exactly the direction the function exists to
prevent. Floor first, cap at rawN last.
2. The horizon cap rejected on the CEILING only, and said so as though that
made the label untruncated. It does not: the horizon ladder also snaps DOWN,
so a pair needing 317 bars is granted 256 and is silently truncated without
ever being flagged CLAMPED. Added SnapHorizonToLadder() / GrantedHorizonBars()
and the scale ladder now reports "needs N gets M" per rung. Rejection stays on
the ceiling alone - matching ReportGeometryExpectancyScan's '!' exactly, which
was the point - because rejecting on the snap-down would select rungs for
landing just above a ladder point rather than for anything about the market.
ComputeBarrierHorizonBars' private copy of the ladder is gone; there is now
one copy, which is the whole reason RequiredHorizonBars was factored out.
3. RETRACTED THE JUSTIFICATION IN 1540ba8's COMMENTS. That commit claimed the
overlap correction was needed because the operating point's null-of-the-
maximum gate fired on 47/73 Perceptron eras (64%) where a family-wise test
should fire on ~5%. Those 73 fits were the first three minutes of a
six-and-a-half-hour run. Over the full run:
PAI 47/3214 = 1.5% HYB 30/1200 = 2.5%
CONV 4/63 = 6.3% LSTM 75/915 = 8.2%
All at or below the null. The gate from 7414570 is working as designed and
PAI's 47 clears were a cold-start transient never repeated in 3,141 later
fits; its threshold over the run's second half has sd 0.01. The overlap
correction is still right - sqrt(p(1-p)/n) on overlapping labels is the wrong
formula - but it fixes no observed failure, and it costs nothing today
because no model is near the deploy line.
WHAT THE FULL RUN DOES CONFIRM, unchanged: the geometry ran away exactly as
described (2.00/6.00 h128 -> 3.49/6.99 h256 -> 4.86/9.71 h384, three passes,
stopping at q90 because the quantile ladder ended), the label stayed long-skewed
at Buy 42.9% / Sell 22.6%, and no checkpoint on any of the four models ever
cleared the deployability floor. Pooled declustered win rates: PAI 31.70%,
HYB 31.02%, LSTM 32.69%, CONV 31.57% - every one 4-6pp below the 37% always-long
chance rate and 1-2.7pp below the 33.7% cost-adjusted break-even.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 07:30:19 -04:00
|
|
|
//| THE horizon ladder, and the only copy of it. It used to be a |
|
|
|
|
|
//| local array inside ComputeBarrierHorizonBars(); the scale ladder |
|
|
|
|
|
//| needs the same snap to report what a rung would actually be |
|
|
|
|
|
//| granted, and a second copy is precisely the drift that let the |
|
|
|
|
|
//| deriver and the expectancy scan disagree about one geometry. |
|
|
|
|
|
//| |
|
|
|
|
|
//| Snaps DOWN, matching ComputeFirstLayerWidth()'s direction: a |
|
|
|
|
|
//| horizon shorter than measured makes the label stricter (more |
|
|
|
|
|
//| Neutral), never more permissive. |
|
|
|
|
|
//+------------------------------------------------------------------+
|
|
|
|
|
int CExpertSignalAIBase::SnapHorizonToLadder(int rawBars)
|
|
|
|
|
{
|
|
|
|
|
int ladder[BARRIER_HORIZON_LADDER_COUNT] = { 12, 16, 24, 32, 48, 64, 96, 128, 192, 256, 384 };
|
|
|
|
|
int raw = rawBars;
|
|
|
|
|
if(raw < BARRIER_HORIZON_MIN)
|
|
|
|
|
raw = BARRIER_HORIZON_MIN;
|
|
|
|
|
if(raw > BARRIER_HORIZON_MAX)
|
|
|
|
|
raw = BARRIER_HORIZON_MAX;
|
|
|
|
|
int snapped = ladder[0];
|
|
|
|
|
for(int k = 0; k < BARRIER_HORIZON_LADDER_COUNT; k++)
|
|
|
|
|
if(ladder[k] <= raw)
|
|
|
|
|
snapped = ladder[k];
|
|
|
|
|
return snapped;
|
|
|
|
|
}
|
|
|
|
|
//+------------------------------------------------------------------+
|
|
|
|
|
//| What this pair would actually be labelled under - see the |
|
|
|
|
|
//| declaration for why this is NOT RequiredHorizonBars(). |
|
|
|
|
|
//+------------------------------------------------------------------+
|
|
|
|
|
int CExpertSignalAIBase::GrantedHorizonBars(double slMult, double tpMult)
|
|
|
|
|
{
|
|
|
|
|
return SnapHorizonToLadder(RequiredHorizonBars(slMult, tpMult));
|
|
|
|
|
}
|
|
|
|
|
//+------------------------------------------------------------------+
|
fix(geometry): the target was small BY CONSTRUCTION - ratio is now policy, scale is measured, ladder ceiling removed
The derivation read the stop from q75 of ADVERSE travel and the target from q50
of FAVOURABLE travel. Over one horizon those distributions are broadly the same
shape, so q75 > q50 MECHANICALLY - the target came out smaller than the stop no
matter what the market did. SP500 H4 shipped stop 3.07 / target 1.70: a 0.55:1
payoff needing 64.3%. That was never a measurement, it was two mismatched
constants.
The reachability line printed beside it - "target on 50.0% of bars, stop on
25.0%" - is exactly 1-q50 and 1-q75. Tautological. It cannot disconfirm
anything, and it read as validation.
WIDTH AND RATIO ARE INDEPENDENT AND ONLY ONE PAYS. EV = edge x width;
ratio is EV-neutral (a driftless walk reaches +m before -k with probability
k/(k+m), which IS break-even). Width is what buys cost efficiency: the spread
is a fixed 0.047*ATR here, so the shipped 4.77*ATR width paid it 21 times per
unit of travel. So:
RATIO = policy. BARRIER_TARGET_RR = 2.0 (user's 1:2). Break-even 33.3%.
SCALE = measured. The stop quantile is chosen from a ladder, WIDEST FIRST,
taking the first rung whose implied 2x target is still reached often
enough to be a trainable class.
That last clause is the difference from the min-reward:risk raise removed in
2026-08-09, which forced target = 2 x stop with NO reachability test, landed on
6.66*ATR reachable on 3.3% of bars, and trained the model to predict something
that essentially never happened. Same ratio; the scale now retreats until the
data says the target is attainable. Every rung is logged.
LADDER CEILING REMOVED. BARRIER_LADDER stopped at 5.00 and the expectancy scan's
"best resolvable pair on width alone" came back as stop 5.05 / target 4.95 - it
pinned to the top rung. A recommendation landing exactly on the edge of its own
search space is a boundary, not a finding: it cannot tell "5 ATR is optimal"
from "5 ATR is all we allowed". Extended to 20*ATR (8 -> 14 rungs). Nothing else
needs editing - every consumer is parameterised by BARRIER_LADDER_COUNT - and
the horizon constraints (decided >= 60%, reachability floor) now bind instead of
a constant.
THE SCAN COULD NOT SEE THE SHIPPED GEOMETRY. ReportBarrierGeometryScan looked
the configured pair up in its integer grid, and DeriveBarrierGeometry produces
CONTINUOUS multiples (3.07/1.70) that can never equal a grid point - so
cfgExcess stayed at its -1.0 sentinel and the report printed "configured 3:2
scores -1.00000", which reads as a catastrophic score and actually means "never
evaluated". Worse, the grid skipped target<stop entirely because it "inverts the
trade's whole premise" - while the derivation was shipping exactly that. The
incumbent is now always scored as a peer (never crowned; it is already in force
and is not an enum pairing the scan could adopt).
BREAK-EVEN NOW INCLUDES THE SPREAD. Every report quoted the frictionless
SL/(SL+TP). On SP500 H4 that read 64.3% while the MEASURED zero-skill rate was
62.1% - a 2.2pp gap that IS the cost, and that made every model look 2.2pp
better than it was. CostAdjustedBreakEvenPct() prices a win at (TP - spread) and
a loss at (SL + spread), matching the expectancy scan's convention exactly so
the two reports cannot disagree.
It also feeds FitDirConfThreshold, which is the correctness half: the operating
point subtracts break-even from precision, so the frictionless figure made every
candidate threshold look better by the width of the spread - 2.2pp against a
measured edge of 2.3pp, i.e. very nearly all of it.
Era line now carries both: "break-even 64.3% frictionless, 66.6% AFTER SPREAD".
Forces a full relabel and retrain. Requested.
NOT COMPILED - user compiles.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 00:34:32 -04:00
|
|
|
//| Break-even INCLUDING the spread - see the declaration comment. |
|
|
|
|
|
//+------------------------------------------------------------------+
|
feat(gate): grade OOS calls on the exit policy actually in force, and move vote combining out of the members and into the orchestrator
Option (a) from the exit-policy question: the certified number must be the traded
number. Plus the modularity correction the user called for on 778b6c0.
1. VOTE COMBINING BELONGS TO THE ORCHESTRATOR, NOT TO A MEMBER.
778b6c0 fixed the last-writer-wins bug on g_LiveAISignedConfidence by having a
member average its siblings through g_warriorEnsemble. That trades a scheduling bug
for a coupling bug, and it is the wrong shape for this EA: every signal runs in its
own instance, minds its own state, and VOTES to the orchestrator, which is the only
thing allowed to combine opinions.
Replaced with a publish/aggregate pair in Variables\ConfidenceBridge.mqh whose split
is enforced by shape rather than by convention:
- PublishAIVote(slot, conf) - a member writes ONLY its own slot, reads nobody's;
- AggregateAIVotes() - called by CExpertSignalCustom::LiveSignedConfidence.
CExpertSignalAIBase::EnsembleLiveSignedConfidence is gone. The orchestrator also
republishes the aggregate into g_LiveAISignedConfidence, because the intelligent
trailing reads that global directly and must act on the same number the exit route
does rather than on a leftover from whichever member ticked last. A solo AI signal
owns slot 0, so the non-ensemble path is unchanged.
2. THE GATE NOW REPLAYS THE REAL EXIT RULE.
SimulateTradeOutcome() walks the same price series with the same fill/barrier/spread
convention as ComputeLabelForBar - deliberately by copy, so a disagreement between
the two can only be a policy effect and never a discrepancy between two pieces of our
own arithmetic - and terminates at the FIRST of stop / target / vote reversal /
horizon. Barriers are tested before the vote on the same bar: intrabar we cannot know
which came first, and the barrier is what the broker executes automatically, so
checking the vote first would credit the exit policy with escapes a real stop would
have taken out of its hands.
It runs AFTER pass 3, not inside it. A vote-flip exit for a trade entered at bar r is
decided by the model's output at bars r-1, r-2, ... - NEWER bars - and pass 3 walks
oldest-to-newest, so at the moment r is graded its own exit does not exist yet. Only
once m_oosDecisionSeries is complete over the whole OOS window can the replay run.
In ensemble mode that series carries the member's adjusted decision and the live exit
reads the ensemble aggregate, which is the coupling the user identified: an LSTM entry
really can be closed by the ensemble turning against it.
3. THE STATISTIC HAS TO CHANGE WITH THE POLICY, AND THAT IS THE REAL FINDING.
A barrier exit pays a fixed R. A vote exit pays whatever the close happens to be. So
the moment vote exits are enabled the payoff is CONTINUOUS, and "win rate vs
break-even" stops being a meaningful test - there is no fixed break-even for a
variable payoff. SimulateTradeOutcome therefore returns R rather than a bool, and the
replay reports expectancy in R with its SE taken from the R distribution (overlap-
deflated on the same EffectiveSampleSize doctrine as every other SE here), not from a
binomial.
This is the same class of error as win-based scoring in 2026-08-09: measuring a
variable-payoff process with a fixed-payoff statistic. Naming it now, while vote exits
are still off, is much cheaper than discovering it after they go on.
4. WHY THIS IS SAFE TO SHIP TODAY.
Min_Vote_Close ships Disabled (101, unreachable on both scales it drives) and reaches
the AI signal through the new ExitPolicy() setter as 1.01, which the setter turns into
0.0 = no vote exit. Under that policy every replayed trade resolves at a barrier and
the simulation is arithmetically the same trade the deploy gate already certifies -
they cannot drift. The report says so explicitly, and prints ONCE per run in that
state; when vote exits are on it prints every era, because then the divergence is the
thing to watch. Nothing about today's numbers moves.
The gate switchover is wired but dormant by construction: it becomes exit-aware the
moment the input is enabled, which is exactly what "the certified number is the traded
number" has to mean.
KNOWN LIMIT, stated rather than hidden: only the AI early-exit route is replayed. The
rule-based averaged-vote close (m_threshold_close) depends on every other filter's
live vote, which pass 3 does not reproduce, so a position the classic filters would
have closed is held to its barrier here. The replay therefore holds LONGER than live
and overstates barrier-reached outcomes. Faithful only while the AI is the dominant
vote - which is the configuration this is being built for.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 16:25:57 -04:00
|
|
|
//+------------------------------------------------------------------+
|
|
|
|
|
//| See the declaration. The trade the EA would ACTUALLY have taken |
|
|
|
|
|
//| from this bar, under the exit policy actually in force. |
|
|
|
|
|
//+------------------------------------------------------------------+
|
|
|
|
|
bool CExpertSignalAIBase::SimulateTradeOutcome(int entryIdx, bool isLong, double &rMultiple,
|
|
|
|
|
int &lifespanBars, bool &endedOnVote)
|
|
|
|
|
{
|
|
|
|
|
rMultiple = 0.0;
|
|
|
|
|
lifespanBars = 0;
|
|
|
|
|
endedOnVote = false;
|
|
|
|
|
double atr = m_ATR.Main(entryIdx);
|
|
|
|
|
if(!MathIsValidNumber(atr) || atr <= 0.0)
|
|
|
|
|
return false;
|
|
|
|
|
double entry = m_Close.GetData(entryIdx);
|
|
|
|
|
if(!MathIsValidNumber(entry) || entry <= 0.0)
|
|
|
|
|
return false;
|
|
|
|
|
double slMult, tpMult;
|
|
|
|
|
BarrierMultiples(slMult, tpMult);
|
|
|
|
|
double risk = slMult * atr;
|
|
|
|
|
double reward = tpMult * atr;
|
|
|
|
|
if(risk <= 0.0)
|
|
|
|
|
return false;
|
|
|
|
|
double spread = (double)m_symbol.Spread() * m_symbol.Point();
|
|
|
|
|
if(!MathIsValidNumber(spread) || spread < 0.0)
|
|
|
|
|
spread = 0.0;
|
|
|
|
|
//--- IDENTICAL fill/barrier convention to ComputeLabelForBar's walk, deliberately and by copy: if the
|
|
|
|
|
//--- two ever disagree about what a trade costs, the "simulated vs hold-to-barrier" comparison this
|
|
|
|
|
//--- function exists to produce measures the discrepancy between two pieces of our own arithmetic
|
|
|
|
|
//--- rather than the effect of the exit policy.
|
|
|
|
|
double fill = isLong ? (entry + spread) : (entry - spread);
|
|
|
|
|
double tpLevel = isLong ? (entry + spread + reward) : (entry - reward - spread);
|
|
|
|
|
double slLevel = isLong ? (entry + spread - risk) : (entry + risk - spread);
|
|
|
|
|
//--- Vote-reversal threshold, 0 when the policy has no vote-driven exit. Mirrors the LIVE rule in
|
feat(vote): thresholds become confidence percentages, on ONE scale everywhere
User request: "the entry/exit thresholds are manual numbers, I would like
them to be confidence percentages, so the current 20 would be only 20%
confidence in a profitable trade."
WHY 20 WAS EVER SENSIBLE. Under UseDatabaseRanking both factors of a filter's
contribution are win rates: the pattern weight is that pattern's measured win
rate (UpdateSignalsWeights -> ApplyPatternWeight) and m_weight is the filter's
average win rate over its patterns, /100. Dividing the sum by the VOTER COUNT
therefore produced a mean of PRODUCTS of two win rates - a genuinely
60%-accurate filter firing a 60% pattern scored 0.60 x 60 = 36. The number was
never on a probability scale, so its magnitude meant nothing on its own.
Dividing by Sum(m_weight) instead makes it a weighted MEAN of win rates, which
is a win rate: result = Sum(w_i*p_i)/Sum(w_i). Every voter at 60% now reads 60;
MACD's double-divergence pattern (weight 100) voting alone reads 100. m_weight
stops being a discount on the probability and becomes how much a filter's
opinion COUNTS - which is what a module weight should always have been.
Default Min_Vote_Open 20 -> 50: not a tightening, the same bar re-expressed.
ONE SCALE, EVERYWHERE - the part that made this bigger than a rescale. Three
other places compared against a 0..1 softmax confidence and would each have
become a fresh currency mismatch the moment the input changed meaning:
* the AI early-exit route (LiveSignedConfidence vs m_ai_exit_threshold) now
reads m_lastAiVote - the AI filters' own weighted mean, undiluted by the
classic side, which is the only reason that route exists - against the
same m_threshold_close the averaged vote uses. m_ai_exit_threshold is
retired rather than left dangling.
* m_oosDecisionSeries now carries the vote, not the confidence, so the exit
SIMULATION stops modelling a close rule the EA does not run.
* ExitPolicy() clamped anything > 1.0 to zero. Passing the unscaled input
through that would have silently switched vote exits off in the
simulation while live went on running them - found before it shipped;
the bound now tracks the scale.
LiveSignedConfidence() is deliberately untouched and still 0..1: MM sizing,
SL/TP scaling and the intelligent trailing want a model confidence, not a win
rate.
CALIBRATION CAVEAT, stated in the code where the claim is made: this is only a
real probability to the extent the pattern weights are. A pattern with fewer
than MIN_TRADES_FOR_WIN_RATE journaled trades keeps its DEFAULT weight - a
designed prior (25/50/75/100 for the AI tiers), not a measurement. Until the
signal DB fills, "60" means "the designed conviction of the patterns that
fired". Closing that gap is the next commit.
Also corrects VOTE_CLOSE_PRESETS' comment, which documented the two scales
this removes.
NOT COMPILED - user compiles in MetaEditor.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-18 15:52:08 -04:00
|
|
|
//--- CExpertSignalCustom::CheckClosePosition: the route fires when the AI VOTE has reversed
|
|
|
|
|
//--- against the position and its magnitude reaches the threshold. Both sides are on the 0-100
|
|
|
|
|
//--- win-rate scale as of 2026-08-18 - m_oosDecisionSeries carries the vote, not the softmax
|
|
|
|
|
//--- confidence it used to, and m_exitVoteThreshold is Min_Vote_Close unscaled. A simulation
|
|
|
|
|
//--- that models a different exit rule than the one that runs is worse than no simulation.
|
|
|
|
|
bool voteExitsOn = (!m_exitHoldToBarrier && m_exitVoteThreshold > 0.0 && m_exitVoteThreshold <= 100.0
|
feat(gate): grade OOS calls on the exit policy actually in force, and move vote combining out of the members and into the orchestrator
Option (a) from the exit-policy question: the certified number must be the traded
number. Plus the modularity correction the user called for on 778b6c0.
1. VOTE COMBINING BELONGS TO THE ORCHESTRATOR, NOT TO A MEMBER.
778b6c0 fixed the last-writer-wins bug on g_LiveAISignedConfidence by having a
member average its siblings through g_warriorEnsemble. That trades a scheduling bug
for a coupling bug, and it is the wrong shape for this EA: every signal runs in its
own instance, minds its own state, and VOTES to the orchestrator, which is the only
thing allowed to combine opinions.
Replaced with a publish/aggregate pair in Variables\ConfidenceBridge.mqh whose split
is enforced by shape rather than by convention:
- PublishAIVote(slot, conf) - a member writes ONLY its own slot, reads nobody's;
- AggregateAIVotes() - called by CExpertSignalCustom::LiveSignedConfidence.
CExpertSignalAIBase::EnsembleLiveSignedConfidence is gone. The orchestrator also
republishes the aggregate into g_LiveAISignedConfidence, because the intelligent
trailing reads that global directly and must act on the same number the exit route
does rather than on a leftover from whichever member ticked last. A solo AI signal
owns slot 0, so the non-ensemble path is unchanged.
2. THE GATE NOW REPLAYS THE REAL EXIT RULE.
SimulateTradeOutcome() walks the same price series with the same fill/barrier/spread
convention as ComputeLabelForBar - deliberately by copy, so a disagreement between
the two can only be a policy effect and never a discrepancy between two pieces of our
own arithmetic - and terminates at the FIRST of stop / target / vote reversal /
horizon. Barriers are tested before the vote on the same bar: intrabar we cannot know
which came first, and the barrier is what the broker executes automatically, so
checking the vote first would credit the exit policy with escapes a real stop would
have taken out of its hands.
It runs AFTER pass 3, not inside it. A vote-flip exit for a trade entered at bar r is
decided by the model's output at bars r-1, r-2, ... - NEWER bars - and pass 3 walks
oldest-to-newest, so at the moment r is graded its own exit does not exist yet. Only
once m_oosDecisionSeries is complete over the whole OOS window can the replay run.
In ensemble mode that series carries the member's adjusted decision and the live exit
reads the ensemble aggregate, which is the coupling the user identified: an LSTM entry
really can be closed by the ensemble turning against it.
3. THE STATISTIC HAS TO CHANGE WITH THE POLICY, AND THAT IS THE REAL FINDING.
A barrier exit pays a fixed R. A vote exit pays whatever the close happens to be. So
the moment vote exits are enabled the payoff is CONTINUOUS, and "win rate vs
break-even" stops being a meaningful test - there is no fixed break-even for a
variable payoff. SimulateTradeOutcome therefore returns R rather than a bool, and the
replay reports expectancy in R with its SE taken from the R distribution (overlap-
deflated on the same EffectiveSampleSize doctrine as every other SE here), not from a
binomial.
This is the same class of error as win-based scoring in 2026-08-09: measuring a
variable-payoff process with a fixed-payoff statistic. Naming it now, while vote exits
are still off, is much cheaper than discovering it after they go on.
4. WHY THIS IS SAFE TO SHIP TODAY.
Min_Vote_Close ships Disabled (101, unreachable on both scales it drives) and reaches
the AI signal through the new ExitPolicy() setter as 1.01, which the setter turns into
0.0 = no vote exit. Under that policy every replayed trade resolves at a barrier and
the simulation is arithmetically the same trade the deploy gate already certifies -
they cannot drift. The report says so explicitly, and prints ONCE per run in that
state; when vote exits are on it prints every era, because then the divergence is the
thing to watch. Nothing about today's numbers moves.
The gate switchover is wired but dormant by construction: it becomes exit-aware the
moment the input is enabled, which is exactly what "the certified number is the traded
number" has to mean.
KNOWN LIMIT, stated rather than hidden: only the AI early-exit route is replayed. The
rule-based averaged-vote close (m_threshold_close) depends on every other filter's
live vote, which pass 3 does not reproduce, so a position the classic filters would
have closed is held to its barrier here. The replay therefore holds LONGER than live
and overstates barrier-reached outcomes. Faithful only while the AI is the dominant
vote - which is the configuration this is being built for.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 16:25:57 -04:00
|
|
|
&& ArraySize(m_oosDecisionSeries) > 0);
|
|
|
|
|
int last = entryIdx - MathMax(m_barrierHorizonBars, 1);
|
|
|
|
|
if(last < 0)
|
|
|
|
|
last = 0;
|
|
|
|
|
for(int t = entryIdx - 1; t >= last; t--)
|
|
|
|
|
{
|
|
|
|
|
double hi = m_High.GetData(t);
|
|
|
|
|
double lo = m_Low.GetData(t);
|
|
|
|
|
double cl = m_Close.GetData(t);
|
|
|
|
|
if(!MathIsValidNumber(hi) || !MathIsValidNumber(lo) || hi == EMPTY_VALUE || lo == EMPTY_VALUE)
|
|
|
|
|
break;
|
|
|
|
|
lifespanBars = entryIdx - t;
|
|
|
|
|
//--- STOP FIRST on a bar that spans both, same pessimism as the label walk.
|
|
|
|
|
if(isLong ? (lo <= slLevel) : (hi >= slLevel))
|
|
|
|
|
{
|
|
|
|
|
rMultiple = -1.0;
|
|
|
|
|
return true;
|
|
|
|
|
}
|
|
|
|
|
if(isLong ? (hi >= tpLevel) : (lo <= tpLevel))
|
|
|
|
|
{
|
|
|
|
|
rMultiple = reward / risk;
|
|
|
|
|
return true;
|
|
|
|
|
}
|
|
|
|
|
//--- VOTE REVERSAL, checked AFTER the barriers on the same bar. Order matters and this is the
|
|
|
|
|
//--- honest one: intrabar we cannot know whether the barrier or the close came first, and the
|
|
|
|
|
//--- barrier is the outcome the broker would have executed automatically, without waiting for a
|
|
|
|
|
//--- bar close. Checking the vote first would credit the exit policy with escapes that a real
|
|
|
|
|
//--- stop would have taken out of its hands.
|
|
|
|
|
if(voteExitsOn && t < ArraySize(m_oosDecisionSeries))
|
|
|
|
|
{
|
|
|
|
|
double vote = m_oosDecisionSeries[t];
|
|
|
|
|
bool reversed = isLong ? (vote < 0.0) : (vote > 0.0);
|
|
|
|
|
if(reversed && MathAbs(vote) >= m_exitVoteThreshold && MathIsValidNumber(cl) && cl > 0.0)
|
|
|
|
|
{
|
|
|
|
|
//--- Closed at THIS bar's close, at whatever P&L that is - which is the whole point: a
|
|
|
|
|
//--- vote exit produces a CONTINUOUS payoff, not a win or a loss, and that is why an
|
|
|
|
|
//--- exit-aware gate cannot go on scoring win-rate against a fixed break-even.
|
|
|
|
|
rMultiple = isLong ? (cl - fill) / risk : (fill - cl) / risk;
|
|
|
|
|
endedOnVote = true;
|
|
|
|
|
return true;
|
|
|
|
|
}
|
|
|
|
|
}
|
|
|
|
|
}
|
|
|
|
|
//--- Ran out of horizon (or history) with neither barrier touched: the trade is closed at the last
|
|
|
|
|
//--- bar the walk could see, exactly as the live horizon-timeout would.
|
|
|
|
|
int lastSeen = entryIdx - MathMax(lifespanBars, 1);
|
|
|
|
|
if(lastSeen < 0)
|
|
|
|
|
lastSeen = 0;
|
|
|
|
|
double closeOut = m_Close.GetData(lastSeen);
|
|
|
|
|
if(!MathIsValidNumber(closeOut) || closeOut <= 0.0)
|
|
|
|
|
return false;
|
|
|
|
|
rMultiple = isLong ? (closeOut - fill) / risk : (fill - closeOut) / risk;
|
|
|
|
|
return true;
|
|
|
|
|
}
|
|
|
|
|
//+------------------------------------------------------------------+
|
|
|
|
|
//| See the declaration. Replays this era's OOS calls under the exit |
|
|
|
|
|
//| policy actually in force, and says how far that lands from the |
|
|
|
|
|
//| hold-to-barrier outcome the gate certifies. |
|
|
|
|
|
//+------------------------------------------------------------------+
|
|
|
|
|
void CExpertSignalAIBase::SimulateExitPolicyOutcomes(void)
|
|
|
|
|
{
|
|
|
|
|
int n = ArraySize(m_oosDecisionSeries);
|
|
|
|
|
for(int r = 0; r < n; r++)
|
|
|
|
|
{
|
fix(shutdown): make ExitPolicy public, and stop every long loop the moment MT5 asks
Two things, one of which was a compile error.
1. ExitPolicy() was declared in the protected block but is pushed in from
Warrior_EA.mq5:770. Moved to public beside the other EA-facing setters.
2. Chart objects surviving OnDeinit. The 4,500 ms teardown budget is measured
from the STOP REQUEST, not from OnDeinit's first line, and OnDeinit cannot
begin until whatever is in flight returns - so a scan still running after
_StopFlag is raised does not delay the cleanup, it SPENDS it, and the purge
never gets its turn.
New CExpertSignalAIBase::ShutdownRequested() = IsStopped() || m_shutdownInProgress.
Deliberately NOT m_trainingStopRequested: that latches, and a latched flag
would permanently disable scans that must run again on the next Start.
Guarded, longest first:
- TuneIndicatorsByFilter - per candidate, restoring the OPERATOR's settings
on the way out (best[] is mutated in place; the tuner otherwise keeps the
last trial's parameters, which nothing chose).
- ReportBarrierGeometryScan - per pairing, breaking to ONE restore point so
m_barrierScanLiveLabels can never be left true (that makes ComputeLabelForBar
read the last candidate's multiples as the configured geometry).
- ReportFeatureLabelInformation / ReportExcursionInformation / lag profile -
nulls ABANDON rather than truncate: fewer draws is not a smaller null, it
is a wrong one, and p shifts toward significance. m_dirEvidence staying
false is the safe direction.
- SimulateExitPolicyOutcomes - zeroes its accumulators so the divergence line
is dropped instead of latching a partial expectancy as the run's only report.
- ReportGeometryExpectancyScan - per ladder rung.
- HttpGet - one choke point for up to a dozen blocking WebRequests per
first-pass Update(). An in-flight request cannot be cancelled; refusing to
start another is the whole remedy.
- PollTraining, OnChartEventHandler's study event, TuneIndicatorsAndTrain -
entry points, so a queued event cannot open an era during teardown.
TuneIndicatorsAndTrain's guard is the first statement, ahead of the
m_tuneFilterDone / g_ensembleChartTuneDone latches.
- OnTick / OnTimer / OnChartEvent.
Training's own bar loops already honoured this (pass 1 per bar, passes 2/2.5/3
yield on a 120 ms budget); the warm-up scans did not, and they are the longest
uninterruptible stretches the EA has.
StopTraining() is unchanged: the operator's Stop still finalises synchronously.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 17:03:11 -04:00
|
|
|
//--- One full forward price walk per directional call, run in a single unchunked pass at the end of
|
|
|
|
|
//--- an era - the longest thing between the last pass-3 yield and Train() returning. Zeroing the
|
|
|
|
|
//--- accumulators on the way out is deliberate: ReportExitPolicyDivergence skips on m_simTrades<=0,
|
|
|
|
|
//--- so a stop drops the line entirely instead of printing an expectancy over the arbitrary prefix
|
|
|
|
|
//--- of trades that happened to run - and it LATCHES m_exitReplayReported, so a partial number
|
|
|
|
|
//--- would be the only one this run ever prints.
|
|
|
|
|
if(ShutdownRequested())
|
|
|
|
|
{
|
|
|
|
|
m_simRSum = 0.0;
|
|
|
|
|
m_simRSumSq = 0.0;
|
|
|
|
|
m_simTrades = 0;
|
|
|
|
|
m_simVoteExits = 0;
|
|
|
|
|
m_simBarrierWins = 0;
|
|
|
|
|
return;
|
|
|
|
|
}
|
feat(gate): grade OOS calls on the exit policy actually in force, and move vote combining out of the members and into the orchestrator
Option (a) from the exit-policy question: the certified number must be the traded
number. Plus the modularity correction the user called for on 778b6c0.
1. VOTE COMBINING BELONGS TO THE ORCHESTRATOR, NOT TO A MEMBER.
778b6c0 fixed the last-writer-wins bug on g_LiveAISignedConfidence by having a
member average its siblings through g_warriorEnsemble. That trades a scheduling bug
for a coupling bug, and it is the wrong shape for this EA: every signal runs in its
own instance, minds its own state, and VOTES to the orchestrator, which is the only
thing allowed to combine opinions.
Replaced with a publish/aggregate pair in Variables\ConfidenceBridge.mqh whose split
is enforced by shape rather than by convention:
- PublishAIVote(slot, conf) - a member writes ONLY its own slot, reads nobody's;
- AggregateAIVotes() - called by CExpertSignalCustom::LiveSignedConfidence.
CExpertSignalAIBase::EnsembleLiveSignedConfidence is gone. The orchestrator also
republishes the aggregate into g_LiveAISignedConfidence, because the intelligent
trailing reads that global directly and must act on the same number the exit route
does rather than on a leftover from whichever member ticked last. A solo AI signal
owns slot 0, so the non-ensemble path is unchanged.
2. THE GATE NOW REPLAYS THE REAL EXIT RULE.
SimulateTradeOutcome() walks the same price series with the same fill/barrier/spread
convention as ComputeLabelForBar - deliberately by copy, so a disagreement between
the two can only be a policy effect and never a discrepancy between two pieces of our
own arithmetic - and terminates at the FIRST of stop / target / vote reversal /
horizon. Barriers are tested before the vote on the same bar: intrabar we cannot know
which came first, and the barrier is what the broker executes automatically, so
checking the vote first would credit the exit policy with escapes a real stop would
have taken out of its hands.
It runs AFTER pass 3, not inside it. A vote-flip exit for a trade entered at bar r is
decided by the model's output at bars r-1, r-2, ... - NEWER bars - and pass 3 walks
oldest-to-newest, so at the moment r is graded its own exit does not exist yet. Only
once m_oosDecisionSeries is complete over the whole OOS window can the replay run.
In ensemble mode that series carries the member's adjusted decision and the live exit
reads the ensemble aggregate, which is the coupling the user identified: an LSTM entry
really can be closed by the ensemble turning against it.
3. THE STATISTIC HAS TO CHANGE WITH THE POLICY, AND THAT IS THE REAL FINDING.
A barrier exit pays a fixed R. A vote exit pays whatever the close happens to be. So
the moment vote exits are enabled the payoff is CONTINUOUS, and "win rate vs
break-even" stops being a meaningful test - there is no fixed break-even for a
variable payoff. SimulateTradeOutcome therefore returns R rather than a bool, and the
replay reports expectancy in R with its SE taken from the R distribution (overlap-
deflated on the same EffectiveSampleSize doctrine as every other SE here), not from a
binomial.
This is the same class of error as win-based scoring in 2026-08-09: measuring a
variable-payoff process with a fixed-payoff statistic. Naming it now, while vote exits
are still off, is much cheaper than discovering it after they go on.
4. WHY THIS IS SAFE TO SHIP TODAY.
Min_Vote_Close ships Disabled (101, unreachable on both scales it drives) and reaches
the AI signal through the new ExitPolicy() setter as 1.01, which the setter turns into
0.0 = no vote exit. Under that policy every replayed trade resolves at a barrier and
the simulation is arithmetically the same trade the deploy gate already certifies -
they cannot drift. The report says so explicitly, and prints ONCE per run in that
state; when vote exits are on it prints every era, because then the divergence is the
thing to watch. Nothing about today's numbers moves.
The gate switchover is wired but dormant by construction: it becomes exit-aware the
moment the input is enabled, which is exactly what "the certified number is the traded
number" has to mean.
KNOWN LIMIT, stated rather than hidden: only the AI early-exit route is replayed. The
rule-based averaged-vote close (m_threshold_close) depends on every other filter's
live vote, which pass 3 does not reproduce, so a position the classic filters would
have closed is held to its barrier here. The replay therefore holds LONGER than live
and overstates barrier-reached outcomes. Faithful only while the AI is the dominant
vote - which is the configuration this is being built for.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 16:25:57 -04:00
|
|
|
double vote = m_oosDecisionSeries[r];
|
|
|
|
|
if(vote == 0.0)
|
|
|
|
|
continue; // abstained - no trade to replay
|
|
|
|
|
bool isLong = (vote > 0.0);
|
|
|
|
|
double rMult = 0.0;
|
|
|
|
|
int life = 0;
|
|
|
|
|
bool onVote = false;
|
|
|
|
|
if(!SimulateTradeOutcome(r, isLong, rMult, life, onVote))
|
|
|
|
|
continue;
|
|
|
|
|
m_simRSum += rMult;
|
|
|
|
|
m_simRSumSq += rMult * rMult;
|
|
|
|
|
m_simTrades++;
|
|
|
|
|
if(onVote)
|
|
|
|
|
m_simVoteExits++;
|
|
|
|
|
//--- What the CERTIFICATE counts on this same call, so the two are compared on identical trades
|
|
|
|
|
//--- rather than on two different populations.
|
|
|
|
|
bool barrierWin = (isLong
|
|
|
|
|
? (r < ArraySize(m_winLongCache) && m_winLongCache[r])
|
|
|
|
|
: (r < ArraySize(m_winShortCache) && m_winShortCache[r]));
|
|
|
|
|
if(barrierWin)
|
|
|
|
|
m_simBarrierWins++;
|
|
|
|
|
}
|
|
|
|
|
}
|
|
|
|
|
//+------------------------------------------------------------------+
|
|
|
|
|
//| See the declaration. The one line that says whether the number |
|
|
|
|
|
//| being certified is still the number that would be traded. |
|
|
|
|
|
//+------------------------------------------------------------------+
|
|
|
|
|
void CExpertSignalAIBase::ReportExitPolicyDivergence(void)
|
|
|
|
|
{
|
|
|
|
|
if(m_simTrades <= 0)
|
|
|
|
|
return;
|
|
|
|
|
bool policyIsBarrier0 = (m_exitHoldToBarrier || m_exitVoteThreshold <= 0.0);
|
|
|
|
|
//--- Every era while vote exits are ON, because then this is load-bearing and its drift is the thing
|
|
|
|
|
//--- to watch. ONCE per run while they are off, because then it is arithmetically guaranteed to agree
|
|
|
|
|
//--- with the certificate and a line per era would be pure noise in an already large journal.
|
|
|
|
|
if(policyIsBarrier0 && m_exitReplayReported)
|
|
|
|
|
return;
|
|
|
|
|
m_exitReplayReported = true;
|
|
|
|
|
double meanR = m_simRSum / m_simTrades;
|
|
|
|
|
//--- SE on the R DISTRIBUTION, not a binomial: once a vote exit can end a trade anywhere between the
|
|
|
|
|
//--- two barriers the payoff is continuous, so "win rate vs break-even" stops being the right
|
|
|
|
|
//--- statistic and expectancy-in-R vs 0 replaces it. Deflated by the label overlap on the same
|
|
|
|
|
//--- doctrine as every other SE here (see EffectiveSampleSize).
|
|
|
|
|
double varR = (m_simRSumSq / m_simTrades) - (meanR * meanR);
|
|
|
|
|
if(varR < 0.0)
|
|
|
|
|
varR = 0.0;
|
|
|
|
|
double effN = EffectiveSampleSize((double)m_simTrades);
|
|
|
|
|
double seR = (effN > 0.0) ? MathSqrt(varR / effN) : 0.0;
|
|
|
|
|
double barrierWinPct = 100.0 * (double)m_simBarrierWins / m_simTrades;
|
|
|
|
|
double votePct = 100.0 * (double)m_simVoteExits / m_simTrades;
|
|
|
|
|
bool policyIsBarrier = (m_exitHoldToBarrier || m_exitVoteThreshold <= 0.0);
|
|
|
|
|
PrintFormat("%s: EXIT-POLICY REPLAY of this era's %d OOS calls - policy in force: %s | simulated"
|
|
|
|
|
" expectancy %+.3f R (2 SE %.3f on %.0f independent trades) | %.0f%% closed by a VOTE"
|
|
|
|
|
" REVERSAL before either barrier | hold-to-barrier win rate on the SAME calls %.1f%%"
|
|
|
|
|
" (break-even %.1f%%). %s",
|
|
|
|
|
ID, m_simTrades,
|
|
|
|
|
policyIsBarrier ? "SL/TP only (no vote exit)" : "vote exit ENABLED",
|
|
|
|
|
meanR, 2.0 * seR, effN, votePct, barrierWinPct, CostAdjustedBreakEvenPct(),
|
|
|
|
|
policyIsBarrier
|
|
|
|
|
? "Vote exits are off, so every trade here resolved at a barrier and this replay is"
|
|
|
|
|
" arithmetically the same trade the deploy gate certifies - the two cannot drift."
|
|
|
|
|
: "VOTE EXITS ARE ON, so these are NOT the trades the deploy gate's win rate describes:"
|
|
|
|
|
" that number grades target-before-stop, and a vote flip inside the horizon is neither."
|
|
|
|
|
" Read the expectancy, not the win rate - a win rate over trades with continuous"
|
|
|
|
|
" payoffs has no fixed break-even to be measured against.");
|
|
|
|
|
}
|
fix(geometry): the target was small BY CONSTRUCTION - ratio is now policy, scale is measured, ladder ceiling removed
The derivation read the stop from q75 of ADVERSE travel and the target from q50
of FAVOURABLE travel. Over one horizon those distributions are broadly the same
shape, so q75 > q50 MECHANICALLY - the target came out smaller than the stop no
matter what the market did. SP500 H4 shipped stop 3.07 / target 1.70: a 0.55:1
payoff needing 64.3%. That was never a measurement, it was two mismatched
constants.
The reachability line printed beside it - "target on 50.0% of bars, stop on
25.0%" - is exactly 1-q50 and 1-q75. Tautological. It cannot disconfirm
anything, and it read as validation.
WIDTH AND RATIO ARE INDEPENDENT AND ONLY ONE PAYS. EV = edge x width;
ratio is EV-neutral (a driftless walk reaches +m before -k with probability
k/(k+m), which IS break-even). Width is what buys cost efficiency: the spread
is a fixed 0.047*ATR here, so the shipped 4.77*ATR width paid it 21 times per
unit of travel. So:
RATIO = policy. BARRIER_TARGET_RR = 2.0 (user's 1:2). Break-even 33.3%.
SCALE = measured. The stop quantile is chosen from a ladder, WIDEST FIRST,
taking the first rung whose implied 2x target is still reached often
enough to be a trainable class.
That last clause is the difference from the min-reward:risk raise removed in
2026-08-09, which forced target = 2 x stop with NO reachability test, landed on
6.66*ATR reachable on 3.3% of bars, and trained the model to predict something
that essentially never happened. Same ratio; the scale now retreats until the
data says the target is attainable. Every rung is logged.
LADDER CEILING REMOVED. BARRIER_LADDER stopped at 5.00 and the expectancy scan's
"best resolvable pair on width alone" came back as stop 5.05 / target 4.95 - it
pinned to the top rung. A recommendation landing exactly on the edge of its own
search space is a boundary, not a finding: it cannot tell "5 ATR is optimal"
from "5 ATR is all we allowed". Extended to 20*ATR (8 -> 14 rungs). Nothing else
needs editing - every consumer is parameterised by BARRIER_LADDER_COUNT - and
the horizon constraints (decided >= 60%, reachability floor) now bind instead of
a constant.
THE SCAN COULD NOT SEE THE SHIPPED GEOMETRY. ReportBarrierGeometryScan looked
the configured pair up in its integer grid, and DeriveBarrierGeometry produces
CONTINUOUS multiples (3.07/1.70) that can never equal a grid point - so
cfgExcess stayed at its -1.0 sentinel and the report printed "configured 3:2
scores -1.00000", which reads as a catastrophic score and actually means "never
evaluated". Worse, the grid skipped target<stop entirely because it "inverts the
trade's whole premise" - while the derivation was shipping exactly that. The
incumbent is now always scored as a peer (never crowned; it is already in force
and is not an enum pairing the scan could adopt).
BREAK-EVEN NOW INCLUDES THE SPREAD. Every report quoted the frictionless
SL/(SL+TP). On SP500 H4 that read 64.3% while the MEASURED zero-skill rate was
62.1% - a 2.2pp gap that IS the cost, and that made every model look 2.2pp
better than it was. CostAdjustedBreakEvenPct() prices a win at (TP - spread) and
a loss at (SL + spread), matching the expectancy scan's convention exactly so
the two reports cannot disagree.
It also feeds FitDirConfThreshold, which is the correctness half: the operating
point subtracts break-even from precision, so the frictionless figure made every
candidate threshold look better by the width of the spread - 2.2pp against a
measured edge of 2.3pp, i.e. very nearly all of it.
Era line now carries both: "break-even 64.3% frictionless, 66.6% AFTER SPREAD".
Forces a full relabel and retrain. Requested.
NOT COMPILED - user compiles.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 00:34:32 -04:00
|
|
|
double CExpertSignalAIBase::CostAdjustedBreakEvenPct(void)
|
|
|
|
|
{
|
|
|
|
|
double slMult = 0.0, tpMult = 0.0;
|
|
|
|
|
BarrierMultiples(slMult, tpMult);
|
|
|
|
|
double frictionless = (slMult + tpMult > 0.0) ? 100.0 * slMult / (slMult + tpMult) : 50.0;
|
|
|
|
|
if(!MathIsValidNumber(m_spreadAtr) || m_spreadAtr <= 0.0)
|
|
|
|
|
return frictionless;
|
|
|
|
|
//--- A long fills at close+spread, so its target needs (TP - spread) of net travel to pay and its stop
|
|
|
|
|
//--- costs (SL + spread) when it trips. Same convention ReportGeometryExpectancyScan prices its ladder
|
|
|
|
|
//--- with, so the two reports cannot disagree about what a trade costs.
|
|
|
|
|
double reward = tpMult - m_spreadAtr;
|
|
|
|
|
double risk = slMult + m_spreadAtr;
|
|
|
|
|
if(reward <= 0.0 || risk <= 0.0)
|
|
|
|
|
return frictionless; // target inside the spread - not tradeable at any hit rate
|
|
|
|
|
return 100.0 * risk / (risk + reward);
|
|
|
|
|
}
|
|
|
|
|
//+------------------------------------------------------------------+
|
feat(ai): triple-barrier labels replace exact-pivot ZigZag targets
The 31:1 class imbalance was self-inflicted by the TARGET, not a property
of the market. Labelling only the exact bar where a ZigZag pivot confirms
gave Buy 1164 / Sell 1164 / Neutral 35841, and every correction mechanism
this codebase accumulated sits downstream of that one choice: the
logit-adjusted loss and its range cap, the prior EMA, the +-3.0 output-bias
seed, balanced-accuracy-then-precision selection with its coverage floor,
the recall floor and its catch-22, the alternation gate, NMS, and the four
oversampling designs that collapsed before them.
The reference this engine is built on (references/neuronetworksbook.pdf
ch. 3.1/3.3) also uses ZigZag, but targets the DIRECTION TO THE NEXT
EXTREMUM on every bar - ~50/50 by construction, with no imbalance to
correct at all. It never had this problem because it never asked "is this
the pivot bar".
Labels are now the triple barrier (Lopez de Prado ch. 3), using the EA's
OWN SL_Mode/TP_Mode: does a trade opened at this bar's close reach its
target before its stop, within a horizon. Buy = long resolves, Sell =
short resolves, Neutral = neither. Consequences:
- dir-precision in the era line stops being a proxy and becomes the win
rate of the strategy under its own exit rules.
- Expected balance ~25/25/50 at the shipped 1:3 (gambler's ruin), i.e.
~2:1 instead of 31:1. Measured and logged at the end of the prebuild.
- Spread is charged on both legs, so it is a NET win rate.
- Intrabar ambiguity resolves to the STOP. OHLC cannot order two touches
inside one bar and the optimistic reading is how a backtested edge
becomes a live loss.
ZigZag stays as input features (EnableSwingContext) and now also supplies
the vertical barrier: the horizon is the median confirmed leg length,
snapped to a coarse ladder. Derived, not configured, and deliberately kept
out of the filename fingerprint - a filename keyed on a measured quantity
orphans a trained model the moment the measurement moves.
Removed, because the premise died with the old target:
- the alternation gate. Correct for pivot labels (a ZigZag cannot emit two
same-type pivots in a row, so a repeat was provably a false fire), and
wrong for barrier labels, which answer each bar independently. It also
took its worst consequence with it: a one-sided model previously got ONE
trade per backtest, a hard blocker on marketplace validation.
- SignalClusterWindow now defaults off - it de-duplicated repeats that are
now real trades. Kept as an opt-in display control.
- LABEL_WINDOW_BARS, the pivot-widening pass, ConfirmedZigZagLabel.
- the era-0 output-bias seed now needs a genuinely dominant class (0.70)
rather than 0.40; at ~50% Neutral a +-3.0 seed is a distortion, not a
correction.
Also fixed, both found while wiring the above:
1. RefreshConvergedSignal sized its buffers from a date delta
(Bars(sym, period, dtStudied, TimeCurrent())). dtStudied is a training
watermark; in the tester it is loaded from a live-chart save AHEAD of
the simulated date, so the interval inverted, Bars() returned ~0, and
the buffer came out at exactly m_historyBars - deep enough for the OHLC
window and far too shallow for the Donchian-50 / 20-bar-return / SMA
extension behind it. Inference silently computed DIFFERENT features
from the ones training learned on, live as well as in the tester. Now
sized from what the feature builder actually needs.
2. The barrier horizon is resolved on the deployed path too. A deployed
model never enters Train(), so it never reached the prebuild, and
OnlineLearnStep reads the horizon as its confirmation delay - left at
the fallback it would have backpropped bars whose barriers had not
resolved. Silent lookahead in the one place that writes to a live model.
SL_Mode/TP_Mode join the weights fingerprint: they define the labels now,
so a model trained at 1:3 must never be silently reused at 1:1. This
re-keys every pre-existing model by design - none were trained on this task.
Inference census extended with the vote gate. LongCondition/ShortCondition
open with a readiness check the refresh counters never see; in the tester it
reduces to "the seeded _optcache.nnw must have LOADED", and if it did not,
every vote is hard-zeroed while the model still answers Buy. The old three
counters would have read that as "the model says Neutral" - false, and a
completely different fix. This is the leading candidate for the
zero-direction backtest and the census can now name it in one run.
Both builds compile 0 errors / 0 warnings. Forces a full retrain.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 20:39:49 -04:00
|
|
|
void CExpertSignalAIBase::BarrierMultiples(double &slMult, double &tpMult)
|
refactor: split CExpertSignalAIBase implementation by responsibility
ExpertSignalAIBase.mqh was 8216 lines: the class declaration followed by 87
method bodies covering training, labelling, feature extraction, persistence,
chart drawing, online learning, the GA auto-tuner and inference, all in one
file. Train() alone is 1492 lines; a change to arrow drawing meant scrolling
past the era loop.
Moved the bodies into Expert\AIBase\, included at the bottom of the original
after the class declaration:
Training.mqh 1607 era loop, plateau ladder, checkpoint select, deploy
Features.mqh 1093 indicator creation + per-bar input feature vector
ChartUI.mqh 634 arrows, arrow persistence, status panel, cleanup
Persistence.mqh 492 .stats/.cfg sidecars, CPU-inference validation, copy
OnlineLearning.mqh 461 live continual learning, EMA shadow, OOS simulator
Labels.mqh 309 ZigZag pivot labels, async label-cache prebuild
AutoTune.mqh 275 genetic tuner (population, crossover, halving)
Inference.mqh 235 softmax, prior calibration, class priors
ExpertSignalAIBase.mqh 8216 -> 3131 (declaration + topology build only)
This is a pure relocation - verified mechanically, not by eye: HEAD's file
reconstructed from the eight partials plus the surviving remainder is
byte-identical to HEAD, span for span (scratchpad verify_split.py). No
declaration moved, no signature changed, no code rewritten, so behaviour is
unchanged by construction.
Compiles 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 00:42:45 -04:00
|
|
|
{
|
feat(labels): measure which barrier is predictable at entry, don't guess
The alignment scan settled the shape of the problem: 4.7x more is
knowable 5 bars into a 128-bar window than at the entry the model
actually trades. A 6xATR target reached over 128 bars is decided
overwhelmingly by what happens DURING the window, so whatever the entry
state knows is buried under 128 bars of later noise. That is a property
of the TARGET, and it is why four different architectures all landed on
precision exactly equal to the base rate - no topology can undo it.
So measure the target. For each SL/TP pairing a user can actually select,
relabel the same sampled bars and score how much the SAME features say
about THAT outcome at entry. Seconds, no training, no topology, and it
runs on the diagnostic path that already exists.
Ranked on excess over its OWN null as a share of its OWN H(Y), never on
raw nats: each geometry has a different class balance, hence a different
finite-sample bias and a different amount of information there to find,
so raw MI would rank the most BALANCED label rather than the most
PREDICTABLE one. The break-even win rate m/(m+k) is printed beside each
so the ranking is read next to the bar the model must clear.
Stated in the output because it is the easy thing to get wrong: chance
precision EQUALS break-even at every geometry, so a tighter target does
not hand you expectancy. It buys predictability - less noise piled on top
of what the entry state knows - which is the one thing changing topology
cannot do.
Read-only by construction: it relabels a sampled copy via
TripleBarrierLabel(), never writes the label cache (which belongs to the
configured geometry), and restores the horizon and overrides it borrowed.
The overrides apply only when BOTH are positive, so a half-set pair can
never silently relabel a live run.
Compiles 0 errors / 0 warnings, standard and Market. Build tag
geometry-scan-v1. Redeploy only - no retrain to READ the ranking.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 14:32:04 -04:00
|
|
|
//--- Scan override (ReportBarrierGeometryScan). Both must be positive or neither applies, so a half-set
|
|
|
|
|
//--- pair can never silently relabel a live run. Restored to 0 by the scan before it returns; nothing
|
|
|
|
|
//--- else writes these, and no persisted state is keyed on them.
|
|
|
|
|
if(m_barrierScanSlMult > 0.0 && m_barrierScanTpMult > 0.0)
|
|
|
|
|
{
|
|
|
|
|
slMult = m_barrierScanSlMult;
|
|
|
|
|
tpMult = m_barrierScanTpMult;
|
|
|
|
|
return;
|
|
|
|
|
}
|
feat: derive the ATR multiples from measured excursions - no hardcoded geometry
The barrier was still two constants. SL_Mode/TP_Mode left the Inputs tab in
3482b6c, but the fallback was a hardcoded 2:6 and the geometry scan only ever
chose from a hardcoded grid {2,3} x {2,3,4,6,8,10}. Picking the least-bad of
eleven guesses is not deriving anything.
WHY THE SCAN WAS THE WRONG INSTRUMENT, now measurable rather than argued. It
ranks pairings by how predictable their OUTCOME is - a question about direction.
The excursion test (2c78f3b) ran on SP500 H1 and direction is the one thing
absent: ASYMMETRY p=0.0846, against RANGE/UP/DOWN all at p=0.0050, with RANGE
scoring 0.01345 vs a 0.00343 null - 4x, where the barrier label sits at 1.01x.
Hence the scan failing its own gate on every run, and its "winner" wandering
2:8 -> 3:8 -> 2:8 -> 2:4 across four runs of the same data. Excursion SIZE is
strongly measurable, so derive the geometry from that instead.
stop = q25 of measured ADVERSE travel (ordinary noise does not reach it)
target = q50 of measured FAVOURABLE travel (reached ~half the time, by
construction, inside the horizon)
Continuous, in ATR units, superseding the enum multiples. Reachability ("target
on X% of bars, stop on Y%") and the implied break-even are printed so the choice
is auditable rather than trusted.
FIXED-POINT ITERATION, not one-shot. ComputeBarrierHorizonBars scales the
horizon with the target (first-passage time grows with the band) and the
excursions are measured OVER the horizon, so target -> horizon -> excursions ->
target is a real loop - deriving once sizes the target from travel measured
under the PREVIOUS horizon. Re-measures until the multiples move <5%, capped at
3 passes, and says so if it does not settle.
Does NOT create expectancy, and the log says as much: chance precision equals
break-even at every geometry (m/(m+k) on both sides). It buys a target the
market reaches and a stop that survives noise. Where Min_Risk_Reward_Ratio
forces a target the market rarely reaches, it WARNS rather than overriding -
the ratio is the user's risk policy, so the honest move is to state its cost.
That is the collision that once rejected 100% of setups.
Pinned in the .cfg as doubles appended AFTER this morning's two ints, so .cfg
files written earlier today still load (their length guard finds no doubles) and
a model that carries them was trained on them and never re-derives.
Also fixes a message from e5ceed6 that claimed "this model resumed from disk"
unconditionally - it printed above a "seeding era 0" line on a brand-new model,
because the branch fires whenever the cache is not built, which is equally true
before a fresh model's first prebuild. A diagnostic that misreports its own
trigger is worse than one that says nothing: it gets quoted back as evidence.
FORCES A FULL RETRAIN (labels change).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-07 12:06:25 -04:00
|
|
|
//--- DERIVED geometry wins over the mode constants. Set once from the measured excursion distribution
|
|
|
|
|
//--- (DeriveBarrierGeometry) and then pinned in the .cfg, so a trained model keeps the barriers it
|
|
|
|
|
//--- learned. Below the scan override deliberately: the scan is exploring hypothetical geometries and
|
|
|
|
|
//--- must still be able to impose one.
|
|
|
|
|
if(m_geometryDerived && m_derivedSlMult > 0.0 && m_derivedTpMult > 0.0)
|
|
|
|
|
{
|
|
|
|
|
slMult = m_derivedSlMult;
|
|
|
|
|
tpMult = m_derivedTpMult;
|
|
|
|
|
return;
|
|
|
|
|
}
|
feat(ai): triple-barrier labels replace exact-pivot ZigZag targets
The 31:1 class imbalance was self-inflicted by the TARGET, not a property
of the market. Labelling only the exact bar where a ZigZag pivot confirms
gave Buy 1164 / Sell 1164 / Neutral 35841, and every correction mechanism
this codebase accumulated sits downstream of that one choice: the
logit-adjusted loss and its range cap, the prior EMA, the +-3.0 output-bias
seed, balanced-accuracy-then-precision selection with its coverage floor,
the recall floor and its catch-22, the alternation gate, NMS, and the four
oversampling designs that collapsed before them.
The reference this engine is built on (references/neuronetworksbook.pdf
ch. 3.1/3.3) also uses ZigZag, but targets the DIRECTION TO THE NEXT
EXTREMUM on every bar - ~50/50 by construction, with no imbalance to
correct at all. It never had this problem because it never asked "is this
the pivot bar".
Labels are now the triple barrier (Lopez de Prado ch. 3), using the EA's
OWN SL_Mode/TP_Mode: does a trade opened at this bar's close reach its
target before its stop, within a horizon. Buy = long resolves, Sell =
short resolves, Neutral = neither. Consequences:
- dir-precision in the era line stops being a proxy and becomes the win
rate of the strategy under its own exit rules.
- Expected balance ~25/25/50 at the shipped 1:3 (gambler's ruin), i.e.
~2:1 instead of 31:1. Measured and logged at the end of the prebuild.
- Spread is charged on both legs, so it is a NET win rate.
- Intrabar ambiguity resolves to the STOP. OHLC cannot order two touches
inside one bar and the optimistic reading is how a backtested edge
becomes a live loss.
ZigZag stays as input features (EnableSwingContext) and now also supplies
the vertical barrier: the horizon is the median confirmed leg length,
snapped to a coarse ladder. Derived, not configured, and deliberately kept
out of the filename fingerprint - a filename keyed on a measured quantity
orphans a trained model the moment the measurement moves.
Removed, because the premise died with the old target:
- the alternation gate. Correct for pivot labels (a ZigZag cannot emit two
same-type pivots in a row, so a repeat was provably a false fire), and
wrong for barrier labels, which answer each bar independently. It also
took its worst consequence with it: a one-sided model previously got ONE
trade per backtest, a hard blocker on marketplace validation.
- SignalClusterWindow now defaults off - it de-duplicated repeats that are
now real trades. Kept as an opt-in display control.
- LABEL_WINDOW_BARS, the pivot-widening pass, ConfirmedZigZagLabel.
- the era-0 output-bias seed now needs a genuinely dominant class (0.70)
rather than 0.40; at ~50% Neutral a +-3.0 seed is a distortion, not a
correction.
Also fixed, both found while wiring the above:
1. RefreshConvergedSignal sized its buffers from a date delta
(Bars(sym, period, dtStudied, TimeCurrent())). dtStudied is a training
watermark; in the tester it is loaded from a live-chart save AHEAD of
the simulated date, so the interval inverted, Bars() returned ~0, and
the buffer came out at exactly m_historyBars - deep enough for the OHLC
window and far too shallow for the Donchian-50 / 20-bar-return / SMA
extension behind it. Inference silently computed DIFFERENT features
from the ones training learned on, live as well as in the tester. Now
sized from what the feature builder actually needs.
2. The barrier horizon is resolved on the deployed path too. A deployed
model never enters Train(), so it never reached the prebuild, and
OnlineLearnStep reads the horizon as its confirmation delay - left at
the fallback it would have backpropped bars whose barriers had not
resolved. Silent lookahead in the one place that writes to a live model.
SL_Mode/TP_Mode join the weights fingerprint: they define the labels now,
so a model trained at 1:3 must never be silently reused at 1:1. This
re-keys every pre-existing model by design - none were trained on this task.
Inference census extended with the vote gate. LongCondition/ShortCondition
open with a readiness check the refresh counters never see; in the tester it
reduces to "the seeded _optcache.nnw must have LOADED", and if it did not,
every vote is hard-zeroed while the model still answers Buy. The old three
counters would have read that as "the model says Neutral" - false, and a
completely different fix. This is the leading candidate for the
zero-direction backtest and the census can now name it in one run.
Both builds compile 0 errors / 0 warnings. Forces a full retrain.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 20:39:49 -04:00
|
|
|
slMult = (m_sl_mode == SL_INTELLIGENT_MODE) ? SL_INTELLIGENT_BASE_MULT : (double)m_sl_mode;
|
|
|
|
|
//--- Same floor OpenLongParams/OpenShortParams apply before sizing anything off the stop, reproduced
|
|
|
|
|
//--- here so the label's risk leg cannot be tighter than the one a real order would receive.
|
|
|
|
|
if(slMult < MIN_SL_ATR_MULTIPLIER)
|
|
|
|
|
slMult = MIN_SL_ATR_MULTIPLIER;
|
|
|
|
|
tpMult = (m_tp_mode == TP_INTELLIGENT_MODE) ? (TP_INTELLIGENT_BASE_RR * slMult) : (double)m_tp_mode;
|
|
|
|
|
if(tpMult <= 0.0)
|
fix: refuse invalid SL/TP, fix the unreachable deploy floor, scale the horizon
Three defects found by reading the 2026-08-01 training logs, all of which
only became visible because the relabel made the numbers mean something.
1. A STALE ENUM TRAINED FOUR MODELS ON THE WRONG TARGET.
`OnInit: trade settings snapshot - SL_Mode=1 TP_Mode=-101`
-101 was TP_PREV_SWING, deleted from TAKE_PROFIT_MODE on 2026-07-31 in
7eb48f5. MetaTrader does not validate a saved enum input against the
enum's current members, so charts saved before that kept the old
integer. BarrierMultiples()'s `if(tpMult <= 0.0) tpMult = slMult;`
then quietly turned it into a 1:1 barrier, and all four topologies
trained ~250 eras against a strategy nobody selected - while the log
reported "target 1.00*ATR" as though it were configured.
Since the relabel these two inputs ARE the label definition, so this
is not a bad trade setting, it is a wrong dataset. ValidateBarrier-
Inputs() now refuses to start (INIT_FAILED + Alert + an explicit fix)
on any value that is not an enum member. Members are enumerated rather
than range-checked because both enums are sparse and carry negative
sentinels, so no min/max test can tell a legal value from a deleted
one - which is the entire failure mode. The fallback survives as
belt-and-braces but now announces itself: a fallback that cannot say
it fired is indistinguishable from correct behaviour.
2. THE DEPLOYABILITY FLOOR BECAME MATHEMATICALLY UNREACHABLE.
`tradeableOK` required `dirPrecPct >= baseRatePct`, where baseRatePct
is Buy+Sell as a share of all bars. At the old exact-pivot target that
was ~6%, so "beat the base rate" read as "beat chance" and the test
looked sound. Triple-barrier labels put it at ~83%, so the gate now
demanded 83% directional precision - impossible by construction.
Observed live: all four topologies cycling "PLATEAU stage 3 ... nothing
safe to deploy" at a perfectly healthy 43-45% precision, with no
checkpoint able to ship however good it got.
Replaced with ZERO-SKILL precision, max(Buy,Sell)/allBars: exactly the
score of the degenerate always-call-one-direction model this floor
exists to reject. Correct at any base rate - ~43% on the current
labels, ~3% on the old rare-pivot ones. The era line now prints
"(chance N%, edge +Mpp)" beside the selection score, because 44%
precision is excellent against a 3% chance level and worthless against
a 43% one, and reading the first as the second is what made tonight's
run look better than it was.
3. THE HORIZON IGNORED THE BARRIER GEOMETRY.
ComputeBarrierHorizonBars() returned the median ZigZag leg, which
measures how long a ~1 ATR move takes and says nothing about how long
the CONFIGURED barrier needs. First-passage time out of [-m,+k] scales
with m*k, so a 1:3 barrier takes ~3x as long as 1:1; the unscaled
horizon would have timed out most 1:3 trades and pushed Neutral
straight back up, re-creating the imbalance the relabel removes.
Now multiplied by slMult*tpMult, calibrated against a real measurement
rather than assumed: the accidental 1:1 run resolved at horizon 12 with
only 16.7% timeouts, so the swing median is the right scale at m*k=1.
Verifiable, not just asserted: the prebuild now counts barriers that
ended on the VERTICAL barrier and reports them as a share of Neutral.
Neutral conflates "timed out" with "stopped out" and only the first
indicts the horizon.
Both builds compile 0 errors / 0 warnings. Forces a retrain - correcting
TP_Mode re-keys the fingerprint (|TB:1:-101 -> |TB:1:3), which is right:
no existing model was trained on the intended target.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 00:30:49 -04:00
|
|
|
{
|
|
|
|
|
//--- UNREACHABLE via the Inputs tab: ValidateBarrierInputs() (Warrior_EA.mq5) refuses to start on
|
|
|
|
|
//--- any value that is not an enum member. It is kept, and made LOUD, because the silent version of
|
|
|
|
|
//--- this line is what let a stale TP_PREV_SWING (-101) train four topologies for ~250 eras on a
|
|
|
|
|
//--- 1:1 barrier while the log cheerfully reported "target 1.00*ATR" as if that were configured.
|
|
|
|
|
//--- A fallback that cannot announce itself is indistinguishable from correct behaviour.
|
|
|
|
|
if(!m_barrierFallbackWarned)
|
|
|
|
|
{
|
|
|
|
|
m_barrierFallbackWarned = true;
|
|
|
|
|
Print(ID + ": ERROR - take-profit mode " + IntegerToString(m_tp_mode) + " is not a valid ATR "
|
|
|
|
|
"multiple; the barrier label is falling back to " + DoubleToString(slMult, 2) + "*ATR (1:1). "
|
|
|
|
|
"This should have been caught at init - the model being trained does NOT match the "
|
|
|
|
|
"configured strategy.");
|
|
|
|
|
}
|
|
|
|
|
tpMult = slMult;
|
|
|
|
|
}
|
feat(ai): triple-barrier labels replace exact-pivot ZigZag targets
The 31:1 class imbalance was self-inflicted by the TARGET, not a property
of the market. Labelling only the exact bar where a ZigZag pivot confirms
gave Buy 1164 / Sell 1164 / Neutral 35841, and every correction mechanism
this codebase accumulated sits downstream of that one choice: the
logit-adjusted loss and its range cap, the prior EMA, the +-3.0 output-bias
seed, balanced-accuracy-then-precision selection with its coverage floor,
the recall floor and its catch-22, the alternation gate, NMS, and the four
oversampling designs that collapsed before them.
The reference this engine is built on (references/neuronetworksbook.pdf
ch. 3.1/3.3) also uses ZigZag, but targets the DIRECTION TO THE NEXT
EXTREMUM on every bar - ~50/50 by construction, with no imbalance to
correct at all. It never had this problem because it never asked "is this
the pivot bar".
Labels are now the triple barrier (Lopez de Prado ch. 3), using the EA's
OWN SL_Mode/TP_Mode: does a trade opened at this bar's close reach its
target before its stop, within a horizon. Buy = long resolves, Sell =
short resolves, Neutral = neither. Consequences:
- dir-precision in the era line stops being a proxy and becomes the win
rate of the strategy under its own exit rules.
- Expected balance ~25/25/50 at the shipped 1:3 (gambler's ruin), i.e.
~2:1 instead of 31:1. Measured and logged at the end of the prebuild.
- Spread is charged on both legs, so it is a NET win rate.
- Intrabar ambiguity resolves to the STOP. OHLC cannot order two touches
inside one bar and the optimistic reading is how a backtested edge
becomes a live loss.
ZigZag stays as input features (EnableSwingContext) and now also supplies
the vertical barrier: the horizon is the median confirmed leg length,
snapped to a coarse ladder. Derived, not configured, and deliberately kept
out of the filename fingerprint - a filename keyed on a measured quantity
orphans a trained model the moment the measurement moves.
Removed, because the premise died with the old target:
- the alternation gate. Correct for pivot labels (a ZigZag cannot emit two
same-type pivots in a row, so a repeat was provably a false fire), and
wrong for barrier labels, which answer each bar independently. It also
took its worst consequence with it: a one-sided model previously got ONE
trade per backtest, a hard blocker on marketplace validation.
- SignalClusterWindow now defaults off - it de-duplicated repeats that are
now real trades. Kept as an opt-in display control.
- LABEL_WINDOW_BARS, the pivot-widening pass, ConfirmedZigZagLabel.
- the era-0 output-bias seed now needs a genuinely dominant class (0.70)
rather than 0.40; at ~50% Neutral a +-3.0 seed is a distortion, not a
correction.
Also fixed, both found while wiring the above:
1. RefreshConvergedSignal sized its buffers from a date delta
(Bars(sym, period, dtStudied, TimeCurrent())). dtStudied is a training
watermark; in the tester it is loaded from a live-chart save AHEAD of
the simulated date, so the interval inverted, Bars() returned ~0, and
the buffer came out at exactly m_historyBars - deep enough for the OHLC
window and far too shallow for the Donchian-50 / 20-bar-return / SMA
extension behind it. Inference silently computed DIFFERENT features
from the ones training learned on, live as well as in the tester. Now
sized from what the feature builder actually needs.
2. The barrier horizon is resolved on the deployed path too. A deployed
model never enters Train(), so it never reached the prebuild, and
OnlineLearnStep reads the horizon as its confirmation delay - left at
the fallback it would have backpropped bars whose barriers had not
resolved. Silent lookahead in the one place that writes to a live model.
SL_Mode/TP_Mode join the weights fingerprint: they define the labels now,
so a model trained at 1:3 must never be silently reused at 1:1. This
re-keys every pre-existing model by design - none were trained on this task.
Inference census extended with the vote gate. LongCondition/ShortCondition
open with a readiness check the refresh counters never see; in the tester it
reduces to "the seeded _optcache.nnw must have LOADED", and if it did not,
every vote is hard-zeroed while the model still answers Buy. The old three
counters would have read that as "the model says Neutral" - false, and a
completely different fix. This is the leading candidate for the
zero-direction backtest and the census can now name it in one run.
Both builds compile 0 errors / 0 warnings. Forces a full retrain.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 20:39:49 -04:00
|
|
|
}
|
|
|
|
|
//+------------------------------------------------------------------+
|
|
|
|
|
//| TRIPLE-BARRIER LABEL for one bar (Lopez de Prado ch. 3). See the |
|
|
|
|
|
//| BARRIER_TIE_GOES_TO_STOP block in Expert\ExpertSignalAIBase.mqh |
|
|
|
|
|
//| for why this replaced the exact-pivot ZigZag target. |
|
|
|
|
|
//| |
|
|
|
|
|
//| Hypothetical entry at bar `idx`'s CLOSE - the same instant the |
|
|
|
|
|
//| feature window ends, so the label answers exactly the question |
|
|
|
|
|
//| the deployed model is asked live: "from what I can see right now, |
|
|
|
|
|
//| does a trade placed here reach its target before its stop?" |
|
|
|
|
|
//| |
|
|
|
|
|
//| Costs are charged. MT5 bar series are BID, so a long fills at ask |
|
|
|
|
|
//| (close + spread) and exits at bid, while a short fills at bid and |
|
|
|
|
|
//| buys back at ask - both legs shifted so the returned outcome is a |
|
|
|
|
|
//| NET result. Spread is taken as the symbol's current value, held |
|
|
|
|
|
//| constant across history: MT5's standard timeseries carries no |
|
|
|
|
|
//| per-bar spread, and a label that ignored the cost entirely would |
|
|
|
|
|
//| report a win rate the account cannot reproduce. |
|
|
|
|
|
//| |
|
|
|
|
|
//| Walks forward in time (toward index 0) for m_barrierHorizonBars. |
|
|
|
|
|
//| Ties inside one bar resolve to the STOP - OHLC cannot order two |
|
|
|
|
|
//| touches within a bar, and the optimistic reading is how a |
|
|
|
|
|
//| backtested edge becomes a live loss. |
|
|
|
|
|
//+------------------------------------------------------------------+
|
|
|
|
|
ENUM_SIGNAL CExpertSignalAIBase::TripleBarrierLabel(int idx)
|
|
|
|
|
{
|
diag: is "optimal SL/TP" learnable? Score the features against excursions
Proposed direction: train the net to predict entry/SL/TP that maximise return
and minimise drawdown, rather than to classify direction. Before rebuilding a
head, measure whether the target is learnable at all.
That question splits into two that behave nothing alike:
HOW FAR price travels (MFE/MAE) - essentially volatility, and volatility
clustering is about the most robust regularity in markets.
WHICH WAY it goes first (the asymmetry) - direction, which is what every
noise-floor verdict in this project has been about.
Expectancy comes ONLY from the second. The first buys position sizing and
drawdown control - worth having under prop-firm limits, but not an edge: exit
management on RANDOM entries already moved the payoff ratio 0.92 -> 5.72 with
expectancy FLAT.
Crucially this is NOT already answered. Every MI figure here scored the
triple-barrier label, i.e. one specific question at one fixed geometry. A
noise-floor result there says nothing about whether excursion MAGNITUDE is
learnable - different target, different answer.
Four targets, and the verdict is the CONTRAST, printed explicitly because the
dangerous misreading of "UP clears" is "we can predict profitable trades":
RANGE (up+dn) - realised volatility, included as a POSITIVE CONTROL that
SHOULD clear. Every prior verdict here lacked a control
expected to pass; a range target at the floor indicts the
measurement, not the market.
UP / DOWN - MFE / MAE.
ASYMMETRY - up-dn, the only one that can pay.
Collected inside the walk the label already does (one max, one min per bar).
The early-out when both barriers resolved is GONE: it would have truncated the
excursions at whichever bar tripped the last barrier, making the measurement a
function of the CURRENT SL/TP - the circularity this is trying to escape. The
loop was already bounded by the horizon, so only the average cost moves.
Discretised into 3 EQUAL-FREQUENCY bins, so every downstream piece (block
permutation, null, p-value) is reused unchanged. Equal-frequency because MFE is
fat-tailed and fixed-width bins would put nearly every row in bin 0; it also
pins H(Y) at ln(3)=1.099 for all four, making them comparable to each other and
to the barrier label's ~1.02 instead of confounded by class balance.
Two bugs fixed in this code before it ever ran, both of which would have
produced a plausible quiet wrong answer rather than an error:
- TripleBarrierLabel early-returns on invalid ATR/close BEFORE the point the
accumulators were reset, so one bar's excursions would be cached under
another bar's index. Cleared at the top now, ahead of every return.
- An unresolvable bar is still flagged as labelled but carries excursions of
exactly 0. Under equal-frequency binning a block of identical zeros drags
the lowest cut onto zero and a third of the sample lands in one
uninformative bin - a depressed score that reads as "not predictable", a
false negative in the direction that would wrongly kill the idea. Rows
where both excursions are zero are dropped; price cannot travel zero both
ways over a whole horizon.
Read-only diagnostic. No topology or label change: no retrain of its own.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-07 10:22:41 -04:00
|
|
|
//--- CLEARED FIRST, ahead of every early return below. These are published to the caller the way
|
|
|
|
|
//--- m_lastBarrierTimedOut is, and an unresolvable bar that returned before touching them would leave
|
|
|
|
|
//--- the PREVIOUS bar's excursions in place for AdvanceBarrierLabelState to cache against this index -
|
|
|
|
|
//--- one bar's outcome filed under another's, which is exactly the kind of silent contamination the
|
|
|
|
|
//--- excursion measurement is being built to avoid.
|
|
|
|
|
m_lastExcUp = 0.0;
|
|
|
|
|
m_lastExcDown = 0.0;
|
fix: both-won bars were labelled "do not trade" - resolve by first touch
Removing the min-reward:risk raise let the MEASURED geometry come back with
the target NEARER than the stop (SP500 H1: target 1.62*ATR at q50 of
favourable, stop 3.33*ATR at q75 of adverse). That reopened a branch the
code called unreachable: price can reach +target and -target inside one
horizon, winning in BOTH directions, and those bars fell through to Neutral.
Neutral has only three producers, both-lost is unreachable (you cannot touch
-3.33 without crossing -1.62 first, which wins the short), and timeouts logged
at 1.0% of Neutral - so ~27% of ALL bars were being handed to the model as the
abstain class when a trade either way would have collected its target. The
cleanest positives in the sample, labelled "do not trade", while the fitted
confidence threshold was being asked to find selectivity in what was left.
Resolved by FIRST TOUCH: the target reached earlier is the trade that would
have closed first. Same forward window, no extra lookahead. Same-bar ties stay
Neutral - OHLC cannot order two touches, and unlike an intrabar stop tie there
is no pessimistic side to fall to, so a guess would inject a coin-flip
direction into the target.
Also:
- count both-won and its same-bar tie subset in the prebuild line, so the
share is measured rather than inferred from arithmetic on a log line
- scope the timeout counter to IS, matching the tally it is reported as a
percentage OF; it was incremented over the whole scan and divided by an
in-sample denominator
- clear m_lastBarrierTimedOut at the top of the walk with the excursions, not
at the bottom - the two early returns published the previous bar's verdict
- mark the pass-1 label line PROVISIONAL. It prints the enum fallback because
geometry can only be derived from excursions that do not exist yet, and it
reads exactly like a config change that failed to take effect
FORCES RETRAIN. Both build variants compile 0 errors, 0 warnings.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 15:35:42 -04:00
|
|
|
//--- MOVED UP from the bottom of the walk (2026-08-09) for exactly the reason written above about the
|
|
|
|
|
//--- excursions: the two early returns below this line return WITHOUT reaching the assignment that
|
|
|
|
|
//--- used to be the only one, so an unresolvable bar published the PREVIOUS bar's timeout verdict. The
|
|
|
|
|
//--- both-won flags are new and are cleared here from the start rather than inheriting that bug.
|
|
|
|
|
m_lastBarrierTimedOut = false;
|
|
|
|
|
m_lastBarrierBothWon = false;
|
|
|
|
|
m_lastBarrierBothWonTied = false;
|
fix: the deploy gate was benchmarking a win rate against a label frequency
The gate rests on an invariant stated at ExpertSignalAIBase.mqh:199 - under a
driftless walk P(touch +k before -m) is m/(m+k), and break-even for a k:m trade
is ALSO m/(m+k), so "beats chance" and "is profitable" are the same test.
That invariant needs reward >= risk, and the measured geometry no longer
satisfies it. With target 1.62*ATR and stop 3.33*ATR, break-even is 67.3%, but
both-won bars were stripped out of Buy and Sell so the label base rate read
37.5%. chancePrecPct is max(BuyTotal,SellTotal)/bars, so the gate was clearing
models nearly 30pp short of break-even: 42% "directional precision" is +4 sigma
against 37.5% and loses money on every single trade. Live since 217b9bc.
Root cause is that label agreement stopped being the same question as trade
profitability. Buy implies winLong, but the converse fails on every both-won
bar, and the label can only name one of two directions that both pay.
So stop asking the model whether it matched a label and start asking whether
its trade paid:
- cache winLong/winShort per bar beside the label, under the same validity
flag; published from the barrier walk before the collapse to 3 classes
- dirPrecPct now counts wins on the side actually called
- chancePrecPct is max(P(winLong), P(winShort)), MEASURED - the textbook
m/(m+k) would credit SP500's drift to the model
- the NMS "what would I have made" pair, the live-fired precision, and the
IS/OOS cumulative win rates all move to the same test. IS and OOS are read
side by side as the overfitting signal, so measuring one in wins and the
other in agreement would put a fixed gap between them that has nothing to do
with generalization
- the confidence threshold is FITTED on wins too, so the operating point
maximises what the gate grades
- per-class label-agreement precision is still computed and logged; it is the
right diagnostic for class separation, just not for a deploy decision
- era line renamed dir-precision -> win-rate, chance -> chance=break-even
Both build variants compile 0 errors, 0 warnings.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 16:00:51 -04:00
|
|
|
m_lastWinLong = false;
|
|
|
|
|
m_lastWinShort = false;
|
fix(labels): overlapping-label sample correction + horizon cap on the scale ladder
Three defects, all surfaced by the 2026-08-17 SP500 H4 run that shipped
stop 4.86 / target 9.71 (width 14.57*ATR, horizon 384).
1. EVERY STANDARD ERROR ASSUMED INDEPENDENT SAMPLES. Triple-barrier labels
started one per bar overlap by the label's lifespan, so n calls are worth
~n/L independent observations (Lopez de Prado, AFML ch. 4 - sample
uniqueness). All three sqrt(p(1-p)/n) sites divided by the RAW count.
The tell: the operating point's null-of-the-maximum gate is family-wise and
should fire on ~5% of eras under the null. Measured fire rates - PAI 47/73
(64%), ConvLSTM 9/24, LSTM 8/21 (38%), CONV 4/62 (6%). CONV, the only model
whose margin distribution admits few bins, sat on the null; the rest cleared
a bar that was too low by ~sqrt(L). PAI's deployed threshold consequently
alternated between the ENDS of its own range era to era (0.10 -> 0.88 ->
0.86 -> 0.66; coverage 16% <-> 73%).
TripleBarrierLabel now records when each label became KNOWABLE - the first
winning touch, or both stops, or the timeout - and the prebuild accumulates
the mean. EffectiveSampleSize() feeds the operating point, the member deploy
gate and the ensemble vote gate. Conservative by construction (n/L is an
upper bound on the damage); gates get harder, never easier.
2. THE SCALE LADDER RAN AWAY, again. Horizon scales as swingMedian*sl*tp, and
since 4d8cb08 reachability is measured OVER that horizon - so a wider rung
buys itself the time that makes it look reachable. Same target -> horizon ->
reach -> target loop the excursion window is kept short to avoid; fixing the
window confusion reopened it through the other door. It walked 128 -> 256 ->
384 bars and stopped at q90, the widest rung there is, with every rung
reading 39-48% against a 20% floor. A floor nothing fails selects nothing.
Rungs whose required horizon exceeds BARRIER_HORIZON_MAX are now rejected -
the same rule ReportGeometryExpectancyScan already applied. It was printing
the shipped pair as CLAMPED and disqualified ('h384!') two lines under the
deriver that chose it: two subsystems, one geometry, opposite verdicts.
3. THE RUNG SNAP DESTROYED THE RATIO IT WAS COMPARING. Both legs snapped
independently to the coarse first-passage grid, re-rating each candidate:
q90 4.86/9.71 -> 5.00/10.00 (2.00), q85 4.07/8.14 -> 5.00/10.00 (IDENTICAL
measurement), q75 3.07/6.13 -> 4.00/6.50 (1.63 - a nearer target). So the
ladder compared win shares taken at ratios from 1.63 to 2.17 and read the
differences as scale. It is why the reach column came out non-monotone in
width (q75 48.5% above q90 42.9%). The stop now snaps to its nearest rung in
log space and the target follows the ratio off it; the pair actually measured
is returned and logged, so a collision reads as a collision.
Also: LadderWinShare guarded against the conditional (fractal) geometry path,
which fills n from m_fracLegCount while leaving idxList empty - a latent
out-of-bounds on a currently-dead path.
New log lines: mean label lifespan and effective n on the label-cache line, the
required-vs-available horizon per rung, and the grid pair the reconciliation
actually measured (its tolerance now scales with the grid skew instead of a flat
5pp).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 01:12:05 -04:00
|
|
|
//--- Cleared with the rest, and for the same reason: an early return must not leave the PREVIOUS bar's
|
|
|
|
|
//--- lifespan for the prebuild to accumulate. 0 = not measured, which the accumulator skips.
|
|
|
|
|
m_lastLabelLifespan = 0;
|
feat(ai): triple-barrier labels replace exact-pivot ZigZag targets
The 31:1 class imbalance was self-inflicted by the TARGET, not a property
of the market. Labelling only the exact bar where a ZigZag pivot confirms
gave Buy 1164 / Sell 1164 / Neutral 35841, and every correction mechanism
this codebase accumulated sits downstream of that one choice: the
logit-adjusted loss and its range cap, the prior EMA, the +-3.0 output-bias
seed, balanced-accuracy-then-precision selection with its coverage floor,
the recall floor and its catch-22, the alternation gate, NMS, and the four
oversampling designs that collapsed before them.
The reference this engine is built on (references/neuronetworksbook.pdf
ch. 3.1/3.3) also uses ZigZag, but targets the DIRECTION TO THE NEXT
EXTREMUM on every bar - ~50/50 by construction, with no imbalance to
correct at all. It never had this problem because it never asked "is this
the pivot bar".
Labels are now the triple barrier (Lopez de Prado ch. 3), using the EA's
OWN SL_Mode/TP_Mode: does a trade opened at this bar's close reach its
target before its stop, within a horizon. Buy = long resolves, Sell =
short resolves, Neutral = neither. Consequences:
- dir-precision in the era line stops being a proxy and becomes the win
rate of the strategy under its own exit rules.
- Expected balance ~25/25/50 at the shipped 1:3 (gambler's ruin), i.e.
~2:1 instead of 31:1. Measured and logged at the end of the prebuild.
- Spread is charged on both legs, so it is a NET win rate.
- Intrabar ambiguity resolves to the STOP. OHLC cannot order two touches
inside one bar and the optimistic reading is how a backtested edge
becomes a live loss.
ZigZag stays as input features (EnableSwingContext) and now also supplies
the vertical barrier: the horizon is the median confirmed leg length,
snapped to a coarse ladder. Derived, not configured, and deliberately kept
out of the filename fingerprint - a filename keyed on a measured quantity
orphans a trained model the moment the measurement moves.
Removed, because the premise died with the old target:
- the alternation gate. Correct for pivot labels (a ZigZag cannot emit two
same-type pivots in a row, so a repeat was provably a false fire), and
wrong for barrier labels, which answer each bar independently. It also
took its worst consequence with it: a one-sided model previously got ONE
trade per backtest, a hard blocker on marketplace validation.
- SignalClusterWindow now defaults off - it de-duplicated repeats that are
now real trades. Kept as an opt-in display control.
- LABEL_WINDOW_BARS, the pivot-widening pass, ConfirmedZigZagLabel.
- the era-0 output-bias seed now needs a genuinely dominant class (0.70)
rather than 0.40; at ~50% Neutral a +-3.0 seed is a distortion, not a
correction.
Also fixed, both found while wiring the above:
1. RefreshConvergedSignal sized its buffers from a date delta
(Bars(sym, period, dtStudied, TimeCurrent())). dtStudied is a training
watermark; in the tester it is loaded from a live-chart save AHEAD of
the simulated date, so the interval inverted, Bars() returned ~0, and
the buffer came out at exactly m_historyBars - deep enough for the OHLC
window and far too shallow for the Donchian-50 / 20-bar-return / SMA
extension behind it. Inference silently computed DIFFERENT features
from the ones training learned on, live as well as in the tester. Now
sized from what the feature builder actually needs.
2. The barrier horizon is resolved on the deployed path too. A deployed
model never enters Train(), so it never reached the prebuild, and
OnlineLearnStep reads the horizon as its confirmation delay - left at
the fallback it would have backpropped bars whose barriers had not
resolved. Silent lookahead in the one place that writes to a live model.
SL_Mode/TP_Mode join the weights fingerprint: they define the labels now,
so a model trained at 1:3 must never be silently reused at 1:1. This
re-keys every pre-existing model by design - none were trained on this task.
Inference census extended with the vote gate. LongCondition/ShortCondition
open with a readiness check the refresh counters never see; in the tester it
reduces to "the seeded _optcache.nnw must have LOADED", and if it did not,
every vote is hard-zeroed while the model still answers Buy. The old three
counters would have read that as "the model says Neutral" - false, and a
completely different fix. This is the leading candidate for the
zero-direction backtest and the census can now name it in one run.
Both builds compile 0 errors / 0 warnings. Forces a full retrain.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 20:39:49 -04:00
|
|
|
double atr = m_ATR.Main(idx);
|
|
|
|
|
if(!MathIsValidNumber(atr) || atr <= 0.0)
|
|
|
|
|
return Neutral; // no volatility scale yet - unresolvable, same practical answer as "no setup"
|
|
|
|
|
double entry = m_Close.GetData(idx);
|
|
|
|
|
if(!MathIsValidNumber(entry) || entry <= 0.0)
|
|
|
|
|
return Neutral;
|
|
|
|
|
double slMult, tpMult;
|
|
|
|
|
BarrierMultiples(slMult, tpMult);
|
|
|
|
|
double risk = slMult * atr;
|
|
|
|
|
double reward = tpMult * atr;
|
|
|
|
|
//--- Round-trip cost, in price. Both sides pay it once.
|
|
|
|
|
double spread = (double)m_symbol.Spread() * m_symbol.Point();
|
|
|
|
|
if(!MathIsValidNumber(spread) || spread < 0.0)
|
|
|
|
|
spread = 0.0;
|
|
|
|
|
//--- Barrier levels expressed in BID terms, which is what m_High/m_Low carry.
|
|
|
|
|
//--- Long fills at close+spread: target needs bid >= fill+reward, stop trips at bid <= fill-risk.
|
|
|
|
|
//--- Short fills at close: target needs bid <= close-reward-spread (it buys back at ask),
|
|
|
|
|
//--- stop trips at bid >= close+risk-spread.
|
|
|
|
|
double longTp = entry + spread + reward;
|
|
|
|
|
double longSl = entry + spread - risk;
|
|
|
|
|
double shortTp = entry - reward - spread;
|
|
|
|
|
double shortSl = entry + risk - spread;
|
|
|
|
|
bool longWon = false, longLost = false, shortWon = false, shortLost = false;
|
fix: both-won bars were labelled "do not trade" - resolve by first touch
Removing the min-reward:risk raise let the MEASURED geometry come back with
the target NEARER than the stop (SP500 H1: target 1.62*ATR at q50 of
favourable, stop 3.33*ATR at q75 of adverse). That reopened a branch the
code called unreachable: price can reach +target and -target inside one
horizon, winning in BOTH directions, and those bars fell through to Neutral.
Neutral has only three producers, both-lost is unreachable (you cannot touch
-3.33 without crossing -1.62 first, which wins the short), and timeouts logged
at 1.0% of Neutral - so ~27% of ALL bars were being handed to the model as the
abstain class when a trade either way would have collected its target. The
cleanest positives in the sample, labelled "do not trade", while the fitted
confidence threshold was being asked to find selectivity in what was left.
Resolved by FIRST TOUCH: the target reached earlier is the trade that would
have closed first. Same forward window, no extra lookahead. Same-bar ties stay
Neutral - OHLC cannot order two touches, and unlike an intrabar stop tie there
is no pessimistic side to fall to, so a guess would inject a coin-flip
direction into the target.
Also:
- count both-won and its same-bar tie subset in the prebuild line, so the
share is measured rather than inferred from arithmetic on a log line
- scope the timeout counter to IS, matching the tally it is reported as a
percentage OF; it was incremented over the whole scan and divided by an
in-sample denominator
- clear m_lastBarrierTimedOut at the top of the walk with the excursions, not
at the bottom - the two early returns published the previous bar's verdict
- mark the pass-1 label line PROVISIONAL. It prints the enum fallback because
geometry can only be derived from excursions that do not exist yet, and it
reads exactly like a config change that failed to take effect
FORCES RETRAIN. Both build variants compile 0 errors, 0 warnings.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 15:35:42 -04:00
|
|
|
//--- Bar index at which each target was FIRST reached, for the both-won resolution below. The walk
|
|
|
|
|
//--- runs t = idx-1 downward, i.e. forward in time, so the LARGER t is the earlier touch.
|
|
|
|
|
int longWonAt = -1, shortWonAt = -1;
|
fix(labels): overlapping-label sample correction + horizon cap on the scale ladder
Three defects, all surfaced by the 2026-08-17 SP500 H4 run that shipped
stop 4.86 / target 9.71 (width 14.57*ATR, horizon 384).
1. EVERY STANDARD ERROR ASSUMED INDEPENDENT SAMPLES. Triple-barrier labels
started one per bar overlap by the label's lifespan, so n calls are worth
~n/L independent observations (Lopez de Prado, AFML ch. 4 - sample
uniqueness). All three sqrt(p(1-p)/n) sites divided by the RAW count.
The tell: the operating point's null-of-the-maximum gate is family-wise and
should fire on ~5% of eras under the null. Measured fire rates - PAI 47/73
(64%), ConvLSTM 9/24, LSTM 8/21 (38%), CONV 4/62 (6%). CONV, the only model
whose margin distribution admits few bins, sat on the null; the rest cleared
a bar that was too low by ~sqrt(L). PAI's deployed threshold consequently
alternated between the ENDS of its own range era to era (0.10 -> 0.88 ->
0.86 -> 0.66; coverage 16% <-> 73%).
TripleBarrierLabel now records when each label became KNOWABLE - the first
winning touch, or both stops, or the timeout - and the prebuild accumulates
the mean. EffectiveSampleSize() feeds the operating point, the member deploy
gate and the ensemble vote gate. Conservative by construction (n/L is an
upper bound on the damage); gates get harder, never easier.
2. THE SCALE LADDER RAN AWAY, again. Horizon scales as swingMedian*sl*tp, and
since 4d8cb08 reachability is measured OVER that horizon - so a wider rung
buys itself the time that makes it look reachable. Same target -> horizon ->
reach -> target loop the excursion window is kept short to avoid; fixing the
window confusion reopened it through the other door. It walked 128 -> 256 ->
384 bars and stopped at q90, the widest rung there is, with every rung
reading 39-48% against a 20% floor. A floor nothing fails selects nothing.
Rungs whose required horizon exceeds BARRIER_HORIZON_MAX are now rejected -
the same rule ReportGeometryExpectancyScan already applied. It was printing
the shipped pair as CLAMPED and disqualified ('h384!') two lines under the
deriver that chose it: two subsystems, one geometry, opposite verdicts.
3. THE RUNG SNAP DESTROYED THE RATIO IT WAS COMPARING. Both legs snapped
independently to the coarse first-passage grid, re-rating each candidate:
q90 4.86/9.71 -> 5.00/10.00 (2.00), q85 4.07/8.14 -> 5.00/10.00 (IDENTICAL
measurement), q75 3.07/6.13 -> 4.00/6.50 (1.63 - a nearer target). So the
ladder compared win shares taken at ratios from 1.63 to 2.17 and read the
differences as scale. It is why the reach column came out non-monotone in
width (q75 48.5% above q90 42.9%). The stop now snaps to its nearest rung in
log space and the target follows the ratio off it; the pair actually measured
is returned and logged, so a collision reads as a collision.
Also: LadderWinShare guarded against the conditional (fractal) geometry path,
which fills n from m_fracLegCount while leaving idxList empty - a latent
out-of-bounds on a currently-dead path.
New log lines: mean label lifespan and effective n on the label-cache line, the
required-vs-available horizon per rung, and the grid pair the reconciliation
actually measured (its tolerance now scales with the grid skew instead of a flat
5pp).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 01:12:05 -04:00
|
|
|
//--- AGE (idx - t, so bars AFTER entry) at which each side first resolved either way. The won-at
|
|
|
|
|
//--- indices above cannot serve: they are bar indices rather than ages, and they say nothing about the
|
|
|
|
|
//--- losing side, which is what fixes a Neutral label's lifespan. See m_lastLabelLifespan.
|
|
|
|
|
int longEndAge = 0, shortEndAge = 0;
|
diag: is "optimal SL/TP" learnable? Score the features against excursions
Proposed direction: train the net to predict entry/SL/TP that maximise return
and minimise drawdown, rather than to classify direction. Before rebuilding a
head, measure whether the target is learnable at all.
That question splits into two that behave nothing alike:
HOW FAR price travels (MFE/MAE) - essentially volatility, and volatility
clustering is about the most robust regularity in markets.
WHICH WAY it goes first (the asymmetry) - direction, which is what every
noise-floor verdict in this project has been about.
Expectancy comes ONLY from the second. The first buys position sizing and
drawdown control - worth having under prop-firm limits, but not an edge: exit
management on RANDOM entries already moved the payoff ratio 0.92 -> 5.72 with
expectancy FLAT.
Crucially this is NOT already answered. Every MI figure here scored the
triple-barrier label, i.e. one specific question at one fixed geometry. A
noise-floor result there says nothing about whether excursion MAGNITUDE is
learnable - different target, different answer.
Four targets, and the verdict is the CONTRAST, printed explicitly because the
dangerous misreading of "UP clears" is "we can predict profitable trades":
RANGE (up+dn) - realised volatility, included as a POSITIVE CONTROL that
SHOULD clear. Every prior verdict here lacked a control
expected to pass; a range target at the floor indicts the
measurement, not the market.
UP / DOWN - MFE / MAE.
ASYMMETRY - up-dn, the only one that can pay.
Collected inside the walk the label already does (one max, one min per bar).
The early-out when both barriers resolved is GONE: it would have truncated the
excursions at whichever bar tripped the last barrier, making the measurement a
function of the CURRENT SL/TP - the circularity this is trying to escape. The
loop was already bounded by the horizon, so only the average cost moves.
Discretised into 3 EQUAL-FREQUENCY bins, so every downstream piece (block
permutation, null, p-value) is reused unchanged. Equal-frequency because MFE is
fat-tailed and fixed-width bins would put nearly every row in bin 0; it also
pins H(Y) at ln(3)=1.099 for all four, making them comparable to each other and
to the barrier label's ~1.02 instead of confounded by class balance.
Two bugs fixed in this code before it ever ran, both of which would have
produced a plausible quiet wrong answer rather than an error:
- TripleBarrierLabel early-returns on invalid ATR/close BEFORE the point the
accumulators were reset, so one bar's excursions would be cached under
another bar's index. Cleared at the top now, ahead of every return.
- An unresolvable bar is still flagged as labelled but carries excursions of
exactly 0. Under equal-frequency binning a block of identical zeros drags
the lowest cut onto zero and a third of the sample lands in one
uninformative bin - a depressed score that reads as "not predictable", a
false negative in the direction that would wrongly kill the idea. Rows
where both excursions are zero are dropped; price cannot travel zero both
ways over a whole horizon.
Read-only diagnostic. No topology or label change: no retrain of its own.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-07 10:22:41 -04:00
|
|
|
//--- Excursion accumulators. Deliberately NOT stopped when a barrier trips: they describe how far
|
|
|
|
|
//--- price travelled over the whole horizon, which is the question a predicted SL/TP needs answered,
|
|
|
|
|
//--- whereas the barriers describe what a trade with THIS geometry would have collected. Truncating
|
|
|
|
|
//--- them at the first touch would bake the current SL/TP back into the measurement of whether a
|
|
|
|
|
//--- different SL/TP is learnable - the circularity the whole exercise is trying to escape.
|
|
|
|
|
double maxHigh = -DBL_MAX, minLow = DBL_MAX; // published values already cleared at the top
|
feat: first-passage ladder + expectancy scan - price every geometry, not just the chosen one
Corrects the premise of the previous plan. Break-even is NOT a ceiling.
If the model shifts the win probability on the bars it selects from
p0 = m/(m+k) to p0 + d, then
EV = (p0+d)*k - (1-p0-d)*m = d*(k+m)
because p0*k - (1-p0)*m is zero by construction. The stop:target RATIO
is expectancy-neutral - a punishing break-even is exactly repaid by the
payoff - and only the real edge d and the TOTAL WIDTH (k+m) move EV.
Width matters because the spread is charged once per trade however wide
the barriers are, so a narrow barrier spends much of its own range on
costs. DeriveBarrierGeometry's own comment already said the ratio buys
nothing; the objective just never followed from it.
Blocker this had to solve first: m_excUpCache/m_excDownCache hold only
MAXIMUM travel each way, and a maximum cannot say which side was
reached FIRST - so any geometry other than the walked one was
undecidable on precisely the bars where both barriers were touched,
~28% of the sample.
- BARRIER_LADDER: per bar, the first-touch AGE for 8 travel distances
in each direction, filled during the walk the labels already run.
Cursors keep it O(1) amortised per walked bar rather than 16
comparisons. Levels are travel FROM ENTRY, not barrier prices, so one
ladder serves both directions and the spread is applied analytically
when a level converts back to an SL/TP multiple - storing prices
would need four ladders and bake today's spread into the cache.
Sized, invalidated and validity-gated with the label caches.
- ReportGeometryExpectancyScan: every ladder pair priced exactly off
that cache - width in ATR and in SPREADS (cost efficiency, knowable
without knowing d), break-even, both base rates, the share of bars
resolved inside the horizon, and EV per unit of edge. Compares the
widest resolvable pair against the quantile rule's pick.
MEASUREMENT ONLY - the quantile rule still chooses. Nothing here can
measure d, and width buys nothing if the wider target is less
predictable. Base rates are printed beside each break-even because a
persistent gap is DRIFT and must not be credited to the model.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-10 12:59:18 -04:00
|
|
|
//--- First-passage ladder for THIS bar (see BARRIER_LADDER). Cursors, not a full rescan: the ladder is
|
|
|
|
|
//--- ascending and travel is monotone in the running extreme, so once a level is passed no lower level
|
|
|
|
|
//--- can be reached later - each level is tested until it trips exactly once, which keeps this O(1)
|
|
|
|
|
//--- amortised per walked bar instead of 2 x BARRIER_LADDER_COUNT comparisons on every one.
|
|
|
|
|
ArrayInitialize(m_lastLadderUpAt, 0);
|
|
|
|
|
ArrayInitialize(m_lastLadderDownAt, 0);
|
|
|
|
|
int upCursor = 0, dnCursor = 0;
|
fix(labels): overlapping-label sample correction + horizon cap on the scale ladder
Three defects, all surfaced by the 2026-08-17 SP500 H4 run that shipped
stop 4.86 / target 9.71 (width 14.57*ATR, horizon 384).
1. EVERY STANDARD ERROR ASSUMED INDEPENDENT SAMPLES. Triple-barrier labels
started one per bar overlap by the label's lifespan, so n calls are worth
~n/L independent observations (Lopez de Prado, AFML ch. 4 - sample
uniqueness). All three sqrt(p(1-p)/n) sites divided by the RAW count.
The tell: the operating point's null-of-the-maximum gate is family-wise and
should fire on ~5% of eras under the null. Measured fire rates - PAI 47/73
(64%), ConvLSTM 9/24, LSTM 8/21 (38%), CONV 4/62 (6%). CONV, the only model
whose margin distribution admits few bins, sat on the null; the rest cleared
a bar that was too low by ~sqrt(L). PAI's deployed threshold consequently
alternated between the ENDS of its own range era to era (0.10 -> 0.88 ->
0.86 -> 0.66; coverage 16% <-> 73%).
TripleBarrierLabel now records when each label became KNOWABLE - the first
winning touch, or both stops, or the timeout - and the prebuild accumulates
the mean. EffectiveSampleSize() feeds the operating point, the member deploy
gate and the ensemble vote gate. Conservative by construction (n/L is an
upper bound on the damage); gates get harder, never easier.
2. THE SCALE LADDER RAN AWAY, again. Horizon scales as swingMedian*sl*tp, and
since 4d8cb08 reachability is measured OVER that horizon - so a wider rung
buys itself the time that makes it look reachable. Same target -> horizon ->
reach -> target loop the excursion window is kept short to avoid; fixing the
window confusion reopened it through the other door. It walked 128 -> 256 ->
384 bars and stopped at q90, the widest rung there is, with every rung
reading 39-48% against a 20% floor. A floor nothing fails selects nothing.
Rungs whose required horizon exceeds BARRIER_HORIZON_MAX are now rejected -
the same rule ReportGeometryExpectancyScan already applied. It was printing
the shipped pair as CLAMPED and disqualified ('h384!') two lines under the
deriver that chose it: two subsystems, one geometry, opposite verdicts.
3. THE RUNG SNAP DESTROYED THE RATIO IT WAS COMPARING. Both legs snapped
independently to the coarse first-passage grid, re-rating each candidate:
q90 4.86/9.71 -> 5.00/10.00 (2.00), q85 4.07/8.14 -> 5.00/10.00 (IDENTICAL
measurement), q75 3.07/6.13 -> 4.00/6.50 (1.63 - a nearer target). So the
ladder compared win shares taken at ratios from 1.63 to 2.17 and read the
differences as scale. It is why the reach column came out non-monotone in
width (q75 48.5% above q90 42.9%). The stop now snaps to its nearest rung in
log space and the target follows the ratio off it; the pair actually measured
is returned and logged, so a collision reads as a collision.
Also: LadderWinShare guarded against the conditional (fractal) geometry path,
which fills n from m_fracLegCount while leaving idxList empty - a latent
out-of-bounds on a currently-dead path.
New log lines: mean label lifespan and effective n on the label-cache line, the
required-vs-available horizon per rung, and the grid pair the reconciliation
actually measured (its tolerance now scales with the grid skew instead of a flat
5pp).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 01:12:05 -04:00
|
|
|
//--- Age of the LAST bar the loop actually visited. Not simply the horizon: the walk breaks early when
|
|
|
|
|
//--- it runs off loaded history, and a timeout lifespan of "the full horizon" would then be longer than
|
|
|
|
|
//--- the window that was examined.
|
|
|
|
|
int walkedAge = 0;
|
fix: excursion window must not depend on the barrier it sizes
DIRECTION IS NOT THERE, and this run is what establishes it. Three symbols:
raw ASYMMETRY clears on all three (p=0.0199 / 0.0050 / 0.0050)
norm ASYMMETRY collapses on all three (p=0.3433 / 0.5075 / 0.2736),
USDCAD landing BELOW its own null
RANGE control strengthens to 3-5x its null everywhere
Divide sigma out and the apparent directional signal vanishes entirely. What
cleared was volatility leaking through an unnormalised difference. Note this
would have passed any replication test: three instruments at p=0.005 is exactly
the evidence one would accept before committing to a rebuild, and the confound
reproduces perfectly. Replication was never going to catch it - only the
normalisation could.
Two defects of mine, both surfaced by the same run.
1. THE GEOMETRY DERIVATION WAS DIVERGING, NOT CONVERGING. It produced a
14.57*ATR stop and a 29.14*ATR target that only 5.7% of bars ever reach.
Excursions were measured over the barrier horizon; the horizon scales with
the target; the target is a quantile of the excursions - so target ->
horizon -> excursions -> target ran away, and "settled" only because the
horizon ladder caps at 384 bars. A saturated runaway, which the iteration
guard could not catch because it watches for OSCILLATION.
Fixed at the root: excursions now accumulate only over m_swingMedianBars -
the UNSCALED median ZigZag leg, a property of the instrument that owes
nothing to the barrier. The barrier walk still runs the full horizon,
because that is how long the trade is held; only the MEASUREMENT used to
size the barrier is confined to a geometry-independent window.
(The Min_Risk_Reward_Ratio warning fired correctly and is what flagged it -
the diagnostic worked while the derivation behind it did not.)
2. THE CONFOUND VERDICT WAS UNREACHABLE. `sizeCleared && !asymCleared` was
tested first and is true whenever size clears - i.e. always - so the branch
that NAMES the volatility confound never printed; all three symbols showed
the generic size-not-direction message instead. Verdict chain rewritten with
the specific case first, and the dangling elses my first patch introduced
removed.
FORCES A FULL RETRAIN (the excursion window changes every derived barrier).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-07 13:57:23 -04:00
|
|
|
//--- Never longer than the horizon actually walked, so the window cannot claim bars the loop below
|
|
|
|
|
//--- does not visit; falls back to the horizon before the swing median has been measured.
|
|
|
|
|
int excWindow = (m_swingMedianBars > 0)
|
|
|
|
|
? (int)MathMin(m_swingMedianBars, MathMax(m_barrierHorizonBars, 1))
|
|
|
|
|
: MathMax(m_barrierHorizonBars, 1);
|
feat(ai): triple-barrier labels replace exact-pivot ZigZag targets
The 31:1 class imbalance was self-inflicted by the TARGET, not a property
of the market. Labelling only the exact bar where a ZigZag pivot confirms
gave Buy 1164 / Sell 1164 / Neutral 35841, and every correction mechanism
this codebase accumulated sits downstream of that one choice: the
logit-adjusted loss and its range cap, the prior EMA, the +-3.0 output-bias
seed, balanced-accuracy-then-precision selection with its coverage floor,
the recall floor and its catch-22, the alternation gate, NMS, and the four
oversampling designs that collapsed before them.
The reference this engine is built on (references/neuronetworksbook.pdf
ch. 3.1/3.3) also uses ZigZag, but targets the DIRECTION TO THE NEXT
EXTREMUM on every bar - ~50/50 by construction, with no imbalance to
correct at all. It never had this problem because it never asked "is this
the pivot bar".
Labels are now the triple barrier (Lopez de Prado ch. 3), using the EA's
OWN SL_Mode/TP_Mode: does a trade opened at this bar's close reach its
target before its stop, within a horizon. Buy = long resolves, Sell =
short resolves, Neutral = neither. Consequences:
- dir-precision in the era line stops being a proxy and becomes the win
rate of the strategy under its own exit rules.
- Expected balance ~25/25/50 at the shipped 1:3 (gambler's ruin), i.e.
~2:1 instead of 31:1. Measured and logged at the end of the prebuild.
- Spread is charged on both legs, so it is a NET win rate.
- Intrabar ambiguity resolves to the STOP. OHLC cannot order two touches
inside one bar and the optimistic reading is how a backtested edge
becomes a live loss.
ZigZag stays as input features (EnableSwingContext) and now also supplies
the vertical barrier: the horizon is the median confirmed leg length,
snapped to a coarse ladder. Derived, not configured, and deliberately kept
out of the filename fingerprint - a filename keyed on a measured quantity
orphans a trained model the moment the measurement moves.
Removed, because the premise died with the old target:
- the alternation gate. Correct for pivot labels (a ZigZag cannot emit two
same-type pivots in a row, so a repeat was provably a false fire), and
wrong for barrier labels, which answer each bar independently. It also
took its worst consequence with it: a one-sided model previously got ONE
trade per backtest, a hard blocker on marketplace validation.
- SignalClusterWindow now defaults off - it de-duplicated repeats that are
now real trades. Kept as an opt-in display control.
- LABEL_WINDOW_BARS, the pivot-widening pass, ConfirmedZigZagLabel.
- the era-0 output-bias seed now needs a genuinely dominant class (0.70)
rather than 0.40; at ~50% Neutral a +-3.0 seed is a distortion, not a
correction.
Also fixed, both found while wiring the above:
1. RefreshConvergedSignal sized its buffers from a date delta
(Bars(sym, period, dtStudied, TimeCurrent())). dtStudied is a training
watermark; in the tester it is loaded from a live-chart save AHEAD of
the simulated date, so the interval inverted, Bars() returned ~0, and
the buffer came out at exactly m_historyBars - deep enough for the OHLC
window and far too shallow for the Donchian-50 / 20-bar-return / SMA
extension behind it. Inference silently computed DIFFERENT features
from the ones training learned on, live as well as in the tester. Now
sized from what the feature builder actually needs.
2. The barrier horizon is resolved on the deployed path too. A deployed
model never enters Train(), so it never reached the prebuild, and
OnlineLearnStep reads the horizon as its confirmation delay - left at
the fallback it would have backpropped bars whose barriers had not
resolved. Silent lookahead in the one place that writes to a live model.
SL_Mode/TP_Mode join the weights fingerprint: they define the labels now,
so a model trained at 1:3 must never be silently reused at 1:1. This
re-keys every pre-existing model by design - none were trained on this task.
Inference census extended with the vote gate. LongCondition/ShortCondition
open with a readiness check the refresh counters never see; in the tester it
reduces to "the seeded _optcache.nnw must have LOADED", and if it did not,
every vote is hard-zeroed while the model still answers Buy. The old three
counters would have read that as "the model says Neutral" - false, and a
completely different fix. This is the leading candidate for the
zero-direction backtest and the census can now name it in one run.
Both builds compile 0 errors / 0 warnings. Forces a full retrain.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 20:39:49 -04:00
|
|
|
int last = idx - MathMax(m_barrierHorizonBars, 1);
|
|
|
|
|
if(last < 0)
|
|
|
|
|
last = 0;
|
|
|
|
|
for(int t = idx - 1; t >= last; t--)
|
|
|
|
|
{
|
|
|
|
|
double hi = m_High.GetData(t);
|
|
|
|
|
double lo = m_Low.GetData(t);
|
|
|
|
|
if(!MathIsValidNumber(hi) || !MathIsValidNumber(lo) || hi == EMPTY_VALUE || lo == EMPTY_VALUE)
|
|
|
|
|
break; // ran off loaded history - whatever resolved so far stands, the rest times out
|
fix: excursion window must not depend on the barrier it sizes
DIRECTION IS NOT THERE, and this run is what establishes it. Three symbols:
raw ASYMMETRY clears on all three (p=0.0199 / 0.0050 / 0.0050)
norm ASYMMETRY collapses on all three (p=0.3433 / 0.5075 / 0.2736),
USDCAD landing BELOW its own null
RANGE control strengthens to 3-5x its null everywhere
Divide sigma out and the apparent directional signal vanishes entirely. What
cleared was volatility leaking through an unnormalised difference. Note this
would have passed any replication test: three instruments at p=0.005 is exactly
the evidence one would accept before committing to a rebuild, and the confound
reproduces perfectly. Replication was never going to catch it - only the
normalisation could.
Two defects of mine, both surfaced by the same run.
1. THE GEOMETRY DERIVATION WAS DIVERGING, NOT CONVERGING. It produced a
14.57*ATR stop and a 29.14*ATR target that only 5.7% of bars ever reach.
Excursions were measured over the barrier horizon; the horizon scales with
the target; the target is a quantile of the excursions - so target ->
horizon -> excursions -> target ran away, and "settled" only because the
horizon ladder caps at 384 bars. A saturated runaway, which the iteration
guard could not catch because it watches for OSCILLATION.
Fixed at the root: excursions now accumulate only over m_swingMedianBars -
the UNSCALED median ZigZag leg, a property of the instrument that owes
nothing to the barrier. The barrier walk still runs the full horizon,
because that is how long the trade is held; only the MEASUREMENT used to
size the barrier is confined to a geometry-independent window.
(The Min_Risk_Reward_Ratio warning fired correctly and is what flagged it -
the diagnostic worked while the derivation behind it did not.)
2. THE CONFOUND VERDICT WAS UNREACHABLE. `sizeCleared && !asymCleared` was
tested first and is true whenever size clears - i.e. always - so the branch
that NAMES the volatility confound never printed; all three symbols showed
the generic size-not-direction message instead. Verdict chain rewritten with
the specific case first, and the dangling elses my first patch introduced
removed.
FORCES A FULL RETRAIN (the excursion window changes every derived barrier).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-07 13:57:23 -04:00
|
|
|
//--- Excursions accumulate only over the REFERENCE WINDOW, not the whole barrier horizon - see
|
|
|
|
|
//--- m_swingMedianBars. The barrier walk below still runs the full horizon, because that is how
|
|
|
|
|
//--- long the trade is actually held; only the MEASUREMENT used to size the barrier is confined to
|
|
|
|
|
//--- a window that does not depend on the barrier.
|
|
|
|
|
if(idx - t <= excWindow)
|
|
|
|
|
{
|
|
|
|
|
if(hi > maxHigh)
|
|
|
|
|
maxHigh = hi;
|
|
|
|
|
if(lo < minLow)
|
|
|
|
|
minLow = lo;
|
|
|
|
|
}
|
feat: first-passage ladder + expectancy scan - price every geometry, not just the chosen one
Corrects the premise of the previous plan. Break-even is NOT a ceiling.
If the model shifts the win probability on the bars it selects from
p0 = m/(m+k) to p0 + d, then
EV = (p0+d)*k - (1-p0-d)*m = d*(k+m)
because p0*k - (1-p0)*m is zero by construction. The stop:target RATIO
is expectancy-neutral - a punishing break-even is exactly repaid by the
payoff - and only the real edge d and the TOTAL WIDTH (k+m) move EV.
Width matters because the spread is charged once per trade however wide
the barriers are, so a narrow barrier spends much of its own range on
costs. DeriveBarrierGeometry's own comment already said the ratio buys
nothing; the objective just never followed from it.
Blocker this had to solve first: m_excUpCache/m_excDownCache hold only
MAXIMUM travel each way, and a maximum cannot say which side was
reached FIRST - so any geometry other than the walked one was
undecidable on precisely the bars where both barriers were touched,
~28% of the sample.
- BARRIER_LADDER: per bar, the first-touch AGE for 8 travel distances
in each direction, filled during the walk the labels already run.
Cursors keep it O(1) amortised per walked bar rather than 16
comparisons. Levels are travel FROM ENTRY, not barrier prices, so one
ladder serves both directions and the spread is applied analytically
when a level converts back to an SL/TP multiple - storing prices
would need four ladders and bake today's spread into the cache.
Sized, invalidated and validity-gated with the label caches.
- ReportGeometryExpectancyScan: every ladder pair priced exactly off
that cache - width in ATR and in SPREADS (cost efficiency, knowable
without knowing d), break-even, both base rates, the share of bars
resolved inside the horizon, and EV per unit of edge. Compares the
widest resolvable pair against the quantile rule's pick.
MEASUREMENT ONLY - the quantile rule still chooses. Nothing here can
measure d, and width buys nothing if the wider target is less
predictable. Base rates are printed beside each break-even because a
persistent gap is DRIFT and must not be credited to the model.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-10 12:59:18 -04:00
|
|
|
//--- First-passage ladder. Runs over the WHOLE horizon, not excWindow: this measures how a trade
|
|
|
|
|
//--- held to its barriers would have resolved, so it must see every bar the trade would have been
|
|
|
|
|
//--- open for. Travel is measured from `entry` (the close) with no spread applied - see
|
|
|
|
|
//--- BARRIER_LADDER for why, and for how a level converts back into an SL/TP multiple.
|
|
|
|
|
int age = idx - t;
|
fix(labels): overlapping-label sample correction + horizon cap on the scale ladder
Three defects, all surfaced by the 2026-08-17 SP500 H4 run that shipped
stop 4.86 / target 9.71 (width 14.57*ATR, horizon 384).
1. EVERY STANDARD ERROR ASSUMED INDEPENDENT SAMPLES. Triple-barrier labels
started one per bar overlap by the label's lifespan, so n calls are worth
~n/L independent observations (Lopez de Prado, AFML ch. 4 - sample
uniqueness). All three sqrt(p(1-p)/n) sites divided by the RAW count.
The tell: the operating point's null-of-the-maximum gate is family-wise and
should fire on ~5% of eras under the null. Measured fire rates - PAI 47/73
(64%), ConvLSTM 9/24, LSTM 8/21 (38%), CONV 4/62 (6%). CONV, the only model
whose margin distribution admits few bins, sat on the null; the rest cleared
a bar that was too low by ~sqrt(L). PAI's deployed threshold consequently
alternated between the ENDS of its own range era to era (0.10 -> 0.88 ->
0.86 -> 0.66; coverage 16% <-> 73%).
TripleBarrierLabel now records when each label became KNOWABLE - the first
winning touch, or both stops, or the timeout - and the prebuild accumulates
the mean. EffectiveSampleSize() feeds the operating point, the member deploy
gate and the ensemble vote gate. Conservative by construction (n/L is an
upper bound on the damage); gates get harder, never easier.
2. THE SCALE LADDER RAN AWAY, again. Horizon scales as swingMedian*sl*tp, and
since 4d8cb08 reachability is measured OVER that horizon - so a wider rung
buys itself the time that makes it look reachable. Same target -> horizon ->
reach -> target loop the excursion window is kept short to avoid; fixing the
window confusion reopened it through the other door. It walked 128 -> 256 ->
384 bars and stopped at q90, the widest rung there is, with every rung
reading 39-48% against a 20% floor. A floor nothing fails selects nothing.
Rungs whose required horizon exceeds BARRIER_HORIZON_MAX are now rejected -
the same rule ReportGeometryExpectancyScan already applied. It was printing
the shipped pair as CLAMPED and disqualified ('h384!') two lines under the
deriver that chose it: two subsystems, one geometry, opposite verdicts.
3. THE RUNG SNAP DESTROYED THE RATIO IT WAS COMPARING. Both legs snapped
independently to the coarse first-passage grid, re-rating each candidate:
q90 4.86/9.71 -> 5.00/10.00 (2.00), q85 4.07/8.14 -> 5.00/10.00 (IDENTICAL
measurement), q75 3.07/6.13 -> 4.00/6.50 (1.63 - a nearer target). So the
ladder compared win shares taken at ratios from 1.63 to 2.17 and read the
differences as scale. It is why the reach column came out non-monotone in
width (q75 48.5% above q90 42.9%). The stop now snaps to its nearest rung in
log space and the target follows the ratio off it; the pair actually measured
is returned and logged, so a collision reads as a collision.
Also: LadderWinShare guarded against the conditional (fractal) geometry path,
which fills n from m_fracLegCount while leaving idxList empty - a latent
out-of-bounds on a currently-dead path.
New log lines: mean label lifespan and effective n on the label-cache line, the
required-vs-available horizon per rung, and the grid pair the reconciliation
actually measured (its tolerance now scales with the grid skew instead of a flat
5pp).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 01:12:05 -04:00
|
|
|
walkedAge = age;
|
feat: first-passage ladder + expectancy scan - price every geometry, not just the chosen one
Corrects the premise of the previous plan. Break-even is NOT a ceiling.
If the model shifts the win probability on the bars it selects from
p0 = m/(m+k) to p0 + d, then
EV = (p0+d)*k - (1-p0-d)*m = d*(k+m)
because p0*k - (1-p0)*m is zero by construction. The stop:target RATIO
is expectancy-neutral - a punishing break-even is exactly repaid by the
payoff - and only the real edge d and the TOTAL WIDTH (k+m) move EV.
Width matters because the spread is charged once per trade however wide
the barriers are, so a narrow barrier spends much of its own range on
costs. DeriveBarrierGeometry's own comment already said the ratio buys
nothing; the objective just never followed from it.
Blocker this had to solve first: m_excUpCache/m_excDownCache hold only
MAXIMUM travel each way, and a maximum cannot say which side was
reached FIRST - so any geometry other than the walked one was
undecidable on precisely the bars where both barriers were touched,
~28% of the sample.
- BARRIER_LADDER: per bar, the first-touch AGE for 8 travel distances
in each direction, filled during the walk the labels already run.
Cursors keep it O(1) amortised per walked bar rather than 16
comparisons. Levels are travel FROM ENTRY, not barrier prices, so one
ladder serves both directions and the spread is applied analytically
when a level converts back to an SL/TP multiple - storing prices
would need four ladders and bake today's spread into the cache.
Sized, invalidated and validity-gated with the label caches.
- ReportGeometryExpectancyScan: every ladder pair priced exactly off
that cache - width in ATR and in SPREADS (cost efficiency, knowable
without knowing d), break-even, both base rates, the share of bars
resolved inside the horizon, and EV per unit of edge. Compares the
widest resolvable pair against the quantile rule's pick.
MEASUREMENT ONLY - the quantile rule still chooses. Nothing here can
measure d, and width buys nothing if the wider target is less
predictable. Base rates are printed beside each break-even because a
persistent gap is DRIFT and must not be credited to the model.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-10 12:59:18 -04:00
|
|
|
while(upCursor < BARRIER_LADDER_COUNT && hi >= entry + BARRIER_LADDER[upCursor] * atr)
|
|
|
|
|
{
|
|
|
|
|
m_lastLadderUpAt[upCursor] = age;
|
|
|
|
|
upCursor++;
|
|
|
|
|
}
|
|
|
|
|
while(dnCursor < BARRIER_LADDER_COUNT && lo <= entry - BARRIER_LADDER[dnCursor] * atr)
|
|
|
|
|
{
|
|
|
|
|
m_lastLadderDownAt[dnCursor] = age;
|
|
|
|
|
dnCursor++;
|
|
|
|
|
}
|
feat(ai): triple-barrier labels replace exact-pivot ZigZag targets
The 31:1 class imbalance was self-inflicted by the TARGET, not a property
of the market. Labelling only the exact bar where a ZigZag pivot confirms
gave Buy 1164 / Sell 1164 / Neutral 35841, and every correction mechanism
this codebase accumulated sits downstream of that one choice: the
logit-adjusted loss and its range cap, the prior EMA, the +-3.0 output-bias
seed, balanced-accuracy-then-precision selection with its coverage floor,
the recall floor and its catch-22, the alternation gate, NMS, and the four
oversampling designs that collapsed before them.
The reference this engine is built on (references/neuronetworksbook.pdf
ch. 3.1/3.3) also uses ZigZag, but targets the DIRECTION TO THE NEXT
EXTREMUM on every bar - ~50/50 by construction, with no imbalance to
correct at all. It never had this problem because it never asked "is this
the pivot bar".
Labels are now the triple barrier (Lopez de Prado ch. 3), using the EA's
OWN SL_Mode/TP_Mode: does a trade opened at this bar's close reach its
target before its stop, within a horizon. Buy = long resolves, Sell =
short resolves, Neutral = neither. Consequences:
- dir-precision in the era line stops being a proxy and becomes the win
rate of the strategy under its own exit rules.
- Expected balance ~25/25/50 at the shipped 1:3 (gambler's ruin), i.e.
~2:1 instead of 31:1. Measured and logged at the end of the prebuild.
- Spread is charged on both legs, so it is a NET win rate.
- Intrabar ambiguity resolves to the STOP. OHLC cannot order two touches
inside one bar and the optimistic reading is how a backtested edge
becomes a live loss.
ZigZag stays as input features (EnableSwingContext) and now also supplies
the vertical barrier: the horizon is the median confirmed leg length,
snapped to a coarse ladder. Derived, not configured, and deliberately kept
out of the filename fingerprint - a filename keyed on a measured quantity
orphans a trained model the moment the measurement moves.
Removed, because the premise died with the old target:
- the alternation gate. Correct for pivot labels (a ZigZag cannot emit two
same-type pivots in a row, so a repeat was provably a false fire), and
wrong for barrier labels, which answer each bar independently. It also
took its worst consequence with it: a one-sided model previously got ONE
trade per backtest, a hard blocker on marketplace validation.
- SignalClusterWindow now defaults off - it de-duplicated repeats that are
now real trades. Kept as an opt-in display control.
- LABEL_WINDOW_BARS, the pivot-widening pass, ConfirmedZigZagLabel.
- the era-0 output-bias seed now needs a genuinely dominant class (0.70)
rather than 0.40; at ~50% Neutral a +-3.0 seed is a distortion, not a
correction.
Also fixed, both found while wiring the above:
1. RefreshConvergedSignal sized its buffers from a date delta
(Bars(sym, period, dtStudied, TimeCurrent())). dtStudied is a training
watermark; in the tester it is loaded from a live-chart save AHEAD of
the simulated date, so the interval inverted, Bars() returned ~0, and
the buffer came out at exactly m_historyBars - deep enough for the OHLC
window and far too shallow for the Donchian-50 / 20-bar-return / SMA
extension behind it. Inference silently computed DIFFERENT features
from the ones training learned on, live as well as in the tester. Now
sized from what the feature builder actually needs.
2. The barrier horizon is resolved on the deployed path too. A deployed
model never enters Train(), so it never reached the prebuild, and
OnlineLearnStep reads the horizon as its confirmation delay - left at
the fallback it would have backpropped bars whose barriers had not
resolved. Silent lookahead in the one place that writes to a live model.
SL_Mode/TP_Mode join the weights fingerprint: they define the labels now,
so a model trained at 1:3 must never be silently reused at 1:1. This
re-keys every pre-existing model by design - none were trained on this task.
Inference census extended with the vote gate. LongCondition/ShortCondition
open with a readiness check the refresh counters never see; in the tester it
reduces to "the seeded _optcache.nnw must have LOADED", and if it did not,
every vote is hard-zeroed while the model still answers Buy. The old three
counters would have read that as "the model says Neutral" - false, and a
completely different fix. This is the leading candidate for the
zero-direction backtest and the census can now name it in one run.
Both builds compile 0 errors / 0 warnings. Forces a full retrain.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 20:39:49 -04:00
|
|
|
//--- Stop tested FIRST on each side, so a bar that spans both barriers is scored as the loss.
|
|
|
|
|
if(!longWon && !longLost)
|
|
|
|
|
{
|
|
|
|
|
if(lo <= longSl)
|
fix(labels): overlapping-label sample correction + horizon cap on the scale ladder
Three defects, all surfaced by the 2026-08-17 SP500 H4 run that shipped
stop 4.86 / target 9.71 (width 14.57*ATR, horizon 384).
1. EVERY STANDARD ERROR ASSUMED INDEPENDENT SAMPLES. Triple-barrier labels
started one per bar overlap by the label's lifespan, so n calls are worth
~n/L independent observations (Lopez de Prado, AFML ch. 4 - sample
uniqueness). All three sqrt(p(1-p)/n) sites divided by the RAW count.
The tell: the operating point's null-of-the-maximum gate is family-wise and
should fire on ~5% of eras under the null. Measured fire rates - PAI 47/73
(64%), ConvLSTM 9/24, LSTM 8/21 (38%), CONV 4/62 (6%). CONV, the only model
whose margin distribution admits few bins, sat on the null; the rest cleared
a bar that was too low by ~sqrt(L). PAI's deployed threshold consequently
alternated between the ENDS of its own range era to era (0.10 -> 0.88 ->
0.86 -> 0.66; coverage 16% <-> 73%).
TripleBarrierLabel now records when each label became KNOWABLE - the first
winning touch, or both stops, or the timeout - and the prebuild accumulates
the mean. EffectiveSampleSize() feeds the operating point, the member deploy
gate and the ensemble vote gate. Conservative by construction (n/L is an
upper bound on the damage); gates get harder, never easier.
2. THE SCALE LADDER RAN AWAY, again. Horizon scales as swingMedian*sl*tp, and
since 4d8cb08 reachability is measured OVER that horizon - so a wider rung
buys itself the time that makes it look reachable. Same target -> horizon ->
reach -> target loop the excursion window is kept short to avoid; fixing the
window confusion reopened it through the other door. It walked 128 -> 256 ->
384 bars and stopped at q90, the widest rung there is, with every rung
reading 39-48% against a 20% floor. A floor nothing fails selects nothing.
Rungs whose required horizon exceeds BARRIER_HORIZON_MAX are now rejected -
the same rule ReportGeometryExpectancyScan already applied. It was printing
the shipped pair as CLAMPED and disqualified ('h384!') two lines under the
deriver that chose it: two subsystems, one geometry, opposite verdicts.
3. THE RUNG SNAP DESTROYED THE RATIO IT WAS COMPARING. Both legs snapped
independently to the coarse first-passage grid, re-rating each candidate:
q90 4.86/9.71 -> 5.00/10.00 (2.00), q85 4.07/8.14 -> 5.00/10.00 (IDENTICAL
measurement), q75 3.07/6.13 -> 4.00/6.50 (1.63 - a nearer target). So the
ladder compared win shares taken at ratios from 1.63 to 2.17 and read the
differences as scale. It is why the reach column came out non-monotone in
width (q75 48.5% above q90 42.9%). The stop now snaps to its nearest rung in
log space and the target follows the ratio off it; the pair actually measured
is returned and logged, so a collision reads as a collision.
Also: LadderWinShare guarded against the conditional (fractal) geometry path,
which fills n from m_fracLegCount while leaving idxList empty - a latent
out-of-bounds on a currently-dead path.
New log lines: mean label lifespan and effective n on the label-cache line, the
required-vs-available horizon per rung, and the grid pair the reconciliation
actually measured (its tolerance now scales with the grid skew instead of a flat
5pp).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 01:12:05 -04:00
|
|
|
{
|
feat(ai): triple-barrier labels replace exact-pivot ZigZag targets
The 31:1 class imbalance was self-inflicted by the TARGET, not a property
of the market. Labelling only the exact bar where a ZigZag pivot confirms
gave Buy 1164 / Sell 1164 / Neutral 35841, and every correction mechanism
this codebase accumulated sits downstream of that one choice: the
logit-adjusted loss and its range cap, the prior EMA, the +-3.0 output-bias
seed, balanced-accuracy-then-precision selection with its coverage floor,
the recall floor and its catch-22, the alternation gate, NMS, and the four
oversampling designs that collapsed before them.
The reference this engine is built on (references/neuronetworksbook.pdf
ch. 3.1/3.3) also uses ZigZag, but targets the DIRECTION TO THE NEXT
EXTREMUM on every bar - ~50/50 by construction, with no imbalance to
correct at all. It never had this problem because it never asked "is this
the pivot bar".
Labels are now the triple barrier (Lopez de Prado ch. 3), using the EA's
OWN SL_Mode/TP_Mode: does a trade opened at this bar's close reach its
target before its stop, within a horizon. Buy = long resolves, Sell =
short resolves, Neutral = neither. Consequences:
- dir-precision in the era line stops being a proxy and becomes the win
rate of the strategy under its own exit rules.
- Expected balance ~25/25/50 at the shipped 1:3 (gambler's ruin), i.e.
~2:1 instead of 31:1. Measured and logged at the end of the prebuild.
- Spread is charged on both legs, so it is a NET win rate.
- Intrabar ambiguity resolves to the STOP. OHLC cannot order two touches
inside one bar and the optimistic reading is how a backtested edge
becomes a live loss.
ZigZag stays as input features (EnableSwingContext) and now also supplies
the vertical barrier: the horizon is the median confirmed leg length,
snapped to a coarse ladder. Derived, not configured, and deliberately kept
out of the filename fingerprint - a filename keyed on a measured quantity
orphans a trained model the moment the measurement moves.
Removed, because the premise died with the old target:
- the alternation gate. Correct for pivot labels (a ZigZag cannot emit two
same-type pivots in a row, so a repeat was provably a false fire), and
wrong for barrier labels, which answer each bar independently. It also
took its worst consequence with it: a one-sided model previously got ONE
trade per backtest, a hard blocker on marketplace validation.
- SignalClusterWindow now defaults off - it de-duplicated repeats that are
now real trades. Kept as an opt-in display control.
- LABEL_WINDOW_BARS, the pivot-widening pass, ConfirmedZigZagLabel.
- the era-0 output-bias seed now needs a genuinely dominant class (0.70)
rather than 0.40; at ~50% Neutral a +-3.0 seed is a distortion, not a
correction.
Also fixed, both found while wiring the above:
1. RefreshConvergedSignal sized its buffers from a date delta
(Bars(sym, period, dtStudied, TimeCurrent())). dtStudied is a training
watermark; in the tester it is loaded from a live-chart save AHEAD of
the simulated date, so the interval inverted, Bars() returned ~0, and
the buffer came out at exactly m_historyBars - deep enough for the OHLC
window and far too shallow for the Donchian-50 / 20-bar-return / SMA
extension behind it. Inference silently computed DIFFERENT features
from the ones training learned on, live as well as in the tester. Now
sized from what the feature builder actually needs.
2. The barrier horizon is resolved on the deployed path too. A deployed
model never enters Train(), so it never reached the prebuild, and
OnlineLearnStep reads the horizon as its confirmation delay - left at
the fallback it would have backpropped bars whose barriers had not
resolved. Silent lookahead in the one place that writes to a live model.
SL_Mode/TP_Mode join the weights fingerprint: they define the labels now,
so a model trained at 1:3 must never be silently reused at 1:1. This
re-keys every pre-existing model by design - none were trained on this task.
Inference census extended with the vote gate. LongCondition/ShortCondition
open with a readiness check the refresh counters never see; in the tester it
reduces to "the seeded _optcache.nnw must have LOADED", and if it did not,
every vote is hard-zeroed while the model still answers Buy. The old three
counters would have read that as "the model says Neutral" - false, and a
completely different fix. This is the leading candidate for the
zero-direction backtest and the census can now name it in one run.
Both builds compile 0 errors / 0 warnings. Forces a full retrain.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 20:39:49 -04:00
|
|
|
longLost = true;
|
fix(labels): overlapping-label sample correction + horizon cap on the scale ladder
Three defects, all surfaced by the 2026-08-17 SP500 H4 run that shipped
stop 4.86 / target 9.71 (width 14.57*ATR, horizon 384).
1. EVERY STANDARD ERROR ASSUMED INDEPENDENT SAMPLES. Triple-barrier labels
started one per bar overlap by the label's lifespan, so n calls are worth
~n/L independent observations (Lopez de Prado, AFML ch. 4 - sample
uniqueness). All three sqrt(p(1-p)/n) sites divided by the RAW count.
The tell: the operating point's null-of-the-maximum gate is family-wise and
should fire on ~5% of eras under the null. Measured fire rates - PAI 47/73
(64%), ConvLSTM 9/24, LSTM 8/21 (38%), CONV 4/62 (6%). CONV, the only model
whose margin distribution admits few bins, sat on the null; the rest cleared
a bar that was too low by ~sqrt(L). PAI's deployed threshold consequently
alternated between the ENDS of its own range era to era (0.10 -> 0.88 ->
0.86 -> 0.66; coverage 16% <-> 73%).
TripleBarrierLabel now records when each label became KNOWABLE - the first
winning touch, or both stops, or the timeout - and the prebuild accumulates
the mean. EffectiveSampleSize() feeds the operating point, the member deploy
gate and the ensemble vote gate. Conservative by construction (n/L is an
upper bound on the damage); gates get harder, never easier.
2. THE SCALE LADDER RAN AWAY, again. Horizon scales as swingMedian*sl*tp, and
since 4d8cb08 reachability is measured OVER that horizon - so a wider rung
buys itself the time that makes it look reachable. Same target -> horizon ->
reach -> target loop the excursion window is kept short to avoid; fixing the
window confusion reopened it through the other door. It walked 128 -> 256 ->
384 bars and stopped at q90, the widest rung there is, with every rung
reading 39-48% against a 20% floor. A floor nothing fails selects nothing.
Rungs whose required horizon exceeds BARRIER_HORIZON_MAX are now rejected -
the same rule ReportGeometryExpectancyScan already applied. It was printing
the shipped pair as CLAMPED and disqualified ('h384!') two lines under the
deriver that chose it: two subsystems, one geometry, opposite verdicts.
3. THE RUNG SNAP DESTROYED THE RATIO IT WAS COMPARING. Both legs snapped
independently to the coarse first-passage grid, re-rating each candidate:
q90 4.86/9.71 -> 5.00/10.00 (2.00), q85 4.07/8.14 -> 5.00/10.00 (IDENTICAL
measurement), q75 3.07/6.13 -> 4.00/6.50 (1.63 - a nearer target). So the
ladder compared win shares taken at ratios from 1.63 to 2.17 and read the
differences as scale. It is why the reach column came out non-monotone in
width (q75 48.5% above q90 42.9%). The stop now snaps to its nearest rung in
log space and the target follows the ratio off it; the pair actually measured
is returned and logged, so a collision reads as a collision.
Also: LadderWinShare guarded against the conditional (fractal) geometry path,
which fills n from m_fracLegCount while leaving idxList empty - a latent
out-of-bounds on a currently-dead path.
New log lines: mean label lifespan and effective n on the label-cache line, the
required-vs-available horizon per rung, and the grid pair the reconciliation
actually measured (its tolerance now scales with the grid skew instead of a flat
5pp).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 01:12:05 -04:00
|
|
|
longEndAge = age;
|
|
|
|
|
}
|
fix: both-won bars were labelled "do not trade" - resolve by first touch
Removing the min-reward:risk raise let the MEASURED geometry come back with
the target NEARER than the stop (SP500 H1: target 1.62*ATR at q50 of
favourable, stop 3.33*ATR at q75 of adverse). That reopened a branch the
code called unreachable: price can reach +target and -target inside one
horizon, winning in BOTH directions, and those bars fell through to Neutral.
Neutral has only three producers, both-lost is unreachable (you cannot touch
-3.33 without crossing -1.62 first, which wins the short), and timeouts logged
at 1.0% of Neutral - so ~27% of ALL bars were being handed to the model as the
abstain class when a trade either way would have collected its target. The
cleanest positives in the sample, labelled "do not trade", while the fitted
confidence threshold was being asked to find selectivity in what was left.
Resolved by FIRST TOUCH: the target reached earlier is the trade that would
have closed first. Same forward window, no extra lookahead. Same-bar ties stay
Neutral - OHLC cannot order two touches, and unlike an intrabar stop tie there
is no pessimistic side to fall to, so a guess would inject a coin-flip
direction into the target.
Also:
- count both-won and its same-bar tie subset in the prebuild line, so the
share is measured rather than inferred from arithmetic on a log line
- scope the timeout counter to IS, matching the tally it is reported as a
percentage OF; it was incremented over the whole scan and divided by an
in-sample denominator
- clear m_lastBarrierTimedOut at the top of the walk with the excursions, not
at the bottom - the two early returns published the previous bar's verdict
- mark the pass-1 label line PROVISIONAL. It prints the enum fallback because
geometry can only be derived from excursions that do not exist yet, and it
reads exactly like a config change that failed to take effect
FORCES RETRAIN. Both build variants compile 0 errors, 0 warnings.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 15:35:42 -04:00
|
|
|
else
|
|
|
|
|
if(hi >= longTp)
|
|
|
|
|
{
|
|
|
|
|
longWon = true;
|
|
|
|
|
longWonAt = t;
|
fix(labels): overlapping-label sample correction + horizon cap on the scale ladder
Three defects, all surfaced by the 2026-08-17 SP500 H4 run that shipped
stop 4.86 / target 9.71 (width 14.57*ATR, horizon 384).
1. EVERY STANDARD ERROR ASSUMED INDEPENDENT SAMPLES. Triple-barrier labels
started one per bar overlap by the label's lifespan, so n calls are worth
~n/L independent observations (Lopez de Prado, AFML ch. 4 - sample
uniqueness). All three sqrt(p(1-p)/n) sites divided by the RAW count.
The tell: the operating point's null-of-the-maximum gate is family-wise and
should fire on ~5% of eras under the null. Measured fire rates - PAI 47/73
(64%), ConvLSTM 9/24, LSTM 8/21 (38%), CONV 4/62 (6%). CONV, the only model
whose margin distribution admits few bins, sat on the null; the rest cleared
a bar that was too low by ~sqrt(L). PAI's deployed threshold consequently
alternated between the ENDS of its own range era to era (0.10 -> 0.88 ->
0.86 -> 0.66; coverage 16% <-> 73%).
TripleBarrierLabel now records when each label became KNOWABLE - the first
winning touch, or both stops, or the timeout - and the prebuild accumulates
the mean. EffectiveSampleSize() feeds the operating point, the member deploy
gate and the ensemble vote gate. Conservative by construction (n/L is an
upper bound on the damage); gates get harder, never easier.
2. THE SCALE LADDER RAN AWAY, again. Horizon scales as swingMedian*sl*tp, and
since 4d8cb08 reachability is measured OVER that horizon - so a wider rung
buys itself the time that makes it look reachable. Same target -> horizon ->
reach -> target loop the excursion window is kept short to avoid; fixing the
window confusion reopened it through the other door. It walked 128 -> 256 ->
384 bars and stopped at q90, the widest rung there is, with every rung
reading 39-48% against a 20% floor. A floor nothing fails selects nothing.
Rungs whose required horizon exceeds BARRIER_HORIZON_MAX are now rejected -
the same rule ReportGeometryExpectancyScan already applied. It was printing
the shipped pair as CLAMPED and disqualified ('h384!') two lines under the
deriver that chose it: two subsystems, one geometry, opposite verdicts.
3. THE RUNG SNAP DESTROYED THE RATIO IT WAS COMPARING. Both legs snapped
independently to the coarse first-passage grid, re-rating each candidate:
q90 4.86/9.71 -> 5.00/10.00 (2.00), q85 4.07/8.14 -> 5.00/10.00 (IDENTICAL
measurement), q75 3.07/6.13 -> 4.00/6.50 (1.63 - a nearer target). So the
ladder compared win shares taken at ratios from 1.63 to 2.17 and read the
differences as scale. It is why the reach column came out non-monotone in
width (q75 48.5% above q90 42.9%). The stop now snaps to its nearest rung in
log space and the target follows the ratio off it; the pair actually measured
is returned and logged, so a collision reads as a collision.
Also: LadderWinShare guarded against the conditional (fractal) geometry path,
which fills n from m_fracLegCount while leaving idxList empty - a latent
out-of-bounds on a currently-dead path.
New log lines: mean label lifespan and effective n on the label-cache line, the
required-vs-available horizon per rung, and the grid pair the reconciliation
actually measured (its tolerance now scales with the grid skew instead of a flat
5pp).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 01:12:05 -04:00
|
|
|
longEndAge = age;
|
fix: both-won bars were labelled "do not trade" - resolve by first touch
Removing the min-reward:risk raise let the MEASURED geometry come back with
the target NEARER than the stop (SP500 H1: target 1.62*ATR at q50 of
favourable, stop 3.33*ATR at q75 of adverse). That reopened a branch the
code called unreachable: price can reach +target and -target inside one
horizon, winning in BOTH directions, and those bars fell through to Neutral.
Neutral has only three producers, both-lost is unreachable (you cannot touch
-3.33 without crossing -1.62 first, which wins the short), and timeouts logged
at 1.0% of Neutral - so ~27% of ALL bars were being handed to the model as the
abstain class when a trade either way would have collected its target. The
cleanest positives in the sample, labelled "do not trade", while the fitted
confidence threshold was being asked to find selectivity in what was left.
Resolved by FIRST TOUCH: the target reached earlier is the trade that would
have closed first. Same forward window, no extra lookahead. Same-bar ties stay
Neutral - OHLC cannot order two touches, and unlike an intrabar stop tie there
is no pessimistic side to fall to, so a guess would inject a coin-flip
direction into the target.
Also:
- count both-won and its same-bar tie subset in the prebuild line, so the
share is measured rather than inferred from arithmetic on a log line
- scope the timeout counter to IS, matching the tally it is reported as a
percentage OF; it was incremented over the whole scan and divided by an
in-sample denominator
- clear m_lastBarrierTimedOut at the top of the walk with the excursions, not
at the bottom - the two early returns published the previous bar's verdict
- mark the pass-1 label line PROVISIONAL. It prints the enum fallback because
geometry can only be derived from excursions that do not exist yet, and it
reads exactly like a config change that failed to take effect
FORCES RETRAIN. Both build variants compile 0 errors, 0 warnings.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 15:35:42 -04:00
|
|
|
}
|
feat(ai): triple-barrier labels replace exact-pivot ZigZag targets
The 31:1 class imbalance was self-inflicted by the TARGET, not a property
of the market. Labelling only the exact bar where a ZigZag pivot confirms
gave Buy 1164 / Sell 1164 / Neutral 35841, and every correction mechanism
this codebase accumulated sits downstream of that one choice: the
logit-adjusted loss and its range cap, the prior EMA, the +-3.0 output-bias
seed, balanced-accuracy-then-precision selection with its coverage floor,
the recall floor and its catch-22, the alternation gate, NMS, and the four
oversampling designs that collapsed before them.
The reference this engine is built on (references/neuronetworksbook.pdf
ch. 3.1/3.3) also uses ZigZag, but targets the DIRECTION TO THE NEXT
EXTREMUM on every bar - ~50/50 by construction, with no imbalance to
correct at all. It never had this problem because it never asked "is this
the pivot bar".
Labels are now the triple barrier (Lopez de Prado ch. 3), using the EA's
OWN SL_Mode/TP_Mode: does a trade opened at this bar's close reach its
target before its stop, within a horizon. Buy = long resolves, Sell =
short resolves, Neutral = neither. Consequences:
- dir-precision in the era line stops being a proxy and becomes the win
rate of the strategy under its own exit rules.
- Expected balance ~25/25/50 at the shipped 1:3 (gambler's ruin), i.e.
~2:1 instead of 31:1. Measured and logged at the end of the prebuild.
- Spread is charged on both legs, so it is a NET win rate.
- Intrabar ambiguity resolves to the STOP. OHLC cannot order two touches
inside one bar and the optimistic reading is how a backtested edge
becomes a live loss.
ZigZag stays as input features (EnableSwingContext) and now also supplies
the vertical barrier: the horizon is the median confirmed leg length,
snapped to a coarse ladder. Derived, not configured, and deliberately kept
out of the filename fingerprint - a filename keyed on a measured quantity
orphans a trained model the moment the measurement moves.
Removed, because the premise died with the old target:
- the alternation gate. Correct for pivot labels (a ZigZag cannot emit two
same-type pivots in a row, so a repeat was provably a false fire), and
wrong for barrier labels, which answer each bar independently. It also
took its worst consequence with it: a one-sided model previously got ONE
trade per backtest, a hard blocker on marketplace validation.
- SignalClusterWindow now defaults off - it de-duplicated repeats that are
now real trades. Kept as an opt-in display control.
- LABEL_WINDOW_BARS, the pivot-widening pass, ConfirmedZigZagLabel.
- the era-0 output-bias seed now needs a genuinely dominant class (0.70)
rather than 0.40; at ~50% Neutral a +-3.0 seed is a distortion, not a
correction.
Also fixed, both found while wiring the above:
1. RefreshConvergedSignal sized its buffers from a date delta
(Bars(sym, period, dtStudied, TimeCurrent())). dtStudied is a training
watermark; in the tester it is loaded from a live-chart save AHEAD of
the simulated date, so the interval inverted, Bars() returned ~0, and
the buffer came out at exactly m_historyBars - deep enough for the OHLC
window and far too shallow for the Donchian-50 / 20-bar-return / SMA
extension behind it. Inference silently computed DIFFERENT features
from the ones training learned on, live as well as in the tester. Now
sized from what the feature builder actually needs.
2. The barrier horizon is resolved on the deployed path too. A deployed
model never enters Train(), so it never reached the prebuild, and
OnlineLearnStep reads the horizon as its confirmation delay - left at
the fallback it would have backpropped bars whose barriers had not
resolved. Silent lookahead in the one place that writes to a live model.
SL_Mode/TP_Mode join the weights fingerprint: they define the labels now,
so a model trained at 1:3 must never be silently reused at 1:1. This
re-keys every pre-existing model by design - none were trained on this task.
Inference census extended with the vote gate. LongCondition/ShortCondition
open with a readiness check the refresh counters never see; in the tester it
reduces to "the seeded _optcache.nnw must have LOADED", and if it did not,
every vote is hard-zeroed while the model still answers Buy. The old three
counters would have read that as "the model says Neutral" - false, and a
completely different fix. This is the leading candidate for the
zero-direction backtest and the census can now name it in one run.
Both builds compile 0 errors / 0 warnings. Forces a full retrain.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 20:39:49 -04:00
|
|
|
}
|
|
|
|
|
if(!shortWon && !shortLost)
|
|
|
|
|
{
|
|
|
|
|
if(hi >= shortSl)
|
fix(labels): overlapping-label sample correction + horizon cap on the scale ladder
Three defects, all surfaced by the 2026-08-17 SP500 H4 run that shipped
stop 4.86 / target 9.71 (width 14.57*ATR, horizon 384).
1. EVERY STANDARD ERROR ASSUMED INDEPENDENT SAMPLES. Triple-barrier labels
started one per bar overlap by the label's lifespan, so n calls are worth
~n/L independent observations (Lopez de Prado, AFML ch. 4 - sample
uniqueness). All three sqrt(p(1-p)/n) sites divided by the RAW count.
The tell: the operating point's null-of-the-maximum gate is family-wise and
should fire on ~5% of eras under the null. Measured fire rates - PAI 47/73
(64%), ConvLSTM 9/24, LSTM 8/21 (38%), CONV 4/62 (6%). CONV, the only model
whose margin distribution admits few bins, sat on the null; the rest cleared
a bar that was too low by ~sqrt(L). PAI's deployed threshold consequently
alternated between the ENDS of its own range era to era (0.10 -> 0.88 ->
0.86 -> 0.66; coverage 16% <-> 73%).
TripleBarrierLabel now records when each label became KNOWABLE - the first
winning touch, or both stops, or the timeout - and the prebuild accumulates
the mean. EffectiveSampleSize() feeds the operating point, the member deploy
gate and the ensemble vote gate. Conservative by construction (n/L is an
upper bound on the damage); gates get harder, never easier.
2. THE SCALE LADDER RAN AWAY, again. Horizon scales as swingMedian*sl*tp, and
since 4d8cb08 reachability is measured OVER that horizon - so a wider rung
buys itself the time that makes it look reachable. Same target -> horizon ->
reach -> target loop the excursion window is kept short to avoid; fixing the
window confusion reopened it through the other door. It walked 128 -> 256 ->
384 bars and stopped at q90, the widest rung there is, with every rung
reading 39-48% against a 20% floor. A floor nothing fails selects nothing.
Rungs whose required horizon exceeds BARRIER_HORIZON_MAX are now rejected -
the same rule ReportGeometryExpectancyScan already applied. It was printing
the shipped pair as CLAMPED and disqualified ('h384!') two lines under the
deriver that chose it: two subsystems, one geometry, opposite verdicts.
3. THE RUNG SNAP DESTROYED THE RATIO IT WAS COMPARING. Both legs snapped
independently to the coarse first-passage grid, re-rating each candidate:
q90 4.86/9.71 -> 5.00/10.00 (2.00), q85 4.07/8.14 -> 5.00/10.00 (IDENTICAL
measurement), q75 3.07/6.13 -> 4.00/6.50 (1.63 - a nearer target). So the
ladder compared win shares taken at ratios from 1.63 to 2.17 and read the
differences as scale. It is why the reach column came out non-monotone in
width (q75 48.5% above q90 42.9%). The stop now snaps to its nearest rung in
log space and the target follows the ratio off it; the pair actually measured
is returned and logged, so a collision reads as a collision.
Also: LadderWinShare guarded against the conditional (fractal) geometry path,
which fills n from m_fracLegCount while leaving idxList empty - a latent
out-of-bounds on a currently-dead path.
New log lines: mean label lifespan and effective n on the label-cache line, the
required-vs-available horizon per rung, and the grid pair the reconciliation
actually measured (its tolerance now scales with the grid skew instead of a flat
5pp).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 01:12:05 -04:00
|
|
|
{
|
feat(ai): triple-barrier labels replace exact-pivot ZigZag targets
The 31:1 class imbalance was self-inflicted by the TARGET, not a property
of the market. Labelling only the exact bar where a ZigZag pivot confirms
gave Buy 1164 / Sell 1164 / Neutral 35841, and every correction mechanism
this codebase accumulated sits downstream of that one choice: the
logit-adjusted loss and its range cap, the prior EMA, the +-3.0 output-bias
seed, balanced-accuracy-then-precision selection with its coverage floor,
the recall floor and its catch-22, the alternation gate, NMS, and the four
oversampling designs that collapsed before them.
The reference this engine is built on (references/neuronetworksbook.pdf
ch. 3.1/3.3) also uses ZigZag, but targets the DIRECTION TO THE NEXT
EXTREMUM on every bar - ~50/50 by construction, with no imbalance to
correct at all. It never had this problem because it never asked "is this
the pivot bar".
Labels are now the triple barrier (Lopez de Prado ch. 3), using the EA's
OWN SL_Mode/TP_Mode: does a trade opened at this bar's close reach its
target before its stop, within a horizon. Buy = long resolves, Sell =
short resolves, Neutral = neither. Consequences:
- dir-precision in the era line stops being a proxy and becomes the win
rate of the strategy under its own exit rules.
- Expected balance ~25/25/50 at the shipped 1:3 (gambler's ruin), i.e.
~2:1 instead of 31:1. Measured and logged at the end of the prebuild.
- Spread is charged on both legs, so it is a NET win rate.
- Intrabar ambiguity resolves to the STOP. OHLC cannot order two touches
inside one bar and the optimistic reading is how a backtested edge
becomes a live loss.
ZigZag stays as input features (EnableSwingContext) and now also supplies
the vertical barrier: the horizon is the median confirmed leg length,
snapped to a coarse ladder. Derived, not configured, and deliberately kept
out of the filename fingerprint - a filename keyed on a measured quantity
orphans a trained model the moment the measurement moves.
Removed, because the premise died with the old target:
- the alternation gate. Correct for pivot labels (a ZigZag cannot emit two
same-type pivots in a row, so a repeat was provably a false fire), and
wrong for barrier labels, which answer each bar independently. It also
took its worst consequence with it: a one-sided model previously got ONE
trade per backtest, a hard blocker on marketplace validation.
- SignalClusterWindow now defaults off - it de-duplicated repeats that are
now real trades. Kept as an opt-in display control.
- LABEL_WINDOW_BARS, the pivot-widening pass, ConfirmedZigZagLabel.
- the era-0 output-bias seed now needs a genuinely dominant class (0.70)
rather than 0.40; at ~50% Neutral a +-3.0 seed is a distortion, not a
correction.
Also fixed, both found while wiring the above:
1. RefreshConvergedSignal sized its buffers from a date delta
(Bars(sym, period, dtStudied, TimeCurrent())). dtStudied is a training
watermark; in the tester it is loaded from a live-chart save AHEAD of
the simulated date, so the interval inverted, Bars() returned ~0, and
the buffer came out at exactly m_historyBars - deep enough for the OHLC
window and far too shallow for the Donchian-50 / 20-bar-return / SMA
extension behind it. Inference silently computed DIFFERENT features
from the ones training learned on, live as well as in the tester. Now
sized from what the feature builder actually needs.
2. The barrier horizon is resolved on the deployed path too. A deployed
model never enters Train(), so it never reached the prebuild, and
OnlineLearnStep reads the horizon as its confirmation delay - left at
the fallback it would have backpropped bars whose barriers had not
resolved. Silent lookahead in the one place that writes to a live model.
SL_Mode/TP_Mode join the weights fingerprint: they define the labels now,
so a model trained at 1:3 must never be silently reused at 1:1. This
re-keys every pre-existing model by design - none were trained on this task.
Inference census extended with the vote gate. LongCondition/ShortCondition
open with a readiness check the refresh counters never see; in the tester it
reduces to "the seeded _optcache.nnw must have LOADED", and if it did not,
every vote is hard-zeroed while the model still answers Buy. The old three
counters would have read that as "the model says Neutral" - false, and a
completely different fix. This is the leading candidate for the
zero-direction backtest and the census can now name it in one run.
Both builds compile 0 errors / 0 warnings. Forces a full retrain.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 20:39:49 -04:00
|
|
|
shortLost = true;
|
fix(labels): overlapping-label sample correction + horizon cap on the scale ladder
Three defects, all surfaced by the 2026-08-17 SP500 H4 run that shipped
stop 4.86 / target 9.71 (width 14.57*ATR, horizon 384).
1. EVERY STANDARD ERROR ASSUMED INDEPENDENT SAMPLES. Triple-barrier labels
started one per bar overlap by the label's lifespan, so n calls are worth
~n/L independent observations (Lopez de Prado, AFML ch. 4 - sample
uniqueness). All three sqrt(p(1-p)/n) sites divided by the RAW count.
The tell: the operating point's null-of-the-maximum gate is family-wise and
should fire on ~5% of eras under the null. Measured fire rates - PAI 47/73
(64%), ConvLSTM 9/24, LSTM 8/21 (38%), CONV 4/62 (6%). CONV, the only model
whose margin distribution admits few bins, sat on the null; the rest cleared
a bar that was too low by ~sqrt(L). PAI's deployed threshold consequently
alternated between the ENDS of its own range era to era (0.10 -> 0.88 ->
0.86 -> 0.66; coverage 16% <-> 73%).
TripleBarrierLabel now records when each label became KNOWABLE - the first
winning touch, or both stops, or the timeout - and the prebuild accumulates
the mean. EffectiveSampleSize() feeds the operating point, the member deploy
gate and the ensemble vote gate. Conservative by construction (n/L is an
upper bound on the damage); gates get harder, never easier.
2. THE SCALE LADDER RAN AWAY, again. Horizon scales as swingMedian*sl*tp, and
since 4d8cb08 reachability is measured OVER that horizon - so a wider rung
buys itself the time that makes it look reachable. Same target -> horizon ->
reach -> target loop the excursion window is kept short to avoid; fixing the
window confusion reopened it through the other door. It walked 128 -> 256 ->
384 bars and stopped at q90, the widest rung there is, with every rung
reading 39-48% against a 20% floor. A floor nothing fails selects nothing.
Rungs whose required horizon exceeds BARRIER_HORIZON_MAX are now rejected -
the same rule ReportGeometryExpectancyScan already applied. It was printing
the shipped pair as CLAMPED and disqualified ('h384!') two lines under the
deriver that chose it: two subsystems, one geometry, opposite verdicts.
3. THE RUNG SNAP DESTROYED THE RATIO IT WAS COMPARING. Both legs snapped
independently to the coarse first-passage grid, re-rating each candidate:
q90 4.86/9.71 -> 5.00/10.00 (2.00), q85 4.07/8.14 -> 5.00/10.00 (IDENTICAL
measurement), q75 3.07/6.13 -> 4.00/6.50 (1.63 - a nearer target). So the
ladder compared win shares taken at ratios from 1.63 to 2.17 and read the
differences as scale. It is why the reach column came out non-monotone in
width (q75 48.5% above q90 42.9%). The stop now snaps to its nearest rung in
log space and the target follows the ratio off it; the pair actually measured
is returned and logged, so a collision reads as a collision.
Also: LadderWinShare guarded against the conditional (fractal) geometry path,
which fills n from m_fracLegCount while leaving idxList empty - a latent
out-of-bounds on a currently-dead path.
New log lines: mean label lifespan and effective n on the label-cache line, the
required-vs-available horizon per rung, and the grid pair the reconciliation
actually measured (its tolerance now scales with the grid skew instead of a flat
5pp).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 01:12:05 -04:00
|
|
|
shortEndAge = age;
|
|
|
|
|
}
|
fix: both-won bars were labelled "do not trade" - resolve by first touch
Removing the min-reward:risk raise let the MEASURED geometry come back with
the target NEARER than the stop (SP500 H1: target 1.62*ATR at q50 of
favourable, stop 3.33*ATR at q75 of adverse). That reopened a branch the
code called unreachable: price can reach +target and -target inside one
horizon, winning in BOTH directions, and those bars fell through to Neutral.
Neutral has only three producers, both-lost is unreachable (you cannot touch
-3.33 without crossing -1.62 first, which wins the short), and timeouts logged
at 1.0% of Neutral - so ~27% of ALL bars were being handed to the model as the
abstain class when a trade either way would have collected its target. The
cleanest positives in the sample, labelled "do not trade", while the fitted
confidence threshold was being asked to find selectivity in what was left.
Resolved by FIRST TOUCH: the target reached earlier is the trade that would
have closed first. Same forward window, no extra lookahead. Same-bar ties stay
Neutral - OHLC cannot order two touches, and unlike an intrabar stop tie there
is no pessimistic side to fall to, so a guess would inject a coin-flip
direction into the target.
Also:
- count both-won and its same-bar tie subset in the prebuild line, so the
share is measured rather than inferred from arithmetic on a log line
- scope the timeout counter to IS, matching the tally it is reported as a
percentage OF; it was incremented over the whole scan and divided by an
in-sample denominator
- clear m_lastBarrierTimedOut at the top of the walk with the excursions, not
at the bottom - the two early returns published the previous bar's verdict
- mark the pass-1 label line PROVISIONAL. It prints the enum fallback because
geometry can only be derived from excursions that do not exist yet, and it
reads exactly like a config change that failed to take effect
FORCES RETRAIN. Both build variants compile 0 errors, 0 warnings.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 15:35:42 -04:00
|
|
|
else
|
|
|
|
|
if(lo <= shortTp)
|
|
|
|
|
{
|
|
|
|
|
shortWon = true;
|
|
|
|
|
shortWonAt = t;
|
fix(labels): overlapping-label sample correction + horizon cap on the scale ladder
Three defects, all surfaced by the 2026-08-17 SP500 H4 run that shipped
stop 4.86 / target 9.71 (width 14.57*ATR, horizon 384).
1. EVERY STANDARD ERROR ASSUMED INDEPENDENT SAMPLES. Triple-barrier labels
started one per bar overlap by the label's lifespan, so n calls are worth
~n/L independent observations (Lopez de Prado, AFML ch. 4 - sample
uniqueness). All three sqrt(p(1-p)/n) sites divided by the RAW count.
The tell: the operating point's null-of-the-maximum gate is family-wise and
should fire on ~5% of eras under the null. Measured fire rates - PAI 47/73
(64%), ConvLSTM 9/24, LSTM 8/21 (38%), CONV 4/62 (6%). CONV, the only model
whose margin distribution admits few bins, sat on the null; the rest cleared
a bar that was too low by ~sqrt(L). PAI's deployed threshold consequently
alternated between the ENDS of its own range era to era (0.10 -> 0.88 ->
0.86 -> 0.66; coverage 16% <-> 73%).
TripleBarrierLabel now records when each label became KNOWABLE - the first
winning touch, or both stops, or the timeout - and the prebuild accumulates
the mean. EffectiveSampleSize() feeds the operating point, the member deploy
gate and the ensemble vote gate. Conservative by construction (n/L is an
upper bound on the damage); gates get harder, never easier.
2. THE SCALE LADDER RAN AWAY, again. Horizon scales as swingMedian*sl*tp, and
since 4d8cb08 reachability is measured OVER that horizon - so a wider rung
buys itself the time that makes it look reachable. Same target -> horizon ->
reach -> target loop the excursion window is kept short to avoid; fixing the
window confusion reopened it through the other door. It walked 128 -> 256 ->
384 bars and stopped at q90, the widest rung there is, with every rung
reading 39-48% against a 20% floor. A floor nothing fails selects nothing.
Rungs whose required horizon exceeds BARRIER_HORIZON_MAX are now rejected -
the same rule ReportGeometryExpectancyScan already applied. It was printing
the shipped pair as CLAMPED and disqualified ('h384!') two lines under the
deriver that chose it: two subsystems, one geometry, opposite verdicts.
3. THE RUNG SNAP DESTROYED THE RATIO IT WAS COMPARING. Both legs snapped
independently to the coarse first-passage grid, re-rating each candidate:
q90 4.86/9.71 -> 5.00/10.00 (2.00), q85 4.07/8.14 -> 5.00/10.00 (IDENTICAL
measurement), q75 3.07/6.13 -> 4.00/6.50 (1.63 - a nearer target). So the
ladder compared win shares taken at ratios from 1.63 to 2.17 and read the
differences as scale. It is why the reach column came out non-monotone in
width (q75 48.5% above q90 42.9%). The stop now snaps to its nearest rung in
log space and the target follows the ratio off it; the pair actually measured
is returned and logged, so a collision reads as a collision.
Also: LadderWinShare guarded against the conditional (fractal) geometry path,
which fills n from m_fracLegCount while leaving idxList empty - a latent
out-of-bounds on a currently-dead path.
New log lines: mean label lifespan and effective n on the label-cache line, the
required-vs-available horizon per rung, and the grid pair the reconciliation
actually measured (its tolerance now scales with the grid skew instead of a flat
5pp).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 01:12:05 -04:00
|
|
|
shortEndAge = age;
|
fix: both-won bars were labelled "do not trade" - resolve by first touch
Removing the min-reward:risk raise let the MEASURED geometry come back with
the target NEARER than the stop (SP500 H1: target 1.62*ATR at q50 of
favourable, stop 3.33*ATR at q75 of adverse). That reopened a branch the
code called unreachable: price can reach +target and -target inside one
horizon, winning in BOTH directions, and those bars fell through to Neutral.
Neutral has only three producers, both-lost is unreachable (you cannot touch
-3.33 without crossing -1.62 first, which wins the short), and timeouts logged
at 1.0% of Neutral - so ~27% of ALL bars were being handed to the model as the
abstain class when a trade either way would have collected its target. The
cleanest positives in the sample, labelled "do not trade", while the fitted
confidence threshold was being asked to find selectivity in what was left.
Resolved by FIRST TOUCH: the target reached earlier is the trade that would
have closed first. Same forward window, no extra lookahead. Same-bar ties stay
Neutral - OHLC cannot order two touches, and unlike an intrabar stop tie there
is no pessimistic side to fall to, so a guess would inject a coin-flip
direction into the target.
Also:
- count both-won and its same-bar tie subset in the prebuild line, so the
share is measured rather than inferred from arithmetic on a log line
- scope the timeout counter to IS, matching the tally it is reported as a
percentage OF; it was incremented over the whole scan and divided by an
in-sample denominator
- clear m_lastBarrierTimedOut at the top of the walk with the excursions, not
at the bottom - the two early returns published the previous bar's verdict
- mark the pass-1 label line PROVISIONAL. It prints the enum fallback because
geometry can only be derived from excursions that do not exist yet, and it
reads exactly like a config change that failed to take effect
FORCES RETRAIN. Both build variants compile 0 errors, 0 warnings.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 15:35:42 -04:00
|
|
|
}
|
feat(ai): triple-barrier labels replace exact-pivot ZigZag targets
The 31:1 class imbalance was self-inflicted by the TARGET, not a property
of the market. Labelling only the exact bar where a ZigZag pivot confirms
gave Buy 1164 / Sell 1164 / Neutral 35841, and every correction mechanism
this codebase accumulated sits downstream of that one choice: the
logit-adjusted loss and its range cap, the prior EMA, the +-3.0 output-bias
seed, balanced-accuracy-then-precision selection with its coverage floor,
the recall floor and its catch-22, the alternation gate, NMS, and the four
oversampling designs that collapsed before them.
The reference this engine is built on (references/neuronetworksbook.pdf
ch. 3.1/3.3) also uses ZigZag, but targets the DIRECTION TO THE NEXT
EXTREMUM on every bar - ~50/50 by construction, with no imbalance to
correct at all. It never had this problem because it never asked "is this
the pivot bar".
Labels are now the triple barrier (Lopez de Prado ch. 3), using the EA's
OWN SL_Mode/TP_Mode: does a trade opened at this bar's close reach its
target before its stop, within a horizon. Buy = long resolves, Sell =
short resolves, Neutral = neither. Consequences:
- dir-precision in the era line stops being a proxy and becomes the win
rate of the strategy under its own exit rules.
- Expected balance ~25/25/50 at the shipped 1:3 (gambler's ruin), i.e.
~2:1 instead of 31:1. Measured and logged at the end of the prebuild.
- Spread is charged on both legs, so it is a NET win rate.
- Intrabar ambiguity resolves to the STOP. OHLC cannot order two touches
inside one bar and the optimistic reading is how a backtested edge
becomes a live loss.
ZigZag stays as input features (EnableSwingContext) and now also supplies
the vertical barrier: the horizon is the median confirmed leg length,
snapped to a coarse ladder. Derived, not configured, and deliberately kept
out of the filename fingerprint - a filename keyed on a measured quantity
orphans a trained model the moment the measurement moves.
Removed, because the premise died with the old target:
- the alternation gate. Correct for pivot labels (a ZigZag cannot emit two
same-type pivots in a row, so a repeat was provably a false fire), and
wrong for barrier labels, which answer each bar independently. It also
took its worst consequence with it: a one-sided model previously got ONE
trade per backtest, a hard blocker on marketplace validation.
- SignalClusterWindow now defaults off - it de-duplicated repeats that are
now real trades. Kept as an opt-in display control.
- LABEL_WINDOW_BARS, the pivot-widening pass, ConfirmedZigZagLabel.
- the era-0 output-bias seed now needs a genuinely dominant class (0.70)
rather than 0.40; at ~50% Neutral a +-3.0 seed is a distortion, not a
correction.
Also fixed, both found while wiring the above:
1. RefreshConvergedSignal sized its buffers from a date delta
(Bars(sym, period, dtStudied, TimeCurrent())). dtStudied is a training
watermark; in the tester it is loaded from a live-chart save AHEAD of
the simulated date, so the interval inverted, Bars() returned ~0, and
the buffer came out at exactly m_historyBars - deep enough for the OHLC
window and far too shallow for the Donchian-50 / 20-bar-return / SMA
extension behind it. Inference silently computed DIFFERENT features
from the ones training learned on, live as well as in the tester. Now
sized from what the feature builder actually needs.
2. The barrier horizon is resolved on the deployed path too. A deployed
model never enters Train(), so it never reached the prebuild, and
OnlineLearnStep reads the horizon as its confirmation delay - left at
the fallback it would have backpropped bars whose barriers had not
resolved. Silent lookahead in the one place that writes to a live model.
SL_Mode/TP_Mode join the weights fingerprint: they define the labels now,
so a model trained at 1:3 must never be silently reused at 1:1. This
re-keys every pre-existing model by design - none were trained on this task.
Inference census extended with the vote gate. LongCondition/ShortCondition
open with a readiness check the refresh counters never see; in the tester it
reduces to "the seeded _optcache.nnw must have LOADED", and if it did not,
every vote is hard-zeroed while the model still answers Buy. The old three
counters would have read that as "the model says Neutral" - false, and a
completely different fix. This is the leading candidate for the
zero-direction backtest and the census can now name it in one run.
Both builds compile 0 errors / 0 warnings. Forces a full retrain.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 20:39:49 -04:00
|
|
|
}
|
diag: is "optimal SL/TP" learnable? Score the features against excursions
Proposed direction: train the net to predict entry/SL/TP that maximise return
and minimise drawdown, rather than to classify direction. Before rebuilding a
head, measure whether the target is learnable at all.
That question splits into two that behave nothing alike:
HOW FAR price travels (MFE/MAE) - essentially volatility, and volatility
clustering is about the most robust regularity in markets.
WHICH WAY it goes first (the asymmetry) - direction, which is what every
noise-floor verdict in this project has been about.
Expectancy comes ONLY from the second. The first buys position sizing and
drawdown control - worth having under prop-firm limits, but not an edge: exit
management on RANDOM entries already moved the payoff ratio 0.92 -> 5.72 with
expectancy FLAT.
Crucially this is NOT already answered. Every MI figure here scored the
triple-barrier label, i.e. one specific question at one fixed geometry. A
noise-floor result there says nothing about whether excursion MAGNITUDE is
learnable - different target, different answer.
Four targets, and the verdict is the CONTRAST, printed explicitly because the
dangerous misreading of "UP clears" is "we can predict profitable trades":
RANGE (up+dn) - realised volatility, included as a POSITIVE CONTROL that
SHOULD clear. Every prior verdict here lacked a control
expected to pass; a range target at the floor indicts the
measurement, not the market.
UP / DOWN - MFE / MAE.
ASYMMETRY - up-dn, the only one that can pay.
Collected inside the walk the label already does (one max, one min per bar).
The early-out when both barriers resolved is GONE: it would have truncated the
excursions at whichever bar tripped the last barrier, making the measurement a
function of the CURRENT SL/TP - the circularity this is trying to escape. The
loop was already bounded by the horizon, so only the average cost moves.
Discretised into 3 EQUAL-FREQUENCY bins, so every downstream piece (block
permutation, null, p-value) is reused unchanged. Equal-frequency because MFE is
fat-tailed and fixed-width bins would put nearly every row in bin 0; it also
pins H(Y) at ln(3)=1.099 for all four, making them comparable to each other and
to the barrier label's ~1.02 instead of confounded by class balance.
Two bugs fixed in this code before it ever ran, both of which would have
produced a plausible quiet wrong answer rather than an error:
- TripleBarrierLabel early-returns on invalid ATR/close BEFORE the point the
accumulators were reset, so one bar's excursions would be cached under
another bar's index. Cleared at the top now, ahead of every return.
- An unresolvable bar is still flagged as labelled but carries excursions of
exactly 0. Under equal-frequency binning a block of identical zeros drags
the lowest cut onto zero and a third of the sample lands in one
uninformative bin - a depressed score that reads as "not predictable", a
false negative in the direction that would wrongly kill the idea. Rows
where both excursions are zero are dropped; price cannot travel zero both
ways over a whole horizon.
Read-only diagnostic. No topology or label change: no retrain of its own.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-07 10:22:41 -04:00
|
|
|
//--- The early-out that used to sit here (both sides resolved -> break) is GONE, because the
|
|
|
|
|
//--- excursion accumulators above must see the whole horizon and it would have truncated them at
|
|
|
|
|
//--- whichever bar happened to trip the last barrier - making the measured excursion a function of
|
|
|
|
|
//--- the current SL/TP, which is exactly the circularity being escaped. The loop was already
|
|
|
|
|
//--- bounded by m_barrierHorizonBars, so the worst case is unchanged and only the average moves.
|
|
|
|
|
}
|
|
|
|
|
if(maxHigh > -DBL_MAX && minLow < DBL_MAX)
|
|
|
|
|
{
|
|
|
|
|
//--- Same spread convention as the barriers: a long fills at close+spread, so its favourable
|
|
|
|
|
//--- excursion is measured from that fill and its adverse excursion likewise. Clamped at zero -
|
|
|
|
|
//--- a horizon whose every high sits below the fill has no favourable excursion, not a negative one.
|
|
|
|
|
m_lastExcUp = MathMax((maxHigh - (entry + spread)) / atr, 0.0);
|
|
|
|
|
m_lastExcDown = MathMax(((entry + spread) - minLow) / atr, 0.0);
|
feat(ai): triple-barrier labels replace exact-pivot ZigZag targets
The 31:1 class imbalance was self-inflicted by the TARGET, not a property
of the market. Labelling only the exact bar where a ZigZag pivot confirms
gave Buy 1164 / Sell 1164 / Neutral 35841, and every correction mechanism
this codebase accumulated sits downstream of that one choice: the
logit-adjusted loss and its range cap, the prior EMA, the +-3.0 output-bias
seed, balanced-accuracy-then-precision selection with its coverage floor,
the recall floor and its catch-22, the alternation gate, NMS, and the four
oversampling designs that collapsed before them.
The reference this engine is built on (references/neuronetworksbook.pdf
ch. 3.1/3.3) also uses ZigZag, but targets the DIRECTION TO THE NEXT
EXTREMUM on every bar - ~50/50 by construction, with no imbalance to
correct at all. It never had this problem because it never asked "is this
the pivot bar".
Labels are now the triple barrier (Lopez de Prado ch. 3), using the EA's
OWN SL_Mode/TP_Mode: does a trade opened at this bar's close reach its
target before its stop, within a horizon. Buy = long resolves, Sell =
short resolves, Neutral = neither. Consequences:
- dir-precision in the era line stops being a proxy and becomes the win
rate of the strategy under its own exit rules.
- Expected balance ~25/25/50 at the shipped 1:3 (gambler's ruin), i.e.
~2:1 instead of 31:1. Measured and logged at the end of the prebuild.
- Spread is charged on both legs, so it is a NET win rate.
- Intrabar ambiguity resolves to the STOP. OHLC cannot order two touches
inside one bar and the optimistic reading is how a backtested edge
becomes a live loss.
ZigZag stays as input features (EnableSwingContext) and now also supplies
the vertical barrier: the horizon is the median confirmed leg length,
snapped to a coarse ladder. Derived, not configured, and deliberately kept
out of the filename fingerprint - a filename keyed on a measured quantity
orphans a trained model the moment the measurement moves.
Removed, because the premise died with the old target:
- the alternation gate. Correct for pivot labels (a ZigZag cannot emit two
same-type pivots in a row, so a repeat was provably a false fire), and
wrong for barrier labels, which answer each bar independently. It also
took its worst consequence with it: a one-sided model previously got ONE
trade per backtest, a hard blocker on marketplace validation.
- SignalClusterWindow now defaults off - it de-duplicated repeats that are
now real trades. Kept as an opt-in display control.
- LABEL_WINDOW_BARS, the pivot-widening pass, ConfirmedZigZagLabel.
- the era-0 output-bias seed now needs a genuinely dominant class (0.70)
rather than 0.40; at ~50% Neutral a +-3.0 seed is a distortion, not a
correction.
Also fixed, both found while wiring the above:
1. RefreshConvergedSignal sized its buffers from a date delta
(Bars(sym, period, dtStudied, TimeCurrent())). dtStudied is a training
watermark; in the tester it is loaded from a live-chart save AHEAD of
the simulated date, so the interval inverted, Bars() returned ~0, and
the buffer came out at exactly m_historyBars - deep enough for the OHLC
window and far too shallow for the Donchian-50 / 20-bar-return / SMA
extension behind it. Inference silently computed DIFFERENT features
from the ones training learned on, live as well as in the tester. Now
sized from what the feature builder actually needs.
2. The barrier horizon is resolved on the deployed path too. A deployed
model never enters Train(), so it never reached the prebuild, and
OnlineLearnStep reads the horizon as its confirmation delay - left at
the fallback it would have backpropped bars whose barriers had not
resolved. Silent lookahead in the one place that writes to a live model.
SL_Mode/TP_Mode join the weights fingerprint: they define the labels now,
so a model trained at 1:3 must never be silently reused at 1:1. This
re-keys every pre-existing model by design - none were trained on this task.
Inference census extended with the vote gate. LongCondition/ShortCondition
open with a readiness check the refresh counters never see; in the tester it
reduces to "the seeded _optcache.nnw must have LOADED", and if it did not,
every vote is hard-zeroed while the model still answers Buy. The old three
counters would have read that as "the model says Neutral" - false, and a
completely different fix. This is the leading candidate for the
zero-direction backtest and the census can now name it in one run.
Both builds compile 0 errors / 0 warnings. Forces a full retrain.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 20:39:49 -04:00
|
|
|
}
|
fix(labels): overlapping-label sample correction + horizon cap on the scale ladder
Three defects, all surfaced by the 2026-08-17 SP500 H4 run that shipped
stop 4.86 / target 9.71 (width 14.57*ATR, horizon 384).
1. EVERY STANDARD ERROR ASSUMED INDEPENDENT SAMPLES. Triple-barrier labels
started one per bar overlap by the label's lifespan, so n calls are worth
~n/L independent observations (Lopez de Prado, AFML ch. 4 - sample
uniqueness). All three sqrt(p(1-p)/n) sites divided by the RAW count.
The tell: the operating point's null-of-the-maximum gate is family-wise and
should fire on ~5% of eras under the null. Measured fire rates - PAI 47/73
(64%), ConvLSTM 9/24, LSTM 8/21 (38%), CONV 4/62 (6%). CONV, the only model
whose margin distribution admits few bins, sat on the null; the rest cleared
a bar that was too low by ~sqrt(L). PAI's deployed threshold consequently
alternated between the ENDS of its own range era to era (0.10 -> 0.88 ->
0.86 -> 0.66; coverage 16% <-> 73%).
TripleBarrierLabel now records when each label became KNOWABLE - the first
winning touch, or both stops, or the timeout - and the prebuild accumulates
the mean. EffectiveSampleSize() feeds the operating point, the member deploy
gate and the ensemble vote gate. Conservative by construction (n/L is an
upper bound on the damage); gates get harder, never easier.
2. THE SCALE LADDER RAN AWAY, again. Horizon scales as swingMedian*sl*tp, and
since 4d8cb08 reachability is measured OVER that horizon - so a wider rung
buys itself the time that makes it look reachable. Same target -> horizon ->
reach -> target loop the excursion window is kept short to avoid; fixing the
window confusion reopened it through the other door. It walked 128 -> 256 ->
384 bars and stopped at q90, the widest rung there is, with every rung
reading 39-48% against a 20% floor. A floor nothing fails selects nothing.
Rungs whose required horizon exceeds BARRIER_HORIZON_MAX are now rejected -
the same rule ReportGeometryExpectancyScan already applied. It was printing
the shipped pair as CLAMPED and disqualified ('h384!') two lines under the
deriver that chose it: two subsystems, one geometry, opposite verdicts.
3. THE RUNG SNAP DESTROYED THE RATIO IT WAS COMPARING. Both legs snapped
independently to the coarse first-passage grid, re-rating each candidate:
q90 4.86/9.71 -> 5.00/10.00 (2.00), q85 4.07/8.14 -> 5.00/10.00 (IDENTICAL
measurement), q75 3.07/6.13 -> 4.00/6.50 (1.63 - a nearer target). So the
ladder compared win shares taken at ratios from 1.63 to 2.17 and read the
differences as scale. It is why the reach column came out non-monotone in
width (q75 48.5% above q90 42.9%). The stop now snaps to its nearest rung in
log space and the target follows the ratio off it; the pair actually measured
is returned and logged, so a collision reads as a collision.
Also: LadderWinShare guarded against the conditional (fractal) geometry path,
which fills n from m_fracLegCount while leaving idxList empty - a latent
out-of-bounds on a currently-dead path.
New log lines: mean label lifespan and effective n on the label-cache line, the
required-vs-available horizon per rung, and the grid pair the reconciliation
actually measured (its tolerance now scales with the grid skew instead of a flat
5pp).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 01:12:05 -04:00
|
|
|
//--- WHEN THIS LABEL BECAME KNOWABLE, which is what the overlap correction needs - see
|
|
|
|
|
//--- m_lastLabelLifespan. The label is fixed by the FIRST target touched, because the both-won branch
|
|
|
|
|
//--- below resolves by first touch: once one side wins, no later touch on the other side can change the
|
|
|
|
|
//--- answer, and an earlier one would already have been recorded. So a bar with a winner is determined at
|
|
|
|
|
//--- that win, however long the other side takes.
|
|
|
|
|
//--- With no winner the answer is Neutral, but it is not KNOWN to be Neutral until every side that could
|
|
|
|
|
//--- still win has stopped out - or, failing that, until the horizon expires and the trade times out.
|
|
|
|
|
if(longWon || shortWon)
|
|
|
|
|
{
|
|
|
|
|
if(longWon && shortWon)
|
|
|
|
|
m_lastLabelLifespan = (int)MathMin(longEndAge, shortEndAge);
|
|
|
|
|
else
|
|
|
|
|
m_lastLabelLifespan = (longWon ? longEndAge : shortEndAge);
|
|
|
|
|
}
|
|
|
|
|
else
|
|
|
|
|
if(longLost && shortLost)
|
|
|
|
|
m_lastLabelLifespan = (int)MathMax(longEndAge, shortEndAge);
|
|
|
|
|
else
|
|
|
|
|
m_lastLabelLifespan = walkedAge; // a live side ran out of horizon: the timeout IS the decision
|
fix: the deploy gate was benchmarking a win rate against a label frequency
The gate rests on an invariant stated at ExpertSignalAIBase.mqh:199 - under a
driftless walk P(touch +k before -m) is m/(m+k), and break-even for a k:m trade
is ALSO m/(m+k), so "beats chance" and "is profitable" are the same test.
That invariant needs reward >= risk, and the measured geometry no longer
satisfies it. With target 1.62*ATR and stop 3.33*ATR, break-even is 67.3%, but
both-won bars were stripped out of Buy and Sell so the label base rate read
37.5%. chancePrecPct is max(BuyTotal,SellTotal)/bars, so the gate was clearing
models nearly 30pp short of break-even: 42% "directional precision" is +4 sigma
against 37.5% and loses money on every single trade. Live since 217b9bc.
Root cause is that label agreement stopped being the same question as trade
profitability. Buy implies winLong, but the converse fails on every both-won
bar, and the label can only name one of two directions that both pay.
So stop asking the model whether it matched a label and start asking whether
its trade paid:
- cache winLong/winShort per bar beside the label, under the same validity
flag; published from the barrier walk before the collapse to 3 classes
- dirPrecPct now counts wins on the side actually called
- chancePrecPct is max(P(winLong), P(winShort)), MEASURED - the textbook
m/(m+k) would credit SP500's drift to the model
- the NMS "what would I have made" pair, the live-fired precision, and the
IS/OOS cumulative win rates all move to the same test. IS and OOS are read
side by side as the overfitting signal, so measuring one in wins and the
other in agreement would put a fixed gap between them that has nothing to do
with generalization
- the confidence threshold is FITTED on wins too, so the operating point
maximises what the gate grades
- per-class label-agreement precision is still computed and logged; it is the
right diagnostic for class separation, just not for a deploy decision
- era line renamed dir-precision -> win-rate, chance -> chance=break-even
Both build variants compile 0 errors, 0 warnings.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 16:00:51 -04:00
|
|
|
//--- Published BEFORE the collapse to a single label, because the collapse cannot be undone afterwards
|
|
|
|
|
//--- and these are what profitability is actually a function of. Every return below this point carries
|
|
|
|
|
//--- them; the early returns above leave them false, which is correct - an unresolvable bar has no
|
|
|
|
|
//--- winning direction.
|
|
|
|
|
m_lastWinLong = longWon;
|
|
|
|
|
m_lastWinShort = shortWon;
|
feat(ai): triple-barrier labels replace exact-pivot ZigZag targets
The 31:1 class imbalance was self-inflicted by the TARGET, not a property
of the market. Labelling only the exact bar where a ZigZag pivot confirms
gave Buy 1164 / Sell 1164 / Neutral 35841, and every correction mechanism
this codebase accumulated sits downstream of that one choice: the
logit-adjusted loss and its range cap, the prior EMA, the +-3.0 output-bias
seed, balanced-accuracy-then-precision selection with its coverage floor,
the recall floor and its catch-22, the alternation gate, NMS, and the four
oversampling designs that collapsed before them.
The reference this engine is built on (references/neuronetworksbook.pdf
ch. 3.1/3.3) also uses ZigZag, but targets the DIRECTION TO THE NEXT
EXTREMUM on every bar - ~50/50 by construction, with no imbalance to
correct at all. It never had this problem because it never asked "is this
the pivot bar".
Labels are now the triple barrier (Lopez de Prado ch. 3), using the EA's
OWN SL_Mode/TP_Mode: does a trade opened at this bar's close reach its
target before its stop, within a horizon. Buy = long resolves, Sell =
short resolves, Neutral = neither. Consequences:
- dir-precision in the era line stops being a proxy and becomes the win
rate of the strategy under its own exit rules.
- Expected balance ~25/25/50 at the shipped 1:3 (gambler's ruin), i.e.
~2:1 instead of 31:1. Measured and logged at the end of the prebuild.
- Spread is charged on both legs, so it is a NET win rate.
- Intrabar ambiguity resolves to the STOP. OHLC cannot order two touches
inside one bar and the optimistic reading is how a backtested edge
becomes a live loss.
ZigZag stays as input features (EnableSwingContext) and now also supplies
the vertical barrier: the horizon is the median confirmed leg length,
snapped to a coarse ladder. Derived, not configured, and deliberately kept
out of the filename fingerprint - a filename keyed on a measured quantity
orphans a trained model the moment the measurement moves.
Removed, because the premise died with the old target:
- the alternation gate. Correct for pivot labels (a ZigZag cannot emit two
same-type pivots in a row, so a repeat was provably a false fire), and
wrong for barrier labels, which answer each bar independently. It also
took its worst consequence with it: a one-sided model previously got ONE
trade per backtest, a hard blocker on marketplace validation.
- SignalClusterWindow now defaults off - it de-duplicated repeats that are
now real trades. Kept as an opt-in display control.
- LABEL_WINDOW_BARS, the pivot-widening pass, ConfirmedZigZagLabel.
- the era-0 output-bias seed now needs a genuinely dominant class (0.70)
rather than 0.40; at ~50% Neutral a +-3.0 seed is a distortion, not a
correction.
Also fixed, both found while wiring the above:
1. RefreshConvergedSignal sized its buffers from a date delta
(Bars(sym, period, dtStudied, TimeCurrent())). dtStudied is a training
watermark; in the tester it is loaded from a live-chart save AHEAD of
the simulated date, so the interval inverted, Bars() returned ~0, and
the buffer came out at exactly m_historyBars - deep enough for the OHLC
window and far too shallow for the Donchian-50 / 20-bar-return / SMA
extension behind it. Inference silently computed DIFFERENT features
from the ones training learned on, live as well as in the tester. Now
sized from what the feature builder actually needs.
2. The barrier horizon is resolved on the deployed path too. A deployed
model never enters Train(), so it never reached the prebuild, and
OnlineLearnStep reads the horizon as its confirmation delay - left at
the fallback it would have backpropped bars whose barriers had not
resolved. Silent lookahead in the one place that writes to a live model.
SL_Mode/TP_Mode join the weights fingerprint: they define the labels now,
so a model trained at 1:3 must never be silently reused at 1:1. This
re-keys every pre-existing model by design - none were trained on this task.
Inference census extended with the vote gate. LongCondition/ShortCondition
open with a readiness check the refresh counters never see; in the tester it
reduces to "the seeded _optcache.nnw must have LOADED", and if it did not,
every vote is hard-zeroed while the model still answers Buy. The old three
counters would have read that as "the model says Neutral" - false, and a
completely different fix. This is the leading candidate for the
zero-direction backtest and the census can now name it in one run.
Both builds compile 0 errors / 0 warnings. Forces a full retrain.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 20:39:49 -04:00
|
|
|
if(longWon && !shortWon)
|
|
|
|
|
return Buy;
|
|
|
|
|
if(shortWon && !longWon)
|
|
|
|
|
return Sell;
|
fix: both-won bars were labelled "do not trade" - resolve by first touch
Removing the min-reward:risk raise let the MEASURED geometry come back with
the target NEARER than the stop (SP500 H1: target 1.62*ATR at q50 of
favourable, stop 3.33*ATR at q75 of adverse). That reopened a branch the
code called unreachable: price can reach +target and -target inside one
horizon, winning in BOTH directions, and those bars fell through to Neutral.
Neutral has only three producers, both-lost is unreachable (you cannot touch
-3.33 without crossing -1.62 first, which wins the short), and timeouts logged
at 1.0% of Neutral - so ~27% of ALL bars were being handed to the model as the
abstain class when a trade either way would have collected its target. The
cleanest positives in the sample, labelled "do not trade", while the fitted
confidence threshold was being asked to find selectivity in what was left.
Resolved by FIRST TOUCH: the target reached earlier is the trade that would
have closed first. Same forward window, no extra lookahead. Same-bar ties stay
Neutral - OHLC cannot order two touches, and unlike an intrabar stop tie there
is no pessimistic side to fall to, so a guess would inject a coin-flip
direction into the target.
Also:
- count both-won and its same-bar tie subset in the prebuild line, so the
share is measured rather than inferred from arithmetic on a log line
- scope the timeout counter to IS, matching the tally it is reported as a
percentage OF; it was incremented over the whole scan and divided by an
in-sample denominator
- clear m_lastBarrierTimedOut at the top of the walk with the excursions, not
at the bottom - the two early returns published the previous bar's verdict
- mark the pass-1 label line PROVISIONAL. It prints the enum fallback because
geometry can only be derived from excursions that do not exist yet, and it
reads exactly like a config change that failed to take effect
FORCES RETRAIN. Both build variants compile 0 errors, 0 warnings.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 15:35:42 -04:00
|
|
|
//--- BOTH TARGETS REACHED. This comment used to say the branch was "unreachable" for every shipped SL/TP
|
|
|
|
|
//--- pairing, and while reward >= risk that was true - reaching one side's stop necessarily crossed the
|
|
|
|
|
//--- other side's nearer target first, so the two outcomes were complementary. Removing the
|
|
|
|
|
//--- minimum-reward:risk raise (2026-08-09) ended that: the MEASURED geometry puts the target at the q50
|
|
|
|
|
//--- of favourable travel and the stop at the q75 of adverse, i.e. target CLOSER than stop, and price
|
|
|
|
|
//--- that swings +target then -target inside one horizon wins in BOTH directions.
|
|
|
|
|
//---
|
|
|
|
|
//--- Falling through to Neutral here was actively harmful, and not marginally: on SP500 H1 it labelled
|
|
|
|
|
//--- ~27% of all bars "do not trade" when a trade in EITHER direction would have collected its target.
|
|
|
|
|
//--- Those are the cleanest positives in the sample, and they were being handed to the model as the
|
|
|
|
|
//--- abstain class - while the confidence threshold downstream was being asked to find selectivity in
|
|
|
|
|
//--- what was left.
|
|
|
|
|
//---
|
|
|
|
|
//--- Resolved by FIRST TOUCH: the target reached earlier is the trade that would have closed first, and
|
|
|
|
|
//--- it is the direction the bar actually moved in before it reversed. Deterministic, and no more
|
|
|
|
|
//--- lookahead than any other part of this walk - it reads the same forward window.
|
|
|
|
|
if(longWon && shortWon)
|
|
|
|
|
{
|
|
|
|
|
m_lastBarrierBothWon = true;
|
|
|
|
|
if(longWonAt > shortWonAt) // larger t = earlier bar, see the declaration
|
|
|
|
|
return Buy;
|
|
|
|
|
if(shortWonAt > longWonAt)
|
|
|
|
|
return Sell;
|
|
|
|
|
//--- Same bar. OHLC carries no intrabar ordering, and the whole file's convention is to refuse the
|
|
|
|
|
//--- ordering it cannot see rather than guess it (BARRIER_TIE_GOES_TO_STOP). Unlike the stop tie
|
|
|
|
|
//--- there is no pessimistic side to fall to - both directions won - so the bar stays Neutral and
|
|
|
|
|
//--- is counted, because a guess here would inject a coin-flip direction into the training target.
|
|
|
|
|
//--- Requires a single bar spanning both targets, ~2x the measured target in range, so it should be
|
|
|
|
|
//--- rare; m_labelPrebuildBothWonTieCount is what confirms that rather than assuming it.
|
|
|
|
|
m_lastBarrierBothWonTied = true;
|
|
|
|
|
return Neutral;
|
|
|
|
|
}
|
|
|
|
|
//--- Neither side resolved AT ALL = the vertical barrier is what ended it. Recorded separately from a
|
|
|
|
|
//--- stop-out because only this outcome says the horizon is too short - see m_lastBarrierTimedOut.
|
fix: refuse invalid SL/TP, fix the unreachable deploy floor, scale the horizon
Three defects found by reading the 2026-08-01 training logs, all of which
only became visible because the relabel made the numbers mean something.
1. A STALE ENUM TRAINED FOUR MODELS ON THE WRONG TARGET.
`OnInit: trade settings snapshot - SL_Mode=1 TP_Mode=-101`
-101 was TP_PREV_SWING, deleted from TAKE_PROFIT_MODE on 2026-07-31 in
7eb48f5. MetaTrader does not validate a saved enum input against the
enum's current members, so charts saved before that kept the old
integer. BarrierMultiples()'s `if(tpMult <= 0.0) tpMult = slMult;`
then quietly turned it into a 1:1 barrier, and all four topologies
trained ~250 eras against a strategy nobody selected - while the log
reported "target 1.00*ATR" as though it were configured.
Since the relabel these two inputs ARE the label definition, so this
is not a bad trade setting, it is a wrong dataset. ValidateBarrier-
Inputs() now refuses to start (INIT_FAILED + Alert + an explicit fix)
on any value that is not an enum member. Members are enumerated rather
than range-checked because both enums are sparse and carry negative
sentinels, so no min/max test can tell a legal value from a deleted
one - which is the entire failure mode. The fallback survives as
belt-and-braces but now announces itself: a fallback that cannot say
it fired is indistinguishable from correct behaviour.
2. THE DEPLOYABILITY FLOOR BECAME MATHEMATICALLY UNREACHABLE.
`tradeableOK` required `dirPrecPct >= baseRatePct`, where baseRatePct
is Buy+Sell as a share of all bars. At the old exact-pivot target that
was ~6%, so "beat the base rate" read as "beat chance" and the test
looked sound. Triple-barrier labels put it at ~83%, so the gate now
demanded 83% directional precision - impossible by construction.
Observed live: all four topologies cycling "PLATEAU stage 3 ... nothing
safe to deploy" at a perfectly healthy 43-45% precision, with no
checkpoint able to ship however good it got.
Replaced with ZERO-SKILL precision, max(Buy,Sell)/allBars: exactly the
score of the degenerate always-call-one-direction model this floor
exists to reject. Correct at any base rate - ~43% on the current
labels, ~3% on the old rare-pivot ones. The era line now prints
"(chance N%, edge +Mpp)" beside the selection score, because 44%
precision is excellent against a 3% chance level and worthless against
a 43% one, and reading the first as the second is what made tonight's
run look better than it was.
3. THE HORIZON IGNORED THE BARRIER GEOMETRY.
ComputeBarrierHorizonBars() returned the median ZigZag leg, which
measures how long a ~1 ATR move takes and says nothing about how long
the CONFIGURED barrier needs. First-passage time out of [-m,+k] scales
with m*k, so a 1:3 barrier takes ~3x as long as 1:1; the unscaled
horizon would have timed out most 1:3 trades and pushed Neutral
straight back up, re-creating the imbalance the relabel removes.
Now multiplied by slMult*tpMult, calibrated against a real measurement
rather than assumed: the accidental 1:1 run resolved at horizon 12 with
only 16.7% timeouts, so the swing median is the right scale at m*k=1.
Verifiable, not just asserted: the prebuild now counts barriers that
ended on the VERTICAL barrier and reports them as a share of Neutral.
Neutral conflates "timed out" with "stopped out" and only the first
indicts the horizon.
Both builds compile 0 errors / 0 warnings. Forces a retrain - correcting
TP_Mode re-keys the fingerprint (|TB:1:-101 -> |TB:1:3), which is right:
no existing model was trained on the intended target.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 00:30:49 -04:00
|
|
|
m_lastBarrierTimedOut = (!longWon && !longLost && !shortWon && !shortLost);
|
fix: both-won bars were labelled "do not trade" - resolve by first touch
Removing the min-reward:risk raise let the MEASURED geometry come back with
the target NEARER than the stop (SP500 H1: target 1.62*ATR at q50 of
favourable, stop 3.33*ATR at q75 of adverse). That reopened a branch the
code called unreachable: price can reach +target and -target inside one
horizon, winning in BOTH directions, and those bars fell through to Neutral.
Neutral has only three producers, both-lost is unreachable (you cannot touch
-3.33 without crossing -1.62 first, which wins the short), and timeouts logged
at 1.0% of Neutral - so ~27% of ALL bars were being handed to the model as the
abstain class when a trade either way would have collected its target. The
cleanest positives in the sample, labelled "do not trade", while the fitted
confidence threshold was being asked to find selectivity in what was left.
Resolved by FIRST TOUCH: the target reached earlier is the trade that would
have closed first. Same forward window, no extra lookahead. Same-bar ties stay
Neutral - OHLC cannot order two touches, and unlike an intrabar stop tie there
is no pessimistic side to fall to, so a guess would inject a coin-flip
direction into the target.
Also:
- count both-won and its same-bar tie subset in the prebuild line, so the
share is measured rather than inferred from arithmetic on a log line
- scope the timeout counter to IS, matching the tally it is reported as a
percentage OF; it was incremented over the whole scan and divided by an
in-sample denominator
- clear m_lastBarrierTimedOut at the top of the walk with the excursions, not
at the bottom - the two early returns published the previous bar's verdict
- mark the pass-1 label line PROVISIONAL. It prints the enum fallback because
geometry can only be derived from excursions that do not exist yet, and it
reads exactly like a config change that failed to take effect
FORCES RETRAIN. Both build variants compile 0 errors, 0 warnings.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 15:35:42 -04:00
|
|
|
return Neutral; // timed out, stopped out, or an unorderable both-won tie - nothing tradeable here
|
feat(ai): triple-barrier labels replace exact-pivot ZigZag targets
The 31:1 class imbalance was self-inflicted by the TARGET, not a property
of the market. Labelling only the exact bar where a ZigZag pivot confirms
gave Buy 1164 / Sell 1164 / Neutral 35841, and every correction mechanism
this codebase accumulated sits downstream of that one choice: the
logit-adjusted loss and its range cap, the prior EMA, the +-3.0 output-bias
seed, balanced-accuracy-then-precision selection with its coverage floor,
the recall floor and its catch-22, the alternation gate, NMS, and the four
oversampling designs that collapsed before them.
The reference this engine is built on (references/neuronetworksbook.pdf
ch. 3.1/3.3) also uses ZigZag, but targets the DIRECTION TO THE NEXT
EXTREMUM on every bar - ~50/50 by construction, with no imbalance to
correct at all. It never had this problem because it never asked "is this
the pivot bar".
Labels are now the triple barrier (Lopez de Prado ch. 3), using the EA's
OWN SL_Mode/TP_Mode: does a trade opened at this bar's close reach its
target before its stop, within a horizon. Buy = long resolves, Sell =
short resolves, Neutral = neither. Consequences:
- dir-precision in the era line stops being a proxy and becomes the win
rate of the strategy under its own exit rules.
- Expected balance ~25/25/50 at the shipped 1:3 (gambler's ruin), i.e.
~2:1 instead of 31:1. Measured and logged at the end of the prebuild.
- Spread is charged on both legs, so it is a NET win rate.
- Intrabar ambiguity resolves to the STOP. OHLC cannot order two touches
inside one bar and the optimistic reading is how a backtested edge
becomes a live loss.
ZigZag stays as input features (EnableSwingContext) and now also supplies
the vertical barrier: the horizon is the median confirmed leg length,
snapped to a coarse ladder. Derived, not configured, and deliberately kept
out of the filename fingerprint - a filename keyed on a measured quantity
orphans a trained model the moment the measurement moves.
Removed, because the premise died with the old target:
- the alternation gate. Correct for pivot labels (a ZigZag cannot emit two
same-type pivots in a row, so a repeat was provably a false fire), and
wrong for barrier labels, which answer each bar independently. It also
took its worst consequence with it: a one-sided model previously got ONE
trade per backtest, a hard blocker on marketplace validation.
- SignalClusterWindow now defaults off - it de-duplicated repeats that are
now real trades. Kept as an opt-in display control.
- LABEL_WINDOW_BARS, the pivot-widening pass, ConfirmedZigZagLabel.
- the era-0 output-bias seed now needs a genuinely dominant class (0.70)
rather than 0.40; at ~50% Neutral a +-3.0 seed is a distortion, not a
correction.
Also fixed, both found while wiring the above:
1. RefreshConvergedSignal sized its buffers from a date delta
(Bars(sym, period, dtStudied, TimeCurrent())). dtStudied is a training
watermark; in the tester it is loaded from a live-chart save AHEAD of
the simulated date, so the interval inverted, Bars() returned ~0, and
the buffer came out at exactly m_historyBars - deep enough for the OHLC
window and far too shallow for the Donchian-50 / 20-bar-return / SMA
extension behind it. Inference silently computed DIFFERENT features
from the ones training learned on, live as well as in the tester. Now
sized from what the feature builder actually needs.
2. The barrier horizon is resolved on the deployed path too. A deployed
model never enters Train(), so it never reached the prebuild, and
OnlineLearnStep reads the horizon as its confirmation delay - left at
the fallback it would have backpropped bars whose barriers had not
resolved. Silent lookahead in the one place that writes to a live model.
SL_Mode/TP_Mode join the weights fingerprint: they define the labels now,
so a model trained at 1:3 must never be silently reused at 1:1. This
re-keys every pre-existing model by design - none were trained on this task.
Inference census extended with the vote gate. LongCondition/ShortCondition
open with a readiness check the refresh counters never see; in the tester it
reduces to "the seeded _optcache.nnw must have LOADED", and if it did not,
every vote is hard-zeroed while the model still answers Buy. The old three
counters would have read that as "the model says Neutral" - false, and a
completely different fix. This is the leading candidate for the
zero-direction backtest and the census can now name it in one run.
Both builds compile 0 errors / 0 warnings. Forces a full retrain.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 20:39:49 -04:00
|
|
|
}
|
|
|
|
|
//+------------------------------------------------------------------+
|
|
|
|
|
//| Median distance in bars between consecutive confirmed ZigZag |
|
|
|
|
|
//| pivots - this symbol/timeframe's own swing horizon, and what the |
|
|
|
|
|
//| vertical barrier is set to. Snapped to a coarse ladder so the |
|
|
|
|
|
//| estimate has to move ~30% to change the answer; see the |
|
|
|
|
|
//| BARRIER_HORIZON_* constants for why the quantization matters more |
|
|
|
|
|
//| than the precision (an unquantized horizon that drifted as |
|
|
|
|
|
//| history downloaded would relabel a partly-trained model's |
|
|
|
|
|
//| targets mid-run). |
|
|
|
|
|
//| |
|
|
|
|
|
//| Reads only pivots old enough to be non-repainting, for the same |
|
|
|
|
|
//| reason every other ZigZag read in this class does. |
|
|
|
|
|
//+------------------------------------------------------------------+
|
|
|
|
|
int CExpertSignalAIBase::ComputeBarrierHorizonBars(int bars)
|
|
|
|
|
{
|
fix(labels): correct EffectiveSampleSize clamp order, share the horizon ladder, retract a false justification
Self-review of 1540ba8 against the FULL 6,930-era log rather than the first
three minutes of it. Three corrections.
1. EffectiveSampleSize() clamped in the wrong order. MathMax(2, MathMin(eff,
rawN)) returns 2 when rawN is 1 - an effective sample LARGER than the raw
one, shrinking the SE in exactly the direction the function exists to
prevent. Floor first, cap at rawN last.
2. The horizon cap rejected on the CEILING only, and said so as though that
made the label untruncated. It does not: the horizon ladder also snaps DOWN,
so a pair needing 317 bars is granted 256 and is silently truncated without
ever being flagged CLAMPED. Added SnapHorizonToLadder() / GrantedHorizonBars()
and the scale ladder now reports "needs N gets M" per rung. Rejection stays on
the ceiling alone - matching ReportGeometryExpectancyScan's '!' exactly, which
was the point - because rejecting on the snap-down would select rungs for
landing just above a ladder point rather than for anything about the market.
ComputeBarrierHorizonBars' private copy of the ladder is gone; there is now
one copy, which is the whole reason RequiredHorizonBars was factored out.
3. RETRACTED THE JUSTIFICATION IN 1540ba8's COMMENTS. That commit claimed the
overlap correction was needed because the operating point's null-of-the-
maximum gate fired on 47/73 Perceptron eras (64%) where a family-wise test
should fire on ~5%. Those 73 fits were the first three minutes of a
six-and-a-half-hour run. Over the full run:
PAI 47/3214 = 1.5% HYB 30/1200 = 2.5%
CONV 4/63 = 6.3% LSTM 75/915 = 8.2%
All at or below the null. The gate from 7414570 is working as designed and
PAI's 47 clears were a cold-start transient never repeated in 3,141 later
fits; its threshold over the run's second half has sd 0.01. The overlap
correction is still right - sqrt(p(1-p)/n) on overlapping labels is the wrong
formula - but it fixes no observed failure, and it costs nothing today
because no model is near the deploy line.
WHAT THE FULL RUN DOES CONFIRM, unchanged: the geometry ran away exactly as
described (2.00/6.00 h128 -> 3.49/6.99 h256 -> 4.86/9.71 h384, three passes,
stopping at q90 because the quantile ladder ended), the label stayed long-skewed
at Buy 42.9% / Sell 22.6%, and no checkpoint on any of the four models ever
cleared the deployability floor. Pooled declustered win rates: PAI 31.70%,
HYB 31.02%, LSTM 32.69%, CONV 31.57% - every one 4-6pp below the 37% always-long
chance rate and 1-2.7pp below the 33.7% cost-adjusted break-even.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 07:30:19 -04:00
|
|
|
//--- The ladder itself now lives in SnapHorizonToLadder(), which this function ends by calling.
|
feat(ai): triple-barrier labels replace exact-pivot ZigZag targets
The 31:1 class imbalance was self-inflicted by the TARGET, not a property
of the market. Labelling only the exact bar where a ZigZag pivot confirms
gave Buy 1164 / Sell 1164 / Neutral 35841, and every correction mechanism
this codebase accumulated sits downstream of that one choice: the
logit-adjusted loss and its range cap, the prior EMA, the +-3.0 output-bias
seed, balanced-accuracy-then-precision selection with its coverage floor,
the recall floor and its catch-22, the alternation gate, NMS, and the four
oversampling designs that collapsed before them.
The reference this engine is built on (references/neuronetworksbook.pdf
ch. 3.1/3.3) also uses ZigZag, but targets the DIRECTION TO THE NEXT
EXTREMUM on every bar - ~50/50 by construction, with no imbalance to
correct at all. It never had this problem because it never asked "is this
the pivot bar".
Labels are now the triple barrier (Lopez de Prado ch. 3), using the EA's
OWN SL_Mode/TP_Mode: does a trade opened at this bar's close reach its
target before its stop, within a horizon. Buy = long resolves, Sell =
short resolves, Neutral = neither. Consequences:
- dir-precision in the era line stops being a proxy and becomes the win
rate of the strategy under its own exit rules.
- Expected balance ~25/25/50 at the shipped 1:3 (gambler's ruin), i.e.
~2:1 instead of 31:1. Measured and logged at the end of the prebuild.
- Spread is charged on both legs, so it is a NET win rate.
- Intrabar ambiguity resolves to the STOP. OHLC cannot order two touches
inside one bar and the optimistic reading is how a backtested edge
becomes a live loss.
ZigZag stays as input features (EnableSwingContext) and now also supplies
the vertical barrier: the horizon is the median confirmed leg length,
snapped to a coarse ladder. Derived, not configured, and deliberately kept
out of the filename fingerprint - a filename keyed on a measured quantity
orphans a trained model the moment the measurement moves.
Removed, because the premise died with the old target:
- the alternation gate. Correct for pivot labels (a ZigZag cannot emit two
same-type pivots in a row, so a repeat was provably a false fire), and
wrong for barrier labels, which answer each bar independently. It also
took its worst consequence with it: a one-sided model previously got ONE
trade per backtest, a hard blocker on marketplace validation.
- SignalClusterWindow now defaults off - it de-duplicated repeats that are
now real trades. Kept as an opt-in display control.
- LABEL_WINDOW_BARS, the pivot-widening pass, ConfirmedZigZagLabel.
- the era-0 output-bias seed now needs a genuinely dominant class (0.70)
rather than 0.40; at ~50% Neutral a +-3.0 seed is a distortion, not a
correction.
Also fixed, both found while wiring the above:
1. RefreshConvergedSignal sized its buffers from a date delta
(Bars(sym, period, dtStudied, TimeCurrent())). dtStudied is a training
watermark; in the tester it is loaded from a live-chart save AHEAD of
the simulated date, so the interval inverted, Bars() returned ~0, and
the buffer came out at exactly m_historyBars - deep enough for the OHLC
window and far too shallow for the Donchian-50 / 20-bar-return / SMA
extension behind it. Inference silently computed DIFFERENT features
from the ones training learned on, live as well as in the tester. Now
sized from what the feature builder actually needs.
2. The barrier horizon is resolved on the deployed path too. A deployed
model never enters Train(), so it never reached the prebuild, and
OnlineLearnStep reads the horizon as its confirmation delay - left at
the fallback it would have backpropped bars whose barriers had not
resolved. Silent lookahead in the one place that writes to a live model.
SL_Mode/TP_Mode join the weights fingerprint: they define the labels now,
so a model trained at 1:3 must never be silently reused at 1:1. This
re-keys every pre-existing model by design - none were trained on this task.
Inference census extended with the vote gate. LongCondition/ShortCondition
open with a readiness check the refresh counters never see; in the tester it
reduces to "the seeded _optcache.nnw must have LOADED", and if it did not,
every vote is hard-zeroed while the model still answers Buy. The old three
counters would have read that as "the model says Neutral" - false, and a
completely different fix. This is the leading candidate for the
zero-direction backtest and the census can now name it in one run.
Both builds compile 0 errors / 0 warnings. Forces a full retrain.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 20:39:49 -04:00
|
|
|
int gaps[];
|
|
|
|
|
ArrayResize(gaps, 0);
|
|
|
|
|
int prevPivot = -1;
|
|
|
|
|
int scanned = 0;
|
|
|
|
|
//--- Oldest-to-newest is irrelevant here (a median has no order dependence), so scan newest-first from
|
|
|
|
|
//--- the first non-repainting bar and stop at the history edge.
|
|
|
|
|
for(int p = MathMax(m_swingConfirmationBars, 1); p < bars && scanned < SWING_SCAN_CAP_BARS * 4; p++, scanned++)
|
|
|
|
|
{
|
|
|
|
|
if(m_Open.GetData(p) == EMPTY_VALUE)
|
|
|
|
|
break;
|
|
|
|
|
if(m_ADZigZag.GetData(0, p) == 0.0)
|
|
|
|
|
continue;
|
|
|
|
|
if(prevPivot >= 0)
|
|
|
|
|
{
|
|
|
|
|
int gap = p - prevPivot;
|
|
|
|
|
if(gap > 0)
|
|
|
|
|
{
|
|
|
|
|
int n = ArraySize(gaps);
|
|
|
|
|
ArrayResize(gaps, n + 1);
|
|
|
|
|
gaps[n] = gap;
|
|
|
|
|
}
|
|
|
|
|
}
|
|
|
|
|
prevPivot = p;
|
|
|
|
|
}
|
|
|
|
|
int count = ArraySize(gaps);
|
fix: refuse invalid SL/TP, fix the unreachable deploy floor, scale the horizon
Three defects found by reading the 2026-08-01 training logs, all of which
only became visible because the relabel made the numbers mean something.
1. A STALE ENUM TRAINED FOUR MODELS ON THE WRONG TARGET.
`OnInit: trade settings snapshot - SL_Mode=1 TP_Mode=-101`
-101 was TP_PREV_SWING, deleted from TAKE_PROFIT_MODE on 2026-07-31 in
7eb48f5. MetaTrader does not validate a saved enum input against the
enum's current members, so charts saved before that kept the old
integer. BarrierMultiples()'s `if(tpMult <= 0.0) tpMult = slMult;`
then quietly turned it into a 1:1 barrier, and all four topologies
trained ~250 eras against a strategy nobody selected - while the log
reported "target 1.00*ATR" as though it were configured.
Since the relabel these two inputs ARE the label definition, so this
is not a bad trade setting, it is a wrong dataset. ValidateBarrier-
Inputs() now refuses to start (INIT_FAILED + Alert + an explicit fix)
on any value that is not an enum member. Members are enumerated rather
than range-checked because both enums are sparse and carry negative
sentinels, so no min/max test can tell a legal value from a deleted
one - which is the entire failure mode. The fallback survives as
belt-and-braces but now announces itself: a fallback that cannot say
it fired is indistinguishable from correct behaviour.
2. THE DEPLOYABILITY FLOOR BECAME MATHEMATICALLY UNREACHABLE.
`tradeableOK` required `dirPrecPct >= baseRatePct`, where baseRatePct
is Buy+Sell as a share of all bars. At the old exact-pivot target that
was ~6%, so "beat the base rate" read as "beat chance" and the test
looked sound. Triple-barrier labels put it at ~83%, so the gate now
demanded 83% directional precision - impossible by construction.
Observed live: all four topologies cycling "PLATEAU stage 3 ... nothing
safe to deploy" at a perfectly healthy 43-45% precision, with no
checkpoint able to ship however good it got.
Replaced with ZERO-SKILL precision, max(Buy,Sell)/allBars: exactly the
score of the degenerate always-call-one-direction model this floor
exists to reject. Correct at any base rate - ~43% on the current
labels, ~3% on the old rare-pivot ones. The era line now prints
"(chance N%, edge +Mpp)" beside the selection score, because 44%
precision is excellent against a 3% chance level and worthless against
a 43% one, and reading the first as the second is what made tonight's
run look better than it was.
3. THE HORIZON IGNORED THE BARRIER GEOMETRY.
ComputeBarrierHorizonBars() returned the median ZigZag leg, which
measures how long a ~1 ATR move takes and says nothing about how long
the CONFIGURED barrier needs. First-passage time out of [-m,+k] scales
with m*k, so a 1:3 barrier takes ~3x as long as 1:1; the unscaled
horizon would have timed out most 1:3 trades and pushed Neutral
straight back up, re-creating the imbalance the relabel removes.
Now multiplied by slMult*tpMult, calibrated against a real measurement
rather than assumed: the accidental 1:1 run resolved at horizon 12 with
only 16.7% timeouts, so the swing median is the right scale at m*k=1.
Verifiable, not just asserted: the prebuild now counts barriers that
ended on the VERTICAL barrier and reports them as a share of Neutral.
Neutral conflates "timed out" with "stopped out" and only the first
indicts the horizon.
Both builds compile 0 errors / 0 warnings. Forces a retrain - correcting
TP_Mode re-keys the fingerprint (|TB:1:-101 -> |TB:1:3), which is right:
no existing model was trained on the intended target.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 00:30:49 -04:00
|
|
|
double swingMedian = BARRIER_HORIZON_FALLBACK;
|
fix: a restart no longer loses the measured geometry or the training window
Terminal restart, 22:25: all four resumed models sat on empty windows with
enum 2:6 barriers. Three interlocking causes, all visible in one log excerpt:
1) THE PRE-SCAN WINDOW WAS SIZED BY THE SAVED WATERMARK. A resumed model's
dtStudied sits at its last studied bar, so Bars(dtStudied, now) ~ 0 and the
resumed-model MI pre-scan built a zero-bar "complete" label cache - logged as
"Buy: 0 | Sell: 0 | Neutral: 0". Train()'s own era start RESETS dtStudied to
the training-window rule before computing its window; the pre-scan did not.
The rule is now factored into TrainWindowStart() and both use it. The scan
also refuses to arm before SERIES_SYNCHRONIZED (it ran in the same second as
OnInit), and deployed models keep their watermark - for them it gates
inference recency, not a training window.
2) THE HORIZON LATCHED ON AN INDICATOR WARM-UP. ComputeBarrierHorizonBars ran
against a ZigZag with 0 calculated legs, fell back, and EnsureBarrierHorizon
latched fallback(32) x slMult x tpMult = 384 for the process lifetime. A
leg-starved horizon is now PROVISIONAL: re-resolved on the next rebuild, the
label cache wiped if it moved (labels from two horizons answer different
questions), and the geometry deriver refuses to run from it - a pair derived
over a warm-up window would get PINNED.
3) THE DERIVED GEOMETRY WAS NEVER PERSISTED. The .cfg is written at model
creation and at weights-reset - both BEFORE era 0 derives - so the measured
pair lived only in memory: every restart read back zeros, adopted nothing,
fell back to the enum barriers, and the era-0-only gate meant a resumed model
could NEVER re-derive. A full day of training on 3.33/1.62 resumed as 2:6.
Now: the settled pair is pinned to the .cfg the moment derivation completes
(one-shot, atomic write), and the derive gate accepts any model with no
pinned pair, not just era 0 - mid-run stability is carried by
m_geometryDerived itself, which never allows a second derivation.
Both build variants compile 0 errors, 0 warnings.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 22:40:43 -04:00
|
|
|
//--- Published so EnsureBarrierHorizon can refuse to LATCH a fallback: right after a terminal
|
|
|
|
|
//--- restart the ZigZag handle has calculated nothing yet, and a horizon computed from 0 legs is
|
|
|
|
|
//--- the indicator's warm-up state, not a property of the instrument.
|
|
|
|
|
m_barrierHorizonLegStarved = (count < BARRIER_HORIZON_MIN_SAMPLES);
|
|
|
|
|
if(!m_barrierHorizonLegStarved)
|
refactor: split CExpertSignalAIBase implementation by responsibility
ExpertSignalAIBase.mqh was 8216 lines: the class declaration followed by 87
method bodies covering training, labelling, feature extraction, persistence,
chart drawing, online learning, the GA auto-tuner and inference, all in one
file. Train() alone is 1492 lines; a change to arrow drawing meant scrolling
past the era loop.
Moved the bodies into Expert\AIBase\, included at the bottom of the original
after the class declaration:
Training.mqh 1607 era loop, plateau ladder, checkpoint select, deploy
Features.mqh 1093 indicator creation + per-bar input feature vector
ChartUI.mqh 634 arrows, arrow persistence, status panel, cleanup
Persistence.mqh 492 .stats/.cfg sidecars, CPU-inference validation, copy
OnlineLearning.mqh 461 live continual learning, EMA shadow, OOS simulator
Labels.mqh 309 ZigZag pivot labels, async label-cache prebuild
AutoTune.mqh 275 genetic tuner (population, crossover, halving)
Inference.mqh 235 softmax, prior calibration, class priors
ExpertSignalAIBase.mqh 8216 -> 3131 (declaration + topology build only)
This is a pure relocation - verified mechanically, not by eye: HEAD's file
reconstructed from the eight partials plus the surviving remainder is
byte-identical to HEAD, span for span (scratchpad verify_split.py). No
declaration moved, no signature changed, no code rewritten, so behaviour is
unchanged by construction.
Compiles 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 00:42:45 -04:00
|
|
|
{
|
feat(ai): triple-barrier labels replace exact-pivot ZigZag targets
The 31:1 class imbalance was self-inflicted by the TARGET, not a property
of the market. Labelling only the exact bar where a ZigZag pivot confirms
gave Buy 1164 / Sell 1164 / Neutral 35841, and every correction mechanism
this codebase accumulated sits downstream of that one choice: the
logit-adjusted loss and its range cap, the prior EMA, the +-3.0 output-bias
seed, balanced-accuracy-then-precision selection with its coverage floor,
the recall floor and its catch-22, the alternation gate, NMS, and the four
oversampling designs that collapsed before them.
The reference this engine is built on (references/neuronetworksbook.pdf
ch. 3.1/3.3) also uses ZigZag, but targets the DIRECTION TO THE NEXT
EXTREMUM on every bar - ~50/50 by construction, with no imbalance to
correct at all. It never had this problem because it never asked "is this
the pivot bar".
Labels are now the triple barrier (Lopez de Prado ch. 3), using the EA's
OWN SL_Mode/TP_Mode: does a trade opened at this bar's close reach its
target before its stop, within a horizon. Buy = long resolves, Sell =
short resolves, Neutral = neither. Consequences:
- dir-precision in the era line stops being a proxy and becomes the win
rate of the strategy under its own exit rules.
- Expected balance ~25/25/50 at the shipped 1:3 (gambler's ruin), i.e.
~2:1 instead of 31:1. Measured and logged at the end of the prebuild.
- Spread is charged on both legs, so it is a NET win rate.
- Intrabar ambiguity resolves to the STOP. OHLC cannot order two touches
inside one bar and the optimistic reading is how a backtested edge
becomes a live loss.
ZigZag stays as input features (EnableSwingContext) and now also supplies
the vertical barrier: the horizon is the median confirmed leg length,
snapped to a coarse ladder. Derived, not configured, and deliberately kept
out of the filename fingerprint - a filename keyed on a measured quantity
orphans a trained model the moment the measurement moves.
Removed, because the premise died with the old target:
- the alternation gate. Correct for pivot labels (a ZigZag cannot emit two
same-type pivots in a row, so a repeat was provably a false fire), and
wrong for barrier labels, which answer each bar independently. It also
took its worst consequence with it: a one-sided model previously got ONE
trade per backtest, a hard blocker on marketplace validation.
- SignalClusterWindow now defaults off - it de-duplicated repeats that are
now real trades. Kept as an opt-in display control.
- LABEL_WINDOW_BARS, the pivot-widening pass, ConfirmedZigZagLabel.
- the era-0 output-bias seed now needs a genuinely dominant class (0.70)
rather than 0.40; at ~50% Neutral a +-3.0 seed is a distortion, not a
correction.
Also fixed, both found while wiring the above:
1. RefreshConvergedSignal sized its buffers from a date delta
(Bars(sym, period, dtStudied, TimeCurrent())). dtStudied is a training
watermark; in the tester it is loaded from a live-chart save AHEAD of
the simulated date, so the interval inverted, Bars() returned ~0, and
the buffer came out at exactly m_historyBars - deep enough for the OHLC
window and far too shallow for the Donchian-50 / 20-bar-return / SMA
extension behind it. Inference silently computed DIFFERENT features
from the ones training learned on, live as well as in the tester. Now
sized from what the feature builder actually needs.
2. The barrier horizon is resolved on the deployed path too. A deployed
model never enters Train(), so it never reached the prebuild, and
OnlineLearnStep reads the horizon as its confirmation delay - left at
the fallback it would have backpropped bars whose barriers had not
resolved. Silent lookahead in the one place that writes to a live model.
SL_Mode/TP_Mode join the weights fingerprint: they define the labels now,
so a model trained at 1:3 must never be silently reused at 1:1. This
re-keys every pre-existing model by design - none were trained on this task.
Inference census extended with the vote gate. LongCondition/ShortCondition
open with a readiness check the refresh counters never see; in the tester it
reduces to "the seeded _optcache.nnw must have LOADED", and if it did not,
every vote is hard-zeroed while the model still answers Buy. The old three
counters would have read that as "the model says Neutral" - false, and a
completely different fix. This is the leading candidate for the
zero-direction backtest and the census can now name it in one run.
Both builds compile 0 errors / 0 warnings. Forces a full retrain.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 20:39:49 -04:00
|
|
|
ArraySort(gaps);
|
fix: refuse invalid SL/TP, fix the unreachable deploy floor, scale the horizon
Three defects found by reading the 2026-08-01 training logs, all of which
only became visible because the relabel made the numbers mean something.
1. A STALE ENUM TRAINED FOUR MODELS ON THE WRONG TARGET.
`OnInit: trade settings snapshot - SL_Mode=1 TP_Mode=-101`
-101 was TP_PREV_SWING, deleted from TAKE_PROFIT_MODE on 2026-07-31 in
7eb48f5. MetaTrader does not validate a saved enum input against the
enum's current members, so charts saved before that kept the old
integer. BarrierMultiples()'s `if(tpMult <= 0.0) tpMult = slMult;`
then quietly turned it into a 1:1 barrier, and all four topologies
trained ~250 eras against a strategy nobody selected - while the log
reported "target 1.00*ATR" as though it were configured.
Since the relabel these two inputs ARE the label definition, so this
is not a bad trade setting, it is a wrong dataset. ValidateBarrier-
Inputs() now refuses to start (INIT_FAILED + Alert + an explicit fix)
on any value that is not an enum member. Members are enumerated rather
than range-checked because both enums are sparse and carry negative
sentinels, so no min/max test can tell a legal value from a deleted
one - which is the entire failure mode. The fallback survives as
belt-and-braces but now announces itself: a fallback that cannot say
it fired is indistinguishable from correct behaviour.
2. THE DEPLOYABILITY FLOOR BECAME MATHEMATICALLY UNREACHABLE.
`tradeableOK` required `dirPrecPct >= baseRatePct`, where baseRatePct
is Buy+Sell as a share of all bars. At the old exact-pivot target that
was ~6%, so "beat the base rate" read as "beat chance" and the test
looked sound. Triple-barrier labels put it at ~83%, so the gate now
demanded 83% directional precision - impossible by construction.
Observed live: all four topologies cycling "PLATEAU stage 3 ... nothing
safe to deploy" at a perfectly healthy 43-45% precision, with no
checkpoint able to ship however good it got.
Replaced with ZERO-SKILL precision, max(Buy,Sell)/allBars: exactly the
score of the degenerate always-call-one-direction model this floor
exists to reject. Correct at any base rate - ~43% on the current
labels, ~3% on the old rare-pivot ones. The era line now prints
"(chance N%, edge +Mpp)" beside the selection score, because 44%
precision is excellent against a 3% chance level and worthless against
a 43% one, and reading the first as the second is what made tonight's
run look better than it was.
3. THE HORIZON IGNORED THE BARRIER GEOMETRY.
ComputeBarrierHorizonBars() returned the median ZigZag leg, which
measures how long a ~1 ATR move takes and says nothing about how long
the CONFIGURED barrier needs. First-passage time out of [-m,+k] scales
with m*k, so a 1:3 barrier takes ~3x as long as 1:1; the unscaled
horizon would have timed out most 1:3 trades and pushed Neutral
straight back up, re-creating the imbalance the relabel removes.
Now multiplied by slMult*tpMult, calibrated against a real measurement
rather than assumed: the accidental 1:1 run resolved at horizon 12 with
only 16.7% timeouts, so the swing median is the right scale at m*k=1.
Verifiable, not just asserted: the prebuild now counts barriers that
ended on the VERTICAL barrier and reports them as a share of Neutral.
Neutral conflates "timed out" with "stopped out" and only the first
indicts the horizon.
Both builds compile 0 errors / 0 warnings. Forces a retrain - correcting
TP_Mode re-keys the fingerprint (|TB:1:-101 -> |TB:1:3), which is right:
no existing model was trained on the intended target.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 00:30:49 -04:00
|
|
|
swingMedian = gaps[count / 2];
|
refactor: split CExpertSignalAIBase implementation by responsibility
ExpertSignalAIBase.mqh was 8216 lines: the class declaration followed by 87
method bodies covering training, labelling, feature extraction, persistence,
chart drawing, online learning, the GA auto-tuner and inference, all in one
file. Train() alone is 1492 lines; a change to arrow drawing meant scrolling
past the era loop.
Moved the bodies into Expert\AIBase\, included at the bottom of the original
after the class declaration:
Training.mqh 1607 era loop, plateau ladder, checkpoint select, deploy
Features.mqh 1093 indicator creation + per-bar input feature vector
ChartUI.mqh 634 arrows, arrow persistence, status panel, cleanup
Persistence.mqh 492 .stats/.cfg sidecars, CPU-inference validation, copy
OnlineLearning.mqh 461 live continual learning, EMA shadow, OOS simulator
Labels.mqh 309 ZigZag pivot labels, async label-cache prebuild
AutoTune.mqh 275 genetic tuner (population, crossover, halving)
Inference.mqh 235 softmax, prior calibration, class priors
ExpertSignalAIBase.mqh 8216 -> 3131 (declaration + topology build only)
This is a pure relocation - verified mechanically, not by eye: HEAD's file
reconstructed from the eight partials plus the surviving remainder is
byte-identical to HEAD, span for span (scratchpad verify_split.py). No
declaration moved, no signature changed, no code rewritten, so behaviour is
unchanged by construction.
Compiles 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 00:42:45 -04:00
|
|
|
}
|
feat(ai): triple-barrier labels replace exact-pivot ZigZag targets
The 31:1 class imbalance was self-inflicted by the TARGET, not a property
of the market. Labelling only the exact bar where a ZigZag pivot confirms
gave Buy 1164 / Sell 1164 / Neutral 35841, and every correction mechanism
this codebase accumulated sits downstream of that one choice: the
logit-adjusted loss and its range cap, the prior EMA, the +-3.0 output-bias
seed, balanced-accuracy-then-precision selection with its coverage floor,
the recall floor and its catch-22, the alternation gate, NMS, and the four
oversampling designs that collapsed before them.
The reference this engine is built on (references/neuronetworksbook.pdf
ch. 3.1/3.3) also uses ZigZag, but targets the DIRECTION TO THE NEXT
EXTREMUM on every bar - ~50/50 by construction, with no imbalance to
correct at all. It never had this problem because it never asked "is this
the pivot bar".
Labels are now the triple barrier (Lopez de Prado ch. 3), using the EA's
OWN SL_Mode/TP_Mode: does a trade opened at this bar's close reach its
target before its stop, within a horizon. Buy = long resolves, Sell =
short resolves, Neutral = neither. Consequences:
- dir-precision in the era line stops being a proxy and becomes the win
rate of the strategy under its own exit rules.
- Expected balance ~25/25/50 at the shipped 1:3 (gambler's ruin), i.e.
~2:1 instead of 31:1. Measured and logged at the end of the prebuild.
- Spread is charged on both legs, so it is a NET win rate.
- Intrabar ambiguity resolves to the STOP. OHLC cannot order two touches
inside one bar and the optimistic reading is how a backtested edge
becomes a live loss.
ZigZag stays as input features (EnableSwingContext) and now also supplies
the vertical barrier: the horizon is the median confirmed leg length,
snapped to a coarse ladder. Derived, not configured, and deliberately kept
out of the filename fingerprint - a filename keyed on a measured quantity
orphans a trained model the moment the measurement moves.
Removed, because the premise died with the old target:
- the alternation gate. Correct for pivot labels (a ZigZag cannot emit two
same-type pivots in a row, so a repeat was provably a false fire), and
wrong for barrier labels, which answer each bar independently. It also
took its worst consequence with it: a one-sided model previously got ONE
trade per backtest, a hard blocker on marketplace validation.
- SignalClusterWindow now defaults off - it de-duplicated repeats that are
now real trades. Kept as an opt-in display control.
- LABEL_WINDOW_BARS, the pivot-widening pass, ConfirmedZigZagLabel.
- the era-0 output-bias seed now needs a genuinely dominant class (0.70)
rather than 0.40; at ~50% Neutral a +-3.0 seed is a distortion, not a
correction.
Also fixed, both found while wiring the above:
1. RefreshConvergedSignal sized its buffers from a date delta
(Bars(sym, period, dtStudied, TimeCurrent())). dtStudied is a training
watermark; in the tester it is loaded from a live-chart save AHEAD of
the simulated date, so the interval inverted, Bars() returned ~0, and
the buffer came out at exactly m_historyBars - deep enough for the OHLC
window and far too shallow for the Donchian-50 / 20-bar-return / SMA
extension behind it. Inference silently computed DIFFERENT features
from the ones training learned on, live as well as in the tester. Now
sized from what the feature builder actually needs.
2. The barrier horizon is resolved on the deployed path too. A deployed
model never enters Train(), so it never reached the prebuild, and
OnlineLearnStep reads the horizon as its confirmation delay - left at
the fallback it would have backpropped bars whose barriers had not
resolved. Silent lookahead in the one place that writes to a live model.
SL_Mode/TP_Mode join the weights fingerprint: they define the labels now,
so a model trained at 1:3 must never be silently reused at 1:1. This
re-keys every pre-existing model by design - none were trained on this task.
Inference census extended with the vote gate. LongCondition/ShortCondition
open with a readiness check the refresh counters never see; in the tester it
reduces to "the seeded _optcache.nnw must have LOADED", and if it did not,
every vote is hard-zeroed while the model still answers Buy. The old three
counters would have read that as "the model says Neutral" - false, and a
completely different fix. This is the leading candidate for the
zero-direction backtest and the census can now name it in one run.
Both builds compile 0 errors / 0 warnings. Forces a full retrain.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 20:39:49 -04:00
|
|
|
else
|
fix: a restart no longer loses the measured geometry or the training window
Terminal restart, 22:25: all four resumed models sat on empty windows with
enum 2:6 barriers. Three interlocking causes, all visible in one log excerpt:
1) THE PRE-SCAN WINDOW WAS SIZED BY THE SAVED WATERMARK. A resumed model's
dtStudied sits at its last studied bar, so Bars(dtStudied, now) ~ 0 and the
resumed-model MI pre-scan built a zero-bar "complete" label cache - logged as
"Buy: 0 | Sell: 0 | Neutral: 0". Train()'s own era start RESETS dtStudied to
the training-window rule before computing its window; the pre-scan did not.
The rule is now factored into TrainWindowStart() and both use it. The scan
also refuses to arm before SERIES_SYNCHRONIZED (it ran in the same second as
OnInit), and deployed models keep their watermark - for them it gates
inference recency, not a training window.
2) THE HORIZON LATCHED ON AN INDICATOR WARM-UP. ComputeBarrierHorizonBars ran
against a ZigZag with 0 calculated legs, fell back, and EnsureBarrierHorizon
latched fallback(32) x slMult x tpMult = 384 for the process lifetime. A
leg-starved horizon is now PROVISIONAL: re-resolved on the next rebuild, the
label cache wiped if it moved (labels from two horizons answer different
questions), and the geometry deriver refuses to run from it - a pair derived
over a warm-up window would get PINNED.
3) THE DERIVED GEOMETRY WAS NEVER PERSISTED. The .cfg is written at model
creation and at weights-reset - both BEFORE era 0 derives - so the measured
pair lived only in memory: every restart read back zeros, adopted nothing,
fell back to the enum barriers, and the era-0-only gate meant a resumed model
could NEVER re-derive. A full day of training on 3.33/1.62 resumed as 2:6.
Now: the settled pair is pinned to the .cfg the moment derivation completes
(one-shot, atomic write), and the derive gate accepts any model with no
pinned pair, not just era 0 - mid-run stability is carried by
m_geometryDerived itself, which never allows a second derivation.
Both build variants compile 0 errors, 0 warnings.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 22:40:43 -04:00
|
|
|
if(!m_horizonStarvedWarned)
|
|
|
|
|
{
|
|
|
|
|
m_horizonStarvedWarned = true;
|
|
|
|
|
Print(ID + ": barrier horizon - only " + IntegerToString(count) + " confirmed ZigZag legs available (need " +
|
|
|
|
|
IntegerToString(BARRIER_HORIZON_MIN_SAMPLES) + "), falling back to " +
|
|
|
|
|
IntegerToString(BARRIER_HORIZON_FALLBACK) + " bars PROVISIONALLY - re-resolved on the "
|
|
|
|
|
"next label-cache rebuild, once the indicator has caught up");
|
|
|
|
|
}
|
fix: refuse invalid SL/TP, fix the unreachable deploy floor, scale the horizon
Three defects found by reading the 2026-08-01 training logs, all of which
only became visible because the relabel made the numbers mean something.
1. A STALE ENUM TRAINED FOUR MODELS ON THE WRONG TARGET.
`OnInit: trade settings snapshot - SL_Mode=1 TP_Mode=-101`
-101 was TP_PREV_SWING, deleted from TAKE_PROFIT_MODE on 2026-07-31 in
7eb48f5. MetaTrader does not validate a saved enum input against the
enum's current members, so charts saved before that kept the old
integer. BarrierMultiples()'s `if(tpMult <= 0.0) tpMult = slMult;`
then quietly turned it into a 1:1 barrier, and all four topologies
trained ~250 eras against a strategy nobody selected - while the log
reported "target 1.00*ATR" as though it were configured.
Since the relabel these two inputs ARE the label definition, so this
is not a bad trade setting, it is a wrong dataset. ValidateBarrier-
Inputs() now refuses to start (INIT_FAILED + Alert + an explicit fix)
on any value that is not an enum member. Members are enumerated rather
than range-checked because both enums are sparse and carry negative
sentinels, so no min/max test can tell a legal value from a deleted
one - which is the entire failure mode. The fallback survives as
belt-and-braces but now announces itself: a fallback that cannot say
it fired is indistinguishable from correct behaviour.
2. THE DEPLOYABILITY FLOOR BECAME MATHEMATICALLY UNREACHABLE.
`tradeableOK` required `dirPrecPct >= baseRatePct`, where baseRatePct
is Buy+Sell as a share of all bars. At the old exact-pivot target that
was ~6%, so "beat the base rate" read as "beat chance" and the test
looked sound. Triple-barrier labels put it at ~83%, so the gate now
demanded 83% directional precision - impossible by construction.
Observed live: all four topologies cycling "PLATEAU stage 3 ... nothing
safe to deploy" at a perfectly healthy 43-45% precision, with no
checkpoint able to ship however good it got.
Replaced with ZERO-SKILL precision, max(Buy,Sell)/allBars: exactly the
score of the degenerate always-call-one-direction model this floor
exists to reject. Correct at any base rate - ~43% on the current
labels, ~3% on the old rare-pivot ones. The era line now prints
"(chance N%, edge +Mpp)" beside the selection score, because 44%
precision is excellent against a 3% chance level and worthless against
a 43% one, and reading the first as the second is what made tonight's
run look better than it was.
3. THE HORIZON IGNORED THE BARRIER GEOMETRY.
ComputeBarrierHorizonBars() returned the median ZigZag leg, which
measures how long a ~1 ATR move takes and says nothing about how long
the CONFIGURED barrier needs. First-passage time out of [-m,+k] scales
with m*k, so a 1:3 barrier takes ~3x as long as 1:1; the unscaled
horizon would have timed out most 1:3 trades and pushed Neutral
straight back up, re-creating the imbalance the relabel removes.
Now multiplied by slMult*tpMult, calibrated against a real measurement
rather than assumed: the accidental 1:1 run resolved at horizon 12 with
only 16.7% timeouts, so the swing median is the right scale at m*k=1.
Verifiable, not just asserted: the prebuild now counts barriers that
ended on the VERTICAL barrier and reports them as a share of Neutral.
Neutral conflates "timed out" with "stopped out" and only the first
indicts the horizon.
Both builds compile 0 errors / 0 warnings. Forces a retrain - correcting
TP_Mode re-keys the fingerprint (|TB:1:-101 -> |TB:1:3), which is right:
no existing model was trained on the intended target.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 00:30:49 -04:00
|
|
|
//--- SCALE BY THE BARRIER GEOMETRY. The swing median alone measures how long a ~1 ATR move takes on
|
|
|
|
|
//--- this instrument; it says nothing about how long the CONFIGURED barrier takes to resolve, and the
|
|
|
|
|
//--- first version of this function ignored that entirely.
|
|
|
|
|
//--- For a driftless random walk leaving the band [-m*ATR, +k*ATR], the expected first-passage time is
|
|
|
|
|
//--- proportional to m*k. So a 1:3 barrier takes ~3x as long to resolve as a 1:1 one, and a horizon
|
|
|
|
|
//--- tuned for 1:1 applied to 1:3 would time out most trades - pushing Neutral straight back up and
|
|
|
|
|
//--- re-creating the imbalance the relabel exists to remove.
|
|
|
|
|
//--- Calibrated against a real measurement rather than assumed: the 2026-08-01 run resolved at m=k=1
|
|
|
|
|
//--- with a 12-bar horizon and only 16.7% timeouts, so the swing median IS the right scale at m*k=1.
|
|
|
|
|
//--- Multiplying by m*k carries that calibration to every other barrier (1:3 -> 36, snapping to 32).
|
|
|
|
|
double slMult, tpMult;
|
|
|
|
|
BarrierMultiples(slMult, tpMult);
|
fix: excursion window must not depend on the barrier it sizes
DIRECTION IS NOT THERE, and this run is what establishes it. Three symbols:
raw ASYMMETRY clears on all three (p=0.0199 / 0.0050 / 0.0050)
norm ASYMMETRY collapses on all three (p=0.3433 / 0.5075 / 0.2736),
USDCAD landing BELOW its own null
RANGE control strengthens to 3-5x its null everywhere
Divide sigma out and the apparent directional signal vanishes entirely. What
cleared was volatility leaking through an unnormalised difference. Note this
would have passed any replication test: three instruments at p=0.005 is exactly
the evidence one would accept before committing to a rebuild, and the confound
reproduces perfectly. Replication was never going to catch it - only the
normalisation could.
Two defects of mine, both surfaced by the same run.
1. THE GEOMETRY DERIVATION WAS DIVERGING, NOT CONVERGING. It produced a
14.57*ATR stop and a 29.14*ATR target that only 5.7% of bars ever reach.
Excursions were measured over the barrier horizon; the horizon scales with
the target; the target is a quantile of the excursions - so target ->
horizon -> excursions -> target ran away, and "settled" only because the
horizon ladder caps at 384 bars. A saturated runaway, which the iteration
guard could not catch because it watches for OSCILLATION.
Fixed at the root: excursions now accumulate only over m_swingMedianBars -
the UNSCALED median ZigZag leg, a property of the instrument that owes
nothing to the barrier. The barrier walk still runs the full horizon,
because that is how long the trade is held; only the MEASUREMENT used to
size the barrier is confined to a geometry-independent window.
(The Min_Risk_Reward_Ratio warning fired correctly and is what flagged it -
the diagnostic worked while the derivation behind it did not.)
2. THE CONFOUND VERDICT WAS UNREACHABLE. `sizeCleared && !asymCleared` was
tested first and is true whenever size clears - i.e. always - so the branch
that NAMES the volatility confound never printed; all three symbols showed
the generic size-not-direction message instead. Verdict chain rewritten with
the specific case first, and the dangling elses my first patch introduced
removed.
FORCES A FULL RETRAIN (the excursion window changes every derived barrier).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-07 13:57:23 -04:00
|
|
|
//--- THE EXCURSION REFERENCE WINDOW, published UNSCALED. This is a property of the instrument (how
|
|
|
|
|
//--- long its typical swing leg lasts) and owes nothing to the barrier, which is exactly what makes it
|
|
|
|
|
//--- usable for sizing the barrier. Sizing a stop off travel measured over the SCALED horizon below
|
|
|
|
|
//--- is circular: horizon grows with the target, excursions grow with the horizon, the target is a
|
|
|
|
|
//--- quantile of the excursions - so target -> horizon -> excursions -> target diverges. Measured
|
|
|
|
|
//--- 2026-08-07 on EURUSD/USDCAD: it ran away to a 14-15*ATR stop and a 29-31*ATR target that only
|
|
|
|
|
//--- 5.7-7.2% of bars ever reached, and "converged" solely because the ladder caps at 384 bars. A
|
|
|
|
|
//--- saturated runaway, not a fixed point - which is why the iteration guard, watching for
|
|
|
|
|
//--- oscillation, did not catch it.
|
|
|
|
|
m_swingMedianBars = (int)MathMax(MathRound(swingMedian), 1);
|
fix: refuse invalid SL/TP, fix the unreachable deploy floor, scale the horizon
Three defects found by reading the 2026-08-01 training logs, all of which
only became visible because the relabel made the numbers mean something.
1. A STALE ENUM TRAINED FOUR MODELS ON THE WRONG TARGET.
`OnInit: trade settings snapshot - SL_Mode=1 TP_Mode=-101`
-101 was TP_PREV_SWING, deleted from TAKE_PROFIT_MODE on 2026-07-31 in
7eb48f5. MetaTrader does not validate a saved enum input against the
enum's current members, so charts saved before that kept the old
integer. BarrierMultiples()'s `if(tpMult <= 0.0) tpMult = slMult;`
then quietly turned it into a 1:1 barrier, and all four topologies
trained ~250 eras against a strategy nobody selected - while the log
reported "target 1.00*ATR" as though it were configured.
Since the relabel these two inputs ARE the label definition, so this
is not a bad trade setting, it is a wrong dataset. ValidateBarrier-
Inputs() now refuses to start (INIT_FAILED + Alert + an explicit fix)
on any value that is not an enum member. Members are enumerated rather
than range-checked because both enums are sparse and carry negative
sentinels, so no min/max test can tell a legal value from a deleted
one - which is the entire failure mode. The fallback survives as
belt-and-braces but now announces itself: a fallback that cannot say
it fired is indistinguishable from correct behaviour.
2. THE DEPLOYABILITY FLOOR BECAME MATHEMATICALLY UNREACHABLE.
`tradeableOK` required `dirPrecPct >= baseRatePct`, where baseRatePct
is Buy+Sell as a share of all bars. At the old exact-pivot target that
was ~6%, so "beat the base rate" read as "beat chance" and the test
looked sound. Triple-barrier labels put it at ~83%, so the gate now
demanded 83% directional precision - impossible by construction.
Observed live: all four topologies cycling "PLATEAU stage 3 ... nothing
safe to deploy" at a perfectly healthy 43-45% precision, with no
checkpoint able to ship however good it got.
Replaced with ZERO-SKILL precision, max(Buy,Sell)/allBars: exactly the
score of the degenerate always-call-one-direction model this floor
exists to reject. Correct at any base rate - ~43% on the current
labels, ~3% on the old rare-pivot ones. The era line now prints
"(chance N%, edge +Mpp)" beside the selection score, because 44%
precision is excellent against a 3% chance level and worthless against
a 43% one, and reading the first as the second is what made tonight's
run look better than it was.
3. THE HORIZON IGNORED THE BARRIER GEOMETRY.
ComputeBarrierHorizonBars() returned the median ZigZag leg, which
measures how long a ~1 ATR move takes and says nothing about how long
the CONFIGURED barrier needs. First-passage time out of [-m,+k] scales
with m*k, so a 1:3 barrier takes ~3x as long as 1:1; the unscaled
horizon would have timed out most 1:3 trades and pushed Neutral
straight back up, re-creating the imbalance the relabel removes.
Now multiplied by slMult*tpMult, calibrated against a real measurement
rather than assumed: the accidental 1:1 run resolved at horizon 12 with
only 16.7% timeouts, so the swing median is the right scale at m*k=1.
Verifiable, not just asserted: the prebuild now counts barriers that
ended on the VERTICAL barrier and reports them as a share of Neutral.
Neutral conflates "timed out" with "stopped out" and only the first
indicts the horizon.
Both builds compile 0 errors / 0 warnings. Forces a retrain - correcting
TP_Mode re-keys the fingerprint (|TB:1:-101 -> |TB:1:3), which is right:
no existing model was trained on the intended target.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 00:30:49 -04:00
|
|
|
int raw = (int)MathRound(swingMedian * slMult * tpMult);
|
fix(labels): the geometry scan rewarded the labels it should reject
First run named 3:10 on all four charts, at 2.3x the configured 2:6. That
answer was wrong and the fault was the ranking statistic.
3:10 wants a horizon of ~swingMedian*30 (~320 bars) and gets
BARRIER_HORIZON_MAX. Clamped, most trades never resolve, the unresolved
remainder all lands in Neutral, and H(Y) collapses. The old statistic
divided the excess BY H(Y) - so a collapsing denominator made the most
degenerate label look like the most predictable one. Every geometry from
2:6 upward was already showing the clamped h128, and the two widest
scored highest, which is the fingerprint of the artefact rather than of
signal.
Two fixes:
Rank on the raw excess in nats. Subtracting each geometry's OWN measured
null already removes the class-balance bias, which is the only thing the
normalisation was ever needed for.
Disqualify clamped geometries outright rather than ranking them down. The
deployed EA holds until SL or TP with no bar limit, so a truncated label
trains the model on a question the strategy never asks. They are still
printed, marked '!', so the disqualification is visible instead of a
silent omission - and the scan now says so explicitly when nothing
eligible is left, because "the limit is the feature set, not the target"
is itself the finding in that case.
The scan also reports each geometry's directional share and timeout share
now. A label nobody can trade is not a candidate however well it scores,
and that has to be visible in the same line as the score.
Compiles 0 errors / 0 warnings. Build tag geometry-scan-v2.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 14:46:25 -04:00
|
|
|
//--- CLAMPED means the barrier this geometry describes needs MORE time than the ceiling allows, so the
|
|
|
|
|
//--- label stops being "does the target come before the stop" and quietly becomes "does the target come
|
|
|
|
|
//--- within BARRIER_HORIZON_MAX bars". The deployed EA has no such bar limit - it holds until SL or TP -
|
|
|
|
|
//--- so a clamped label trains the model on a question the strategy never asks, and the unresolved
|
|
|
|
|
//--- remainder all lands in Neutral. Recorded rather than merely clamped because the geometry scan must
|
|
|
|
|
//--- be able to disqualify these: they LOOK informative precisely because a Neutral-dominated label has
|
|
|
|
|
//--- little entropy left to explain.
|
|
|
|
|
m_barrierHorizonClamped = (raw > BARRIER_HORIZON_MAX);
|
fix(labels): correct EffectiveSampleSize clamp order, share the horizon ladder, retract a false justification
Self-review of 1540ba8 against the FULL 6,930-era log rather than the first
three minutes of it. Three corrections.
1. EffectiveSampleSize() clamped in the wrong order. MathMax(2, MathMin(eff,
rawN)) returns 2 when rawN is 1 - an effective sample LARGER than the raw
one, shrinking the SE in exactly the direction the function exists to
prevent. Floor first, cap at rawN last.
2. The horizon cap rejected on the CEILING only, and said so as though that
made the label untruncated. It does not: the horizon ladder also snaps DOWN,
so a pair needing 317 bars is granted 256 and is silently truncated without
ever being flagged CLAMPED. Added SnapHorizonToLadder() / GrantedHorizonBars()
and the scale ladder now reports "needs N gets M" per rung. Rejection stays on
the ceiling alone - matching ReportGeometryExpectancyScan's '!' exactly, which
was the point - because rejecting on the snap-down would select rungs for
landing just above a ladder point rather than for anything about the market.
ComputeBarrierHorizonBars' private copy of the ladder is gone; there is now
one copy, which is the whole reason RequiredHorizonBars was factored out.
3. RETRACTED THE JUSTIFICATION IN 1540ba8's COMMENTS. That commit claimed the
overlap correction was needed because the operating point's null-of-the-
maximum gate fired on 47/73 Perceptron eras (64%) where a family-wise test
should fire on ~5%. Those 73 fits were the first three minutes of a
six-and-a-half-hour run. Over the full run:
PAI 47/3214 = 1.5% HYB 30/1200 = 2.5%
CONV 4/63 = 6.3% LSTM 75/915 = 8.2%
All at or below the null. The gate from 7414570 is working as designed and
PAI's 47 clears were a cold-start transient never repeated in 3,141 later
fits; its threshold over the run's second half has sd 0.01. The overlap
correction is still right - sqrt(p(1-p)/n) on overlapping labels is the wrong
formula - but it fixes no observed failure, and it costs nothing today
because no model is near the deploy line.
WHAT THE FULL RUN DOES CONFIRM, unchanged: the geometry ran away exactly as
described (2.00/6.00 h128 -> 3.49/6.99 h256 -> 4.86/9.71 h384, three passes,
stopping at q90 because the quantile ladder ended), the label stayed long-skewed
at Buy 42.9% / Sell 22.6%, and no checkpoint on any of the four models ever
cleared the deployability floor. Pooled declustered win rates: PAI 31.70%,
HYB 31.02%, LSTM 32.69%, CONV 31.57% - every one 4-6pp below the 37% always-long
chance rate and 1-2.7pp below the 33.7% cost-adjusted break-even.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 07:30:19 -04:00
|
|
|
//--- Clamp + snap live in SnapHorizonToLadder() so the scale ladder can ask the same question about a
|
|
|
|
|
//--- candidate rung without a second copy of the ladder - see its header.
|
|
|
|
|
return SnapHorizonToLadder(raw);
|
feat(ai): triple-barrier labels replace exact-pivot ZigZag targets
The 31:1 class imbalance was self-inflicted by the TARGET, not a property
of the market. Labelling only the exact bar where a ZigZag pivot confirms
gave Buy 1164 / Sell 1164 / Neutral 35841, and every correction mechanism
this codebase accumulated sits downstream of that one choice: the
logit-adjusted loss and its range cap, the prior EMA, the +-3.0 output-bias
seed, balanced-accuracy-then-precision selection with its coverage floor,
the recall floor and its catch-22, the alternation gate, NMS, and the four
oversampling designs that collapsed before them.
The reference this engine is built on (references/neuronetworksbook.pdf
ch. 3.1/3.3) also uses ZigZag, but targets the DIRECTION TO THE NEXT
EXTREMUM on every bar - ~50/50 by construction, with no imbalance to
correct at all. It never had this problem because it never asked "is this
the pivot bar".
Labels are now the triple barrier (Lopez de Prado ch. 3), using the EA's
OWN SL_Mode/TP_Mode: does a trade opened at this bar's close reach its
target before its stop, within a horizon. Buy = long resolves, Sell =
short resolves, Neutral = neither. Consequences:
- dir-precision in the era line stops being a proxy and becomes the win
rate of the strategy under its own exit rules.
- Expected balance ~25/25/50 at the shipped 1:3 (gambler's ruin), i.e.
~2:1 instead of 31:1. Measured and logged at the end of the prebuild.
- Spread is charged on both legs, so it is a NET win rate.
- Intrabar ambiguity resolves to the STOP. OHLC cannot order two touches
inside one bar and the optimistic reading is how a backtested edge
becomes a live loss.
ZigZag stays as input features (EnableSwingContext) and now also supplies
the vertical barrier: the horizon is the median confirmed leg length,
snapped to a coarse ladder. Derived, not configured, and deliberately kept
out of the filename fingerprint - a filename keyed on a measured quantity
orphans a trained model the moment the measurement moves.
Removed, because the premise died with the old target:
- the alternation gate. Correct for pivot labels (a ZigZag cannot emit two
same-type pivots in a row, so a repeat was provably a false fire), and
wrong for barrier labels, which answer each bar independently. It also
took its worst consequence with it: a one-sided model previously got ONE
trade per backtest, a hard blocker on marketplace validation.
- SignalClusterWindow now defaults off - it de-duplicated repeats that are
now real trades. Kept as an opt-in display control.
- LABEL_WINDOW_BARS, the pivot-widening pass, ConfirmedZigZagLabel.
- the era-0 output-bias seed now needs a genuinely dominant class (0.70)
rather than 0.40; at ~50% Neutral a +-3.0 seed is a distortion, not a
correction.
Also fixed, both found while wiring the above:
1. RefreshConvergedSignal sized its buffers from a date delta
(Bars(sym, period, dtStudied, TimeCurrent())). dtStudied is a training
watermark; in the tester it is loaded from a live-chart save AHEAD of
the simulated date, so the interval inverted, Bars() returned ~0, and
the buffer came out at exactly m_historyBars - deep enough for the OHLC
window and far too shallow for the Donchian-50 / 20-bar-return / SMA
extension behind it. Inference silently computed DIFFERENT features
from the ones training learned on, live as well as in the tester. Now
sized from what the feature builder actually needs.
2. The barrier horizon is resolved on the deployed path too. A deployed
model never enters Train(), so it never reached the prebuild, and
OnlineLearnStep reads the horizon as its confirmation delay - left at
the fallback it would have backpropped bars whose barriers had not
resolved. Silent lookahead in the one place that writes to a live model.
SL_Mode/TP_Mode join the weights fingerprint: they define the labels now,
so a model trained at 1:3 must never be silently reused at 1:1. This
re-keys every pre-existing model by design - none were trained on this task.
Inference census extended with the vote gate. LongCondition/ShortCondition
open with a readiness check the refresh counters never see; in the tester it
reduces to "the seeded _optcache.nnw must have LOADED", and if it did not,
every vote is hard-zeroed while the model still answers Buy. The old three
counters would have read that as "the model says Neutral" - false, and a
completely different fix. This is the leading candidate for the
zero-direction backtest and the census can now name it in one run.
Both builds compile 0 errors / 0 warnings. Forces a full retrain.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 20:39:49 -04:00
|
|
|
}
|
|
|
|
|
//+------------------------------------------------------------------+
|
|
|
|
|
//| Resolves m_barrierHorizonBars once per process and logs the whole |
|
|
|
|
|
//| label definition. Called from BOTH the training prebuild and the |
|
|
|
|
|
//| deployed inference path - see the declaration for why a deployed |
|
|
|
|
|
//| model that skipped this would silently learn online from bars |
|
|
|
|
|
//| whose barriers had not resolved. |
|
|
|
|
|
//+------------------------------------------------------------------+
|
|
|
|
|
void CExpertSignalAIBase::EnsureBarrierHorizon(int bars)
|
|
|
|
|
{
|
|
|
|
|
if(m_barrierHorizonResolved)
|
|
|
|
|
return;
|
fix: a restart no longer loses the measured geometry or the training window
Terminal restart, 22:25: all four resumed models sat on empty windows with
enum 2:6 barriers. Three interlocking causes, all visible in one log excerpt:
1) THE PRE-SCAN WINDOW WAS SIZED BY THE SAVED WATERMARK. A resumed model's
dtStudied sits at its last studied bar, so Bars(dtStudied, now) ~ 0 and the
resumed-model MI pre-scan built a zero-bar "complete" label cache - logged as
"Buy: 0 | Sell: 0 | Neutral: 0". Train()'s own era start RESETS dtStudied to
the training-window rule before computing its window; the pre-scan did not.
The rule is now factored into TrainWindowStart() and both use it. The scan
also refuses to arm before SERIES_SYNCHRONIZED (it ran in the same second as
OnInit), and deployed models keep their watermark - for them it gates
inference recency, not a training window.
2) THE HORIZON LATCHED ON AN INDICATOR WARM-UP. ComputeBarrierHorizonBars ran
against a ZigZag with 0 calculated legs, fell back, and EnsureBarrierHorizon
latched fallback(32) x slMult x tpMult = 384 for the process lifetime. A
leg-starved horizon is now PROVISIONAL: re-resolved on the next rebuild, the
label cache wiped if it moved (labels from two horizons answer different
questions), and the geometry deriver refuses to run from it - a pair derived
over a warm-up window would get PINNED.
3) THE DERIVED GEOMETRY WAS NEVER PERSISTED. The .cfg is written at model
creation and at weights-reset - both BEFORE era 0 derives - so the measured
pair lived only in memory: every restart read back zeros, adopted nothing,
fell back to the enum barriers, and the era-0-only gate meant a resumed model
could NEVER re-derive. A full day of training on 3.33/1.62 resumed as 2:6.
Now: the settled pair is pinned to the .cfg the moment derivation completes
(one-shot, atomic write), and the derive gate accepts any model with no
pinned pair, not just era 0 - mid-run stability is carried by
m_geometryDerived itself, which never allows a second derivation.
Both build variants compile 0 errors, 0 warnings.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 22:40:43 -04:00
|
|
|
int prevHorizon = m_barrierHorizonBars;
|
feat(ai): triple-barrier labels replace exact-pivot ZigZag targets
The 31:1 class imbalance was self-inflicted by the TARGET, not a property
of the market. Labelling only the exact bar where a ZigZag pivot confirms
gave Buy 1164 / Sell 1164 / Neutral 35841, and every correction mechanism
this codebase accumulated sits downstream of that one choice: the
logit-adjusted loss and its range cap, the prior EMA, the +-3.0 output-bias
seed, balanced-accuracy-then-precision selection with its coverage floor,
the recall floor and its catch-22, the alternation gate, NMS, and the four
oversampling designs that collapsed before them.
The reference this engine is built on (references/neuronetworksbook.pdf
ch. 3.1/3.3) also uses ZigZag, but targets the DIRECTION TO THE NEXT
EXTREMUM on every bar - ~50/50 by construction, with no imbalance to
correct at all. It never had this problem because it never asked "is this
the pivot bar".
Labels are now the triple barrier (Lopez de Prado ch. 3), using the EA's
OWN SL_Mode/TP_Mode: does a trade opened at this bar's close reach its
target before its stop, within a horizon. Buy = long resolves, Sell =
short resolves, Neutral = neither. Consequences:
- dir-precision in the era line stops being a proxy and becomes the win
rate of the strategy under its own exit rules.
- Expected balance ~25/25/50 at the shipped 1:3 (gambler's ruin), i.e.
~2:1 instead of 31:1. Measured and logged at the end of the prebuild.
- Spread is charged on both legs, so it is a NET win rate.
- Intrabar ambiguity resolves to the STOP. OHLC cannot order two touches
inside one bar and the optimistic reading is how a backtested edge
becomes a live loss.
ZigZag stays as input features (EnableSwingContext) and now also supplies
the vertical barrier: the horizon is the median confirmed leg length,
snapped to a coarse ladder. Derived, not configured, and deliberately kept
out of the filename fingerprint - a filename keyed on a measured quantity
orphans a trained model the moment the measurement moves.
Removed, because the premise died with the old target:
- the alternation gate. Correct for pivot labels (a ZigZag cannot emit two
same-type pivots in a row, so a repeat was provably a false fire), and
wrong for barrier labels, which answer each bar independently. It also
took its worst consequence with it: a one-sided model previously got ONE
trade per backtest, a hard blocker on marketplace validation.
- SignalClusterWindow now defaults off - it de-duplicated repeats that are
now real trades. Kept as an opt-in display control.
- LABEL_WINDOW_BARS, the pivot-widening pass, ConfirmedZigZagLabel.
- the era-0 output-bias seed now needs a genuinely dominant class (0.70)
rather than 0.40; at ~50% Neutral a +-3.0 seed is a distortion, not a
correction.
Also fixed, both found while wiring the above:
1. RefreshConvergedSignal sized its buffers from a date delta
(Bars(sym, period, dtStudied, TimeCurrent())). dtStudied is a training
watermark; in the tester it is loaded from a live-chart save AHEAD of
the simulated date, so the interval inverted, Bars() returned ~0, and
the buffer came out at exactly m_historyBars - deep enough for the OHLC
window and far too shallow for the Donchian-50 / 20-bar-return / SMA
extension behind it. Inference silently computed DIFFERENT features
from the ones training learned on, live as well as in the tester. Now
sized from what the feature builder actually needs.
2. The barrier horizon is resolved on the deployed path too. A deployed
model never enters Train(), so it never reached the prebuild, and
OnlineLearnStep reads the horizon as its confirmation delay - left at
the fallback it would have backpropped bars whose barriers had not
resolved. Silent lookahead in the one place that writes to a live model.
SL_Mode/TP_Mode join the weights fingerprint: they define the labels now,
so a model trained at 1:3 must never be silently reused at 1:1. This
re-keys every pre-existing model by design - none were trained on this task.
Inference census extended with the vote gate. LongCondition/ShortCondition
open with a readiness check the refresh counters never see; in the tester it
reduces to "the seeded _optcache.nnw must have LOADED", and if it did not,
every vote is hard-zeroed while the model still answers Buy. The old three
counters would have read that as "the model says Neutral" - false, and a
completely different fix. This is the leading candidate for the
zero-direction backtest and the census can now name it in one run.
Both builds compile 0 errors / 0 warnings. Forces a full retrain.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 20:39:49 -04:00
|
|
|
m_barrierHorizonBars = ComputeBarrierHorizonBars(bars);
|
fix: a restart no longer loses the measured geometry or the training window
Terminal restart, 22:25: all four resumed models sat on empty windows with
enum 2:6 barriers. Three interlocking causes, all visible in one log excerpt:
1) THE PRE-SCAN WINDOW WAS SIZED BY THE SAVED WATERMARK. A resumed model's
dtStudied sits at its last studied bar, so Bars(dtStudied, now) ~ 0 and the
resumed-model MI pre-scan built a zero-bar "complete" label cache - logged as
"Buy: 0 | Sell: 0 | Neutral: 0". Train()'s own era start RESETS dtStudied to
the training-window rule before computing its window; the pre-scan did not.
The rule is now factored into TrainWindowStart() and both use it. The scan
also refuses to arm before SERIES_SYNCHRONIZED (it ran in the same second as
OnInit), and deployed models keep their watermark - for them it gates
inference recency, not a training window.
2) THE HORIZON LATCHED ON AN INDICATOR WARM-UP. ComputeBarrierHorizonBars ran
against a ZigZag with 0 calculated legs, fell back, and EnsureBarrierHorizon
latched fallback(32) x slMult x tpMult = 384 for the process lifetime. A
leg-starved horizon is now PROVISIONAL: re-resolved on the next rebuild, the
label cache wiped if it moved (labels from two horizons answer different
questions), and the geometry deriver refuses to run from it - a pair derived
over a warm-up window would get PINNED.
3) THE DERIVED GEOMETRY WAS NEVER PERSISTED. The .cfg is written at model
creation and at weights-reset - both BEFORE era 0 derives - so the measured
pair lived only in memory: every restart read back zeros, adopted nothing,
fell back to the enum barriers, and the era-0-only gate meant a resumed model
could NEVER re-derive. A full day of training on 3.33/1.62 resumed as 2:6.
Now: the settled pair is pinned to the .cfg the moment derivation completes
(one-shot, atomic write), and the derive gate accepts any model with no
pinned pair, not just era 0 - mid-run stability is carried by
m_geometryDerived itself, which never allows a second derivation.
Both build variants compile 0 errors, 0 warnings.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 22:40:43 -04:00
|
|
|
//--- A leg-starved computation is the fallback, not a measurement - keep it PROVISIONAL so the next
|
|
|
|
|
//--- full rebuild recomputes it, instead of latching an indicator warm-up artifact for the process
|
|
|
|
|
//--- lifetime (see m_barrierHorizonLegStarved).
|
|
|
|
|
m_barrierHorizonResolved = !m_barrierHorizonLegStarved;
|
|
|
|
|
//--- If a re-resolution actually MOVED the horizon, any label cached under the old one answers a
|
|
|
|
|
//--- different question - wipe, and let the prebuild refill under one rule. The wipe is cheap
|
|
|
|
|
//--- relative to training on labels from two different horizons, which is the exact failure the
|
|
|
|
|
//--- once-per-process latch exists to prevent.
|
|
|
|
|
if(m_barrierHorizonBars != prevHorizon && ArraySize(m_labelCacheHasValue) > 0)
|
|
|
|
|
{
|
|
|
|
|
ArrayInitialize(m_labelCacheHasValue, false);
|
|
|
|
|
m_labelCachePrebuilt = false;
|
|
|
|
|
}
|
feat(ai): triple-barrier labels replace exact-pivot ZigZag targets
The 31:1 class imbalance was self-inflicted by the TARGET, not a property
of the market. Labelling only the exact bar where a ZigZag pivot confirms
gave Buy 1164 / Sell 1164 / Neutral 35841, and every correction mechanism
this codebase accumulated sits downstream of that one choice: the
logit-adjusted loss and its range cap, the prior EMA, the +-3.0 output-bias
seed, balanced-accuracy-then-precision selection with its coverage floor,
the recall floor and its catch-22, the alternation gate, NMS, and the four
oversampling designs that collapsed before them.
The reference this engine is built on (references/neuronetworksbook.pdf
ch. 3.1/3.3) also uses ZigZag, but targets the DIRECTION TO THE NEXT
EXTREMUM on every bar - ~50/50 by construction, with no imbalance to
correct at all. It never had this problem because it never asked "is this
the pivot bar".
Labels are now the triple barrier (Lopez de Prado ch. 3), using the EA's
OWN SL_Mode/TP_Mode: does a trade opened at this bar's close reach its
target before its stop, within a horizon. Buy = long resolves, Sell =
short resolves, Neutral = neither. Consequences:
- dir-precision in the era line stops being a proxy and becomes the win
rate of the strategy under its own exit rules.
- Expected balance ~25/25/50 at the shipped 1:3 (gambler's ruin), i.e.
~2:1 instead of 31:1. Measured and logged at the end of the prebuild.
- Spread is charged on both legs, so it is a NET win rate.
- Intrabar ambiguity resolves to the STOP. OHLC cannot order two touches
inside one bar and the optimistic reading is how a backtested edge
becomes a live loss.
ZigZag stays as input features (EnableSwingContext) and now also supplies
the vertical barrier: the horizon is the median confirmed leg length,
snapped to a coarse ladder. Derived, not configured, and deliberately kept
out of the filename fingerprint - a filename keyed on a measured quantity
orphans a trained model the moment the measurement moves.
Removed, because the premise died with the old target:
- the alternation gate. Correct for pivot labels (a ZigZag cannot emit two
same-type pivots in a row, so a repeat was provably a false fire), and
wrong for barrier labels, which answer each bar independently. It also
took its worst consequence with it: a one-sided model previously got ONE
trade per backtest, a hard blocker on marketplace validation.
- SignalClusterWindow now defaults off - it de-duplicated repeats that are
now real trades. Kept as an opt-in display control.
- LABEL_WINDOW_BARS, the pivot-widening pass, ConfirmedZigZagLabel.
- the era-0 output-bias seed now needs a genuinely dominant class (0.70)
rather than 0.40; at ~50% Neutral a +-3.0 seed is a distortion, not a
correction.
Also fixed, both found while wiring the above:
1. RefreshConvergedSignal sized its buffers from a date delta
(Bars(sym, period, dtStudied, TimeCurrent())). dtStudied is a training
watermark; in the tester it is loaded from a live-chart save AHEAD of
the simulated date, so the interval inverted, Bars() returned ~0, and
the buffer came out at exactly m_historyBars - deep enough for the OHLC
window and far too shallow for the Donchian-50 / 20-bar-return / SMA
extension behind it. Inference silently computed DIFFERENT features
from the ones training learned on, live as well as in the tester. Now
sized from what the feature builder actually needs.
2. The barrier horizon is resolved on the deployed path too. A deployed
model never enters Train(), so it never reached the prebuild, and
OnlineLearnStep reads the horizon as its confirmation delay - left at
the fallback it would have backpropped bars whose barriers had not
resolved. Silent lookahead in the one place that writes to a live model.
SL_Mode/TP_Mode join the weights fingerprint: they define the labels now,
so a model trained at 1:3 must never be silently reused at 1:1. This
re-keys every pre-existing model by design - none were trained on this task.
Inference census extended with the vote gate. LongCondition/ShortCondition
open with a readiness check the refresh counters never see; in the tester it
reduces to "the seeded _optcache.nnw must have LOADED", and if it did not,
every vote is hard-zeroed while the model still answers Buy. The old three
counters would have read that as "the model says Neutral" - false, and a
completely different fix. This is the leading candidate for the
zero-direction backtest and the census can now name it in one run.
Both builds compile 0 errors / 0 warnings. Forces a full retrain.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 20:39:49 -04:00
|
|
|
double slMultLog, tpMultLog;
|
|
|
|
|
BarrierMultiples(slMultLog, tpMultLog);
|
fix: both-won bars were labelled "do not trade" - resolve by first touch
Removing the min-reward:risk raise let the MEASURED geometry come back with
the target NEARER than the stop (SP500 H1: target 1.62*ATR at q50 of
favourable, stop 3.33*ATR at q75 of adverse). That reopened a branch the
code called unreachable: price can reach +target and -target inside one
horizon, winning in BOTH directions, and those bars fell through to Neutral.
Neutral has only three producers, both-lost is unreachable (you cannot touch
-3.33 without crossing -1.62 first, which wins the short), and timeouts logged
at 1.0% of Neutral - so ~27% of ALL bars were being handed to the model as the
abstain class when a trade either way would have collected its target. The
cleanest positives in the sample, labelled "do not trade", while the fitted
confidence threshold was being asked to find selectivity in what was left.
Resolved by FIRST TOUCH: the target reached earlier is the trade that would
have closed first. Same forward window, no extra lookahead. Same-bar ties stay
Neutral - OHLC cannot order two touches, and unlike an intrabar stop tie there
is no pessimistic side to fall to, so a guess would inject a coin-flip
direction into the target.
Also:
- count both-won and its same-bar tie subset in the prebuild line, so the
share is measured rather than inferred from arithmetic on a log line
- scope the timeout counter to IS, matching the tally it is reported as a
percentage OF; it was incremented over the whole scan and divided by an
in-sample denominator
- clear m_lastBarrierTimedOut at the top of the walk with the excursions, not
at the bottom - the two early returns published the previous bar's verdict
- mark the pass-1 label line PROVISIONAL. It prints the enum fallback because
geometry can only be derived from excursions that do not exist yet, and it
reads exactly like a config change that failed to take effect
FORCES RETRAIN. Both build variants compile 0 errors, 0 warnings.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 15:35:42 -04:00
|
|
|
//--- PROVISIONAL vs FINAL. The geometry can only be derived from measured excursions, and excursions only
|
|
|
|
|
//--- exist once bars have been labelled, so the first pass necessarily labels with the enum fallback and
|
|
|
|
|
//--- prints it here. It is then discarded: DeriveBarrierGeometry() runs, the horizon re-resolves, the
|
|
|
|
|
//--- cache is wiped and this line prints again with the measured pair. Reading the log without knowing
|
|
|
|
|
//--- that, the first line looks exactly like a config change that failed to take effect - which is how it
|
|
|
|
|
//--- was in fact read. Say which one this is.
|
|
|
|
|
string stage = m_geometryDerived
|
|
|
|
|
? " | MEASURED geometry, this is what trains"
|
|
|
|
|
: " | PROVISIONAL - enum fallback for the measurement pass only, superseded by the "
|
|
|
|
|
"DERIVED pair logged next";
|
fix: a restart no longer loses the measured geometry or the training window
Terminal restart, 22:25: all four resumed models sat on empty windows with
enum 2:6 barriers. Three interlocking causes, all visible in one log excerpt:
1) THE PRE-SCAN WINDOW WAS SIZED BY THE SAVED WATERMARK. A resumed model's
dtStudied sits at its last studied bar, so Bars(dtStudied, now) ~ 0 and the
resumed-model MI pre-scan built a zero-bar "complete" label cache - logged as
"Buy: 0 | Sell: 0 | Neutral: 0". Train()'s own era start RESETS dtStudied to
the training-window rule before computing its window; the pre-scan did not.
The rule is now factored into TrainWindowStart() and both use it. The scan
also refuses to arm before SERIES_SYNCHRONIZED (it ran in the same second as
OnInit), and deployed models keep their watermark - for them it gates
inference recency, not a training window.
2) THE HORIZON LATCHED ON AN INDICATOR WARM-UP. ComputeBarrierHorizonBars ran
against a ZigZag with 0 calculated legs, fell back, and EnsureBarrierHorizon
latched fallback(32) x slMult x tpMult = 384 for the process lifetime. A
leg-starved horizon is now PROVISIONAL: re-resolved on the next rebuild, the
label cache wiped if it moved (labels from two horizons answer different
questions), and the geometry deriver refuses to run from it - a pair derived
over a warm-up window would get PINNED.
3) THE DERIVED GEOMETRY WAS NEVER PERSISTED. The .cfg is written at model
creation and at weights-reset - both BEFORE era 0 derives - so the measured
pair lived only in memory: every restart read back zeros, adopted nothing,
fell back to the enum barriers, and the era-0-only gate meant a resumed model
could NEVER re-derive. A full day of training on 3.33/1.62 resumed as 2:6.
Now: the settled pair is pinned to the .cfg the moment derivation completes
(one-shot, atomic write), and the derive gate accepts any model with no
pinned pair, not just era 0 - mid-run stability is carried by
m_geometryDerived itself, which never allows a second derivation.
Both build variants compile 0 errors, 0 warnings.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 22:40:43 -04:00
|
|
|
if(m_barrierHorizonLegStarved)
|
|
|
|
|
stage += " | horizon PROVISIONAL (ZigZag still warming up, re-resolved on the next rebuild)";
|
feat(ai): triple-barrier labels replace exact-pivot ZigZag targets
The 31:1 class imbalance was self-inflicted by the TARGET, not a property
of the market. Labelling only the exact bar where a ZigZag pivot confirms
gave Buy 1164 / Sell 1164 / Neutral 35841, and every correction mechanism
this codebase accumulated sits downstream of that one choice: the
logit-adjusted loss and its range cap, the prior EMA, the +-3.0 output-bias
seed, balanced-accuracy-then-precision selection with its coverage floor,
the recall floor and its catch-22, the alternation gate, NMS, and the four
oversampling designs that collapsed before them.
The reference this engine is built on (references/neuronetworksbook.pdf
ch. 3.1/3.3) also uses ZigZag, but targets the DIRECTION TO THE NEXT
EXTREMUM on every bar - ~50/50 by construction, with no imbalance to
correct at all. It never had this problem because it never asked "is this
the pivot bar".
Labels are now the triple barrier (Lopez de Prado ch. 3), using the EA's
OWN SL_Mode/TP_Mode: does a trade opened at this bar's close reach its
target before its stop, within a horizon. Buy = long resolves, Sell =
short resolves, Neutral = neither. Consequences:
- dir-precision in the era line stops being a proxy and becomes the win
rate of the strategy under its own exit rules.
- Expected balance ~25/25/50 at the shipped 1:3 (gambler's ruin), i.e.
~2:1 instead of 31:1. Measured and logged at the end of the prebuild.
- Spread is charged on both legs, so it is a NET win rate.
- Intrabar ambiguity resolves to the STOP. OHLC cannot order two touches
inside one bar and the optimistic reading is how a backtested edge
becomes a live loss.
ZigZag stays as input features (EnableSwingContext) and now also supplies
the vertical barrier: the horizon is the median confirmed leg length,
snapped to a coarse ladder. Derived, not configured, and deliberately kept
out of the filename fingerprint - a filename keyed on a measured quantity
orphans a trained model the moment the measurement moves.
Removed, because the premise died with the old target:
- the alternation gate. Correct for pivot labels (a ZigZag cannot emit two
same-type pivots in a row, so a repeat was provably a false fire), and
wrong for barrier labels, which answer each bar independently. It also
took its worst consequence with it: a one-sided model previously got ONE
trade per backtest, a hard blocker on marketplace validation.
- SignalClusterWindow now defaults off - it de-duplicated repeats that are
now real trades. Kept as an opt-in display control.
- LABEL_WINDOW_BARS, the pivot-widening pass, ConfirmedZigZagLabel.
- the era-0 output-bias seed now needs a genuinely dominant class (0.70)
rather than 0.40; at ~50% Neutral a +-3.0 seed is a distortion, not a
correction.
Also fixed, both found while wiring the above:
1. RefreshConvergedSignal sized its buffers from a date delta
(Bars(sym, period, dtStudied, TimeCurrent())). dtStudied is a training
watermark; in the tester it is loaded from a live-chart save AHEAD of
the simulated date, so the interval inverted, Bars() returned ~0, and
the buffer came out at exactly m_historyBars - deep enough for the OHLC
window and far too shallow for the Donchian-50 / 20-bar-return / SMA
extension behind it. Inference silently computed DIFFERENT features
from the ones training learned on, live as well as in the tester. Now
sized from what the feature builder actually needs.
2. The barrier horizon is resolved on the deployed path too. A deployed
model never enters Train(), so it never reached the prebuild, and
OnlineLearnStep reads the horizon as its confirmation delay - left at
the fallback it would have backpropped bars whose barriers had not
resolved. Silent lookahead in the one place that writes to a live model.
SL_Mode/TP_Mode join the weights fingerprint: they define the labels now,
so a model trained at 1:3 must never be silently reused at 1:1. This
re-keys every pre-existing model by design - none were trained on this task.
Inference census extended with the vote gate. LongCondition/ShortCondition
open with a readiness check the refresh counters never see; in the tester it
reduces to "the seeded _optcache.nnw must have LOADED", and if it did not,
every vote is hard-zeroed while the model still answers Buy. The old three
counters would have read that as "the model says Neutral" - false, and a
completely different fix. This is the leading candidate for the
zero-direction backtest and the census can now name it in one run.
Both builds compile 0 errors / 0 warnings. Forces a full retrain.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 20:39:49 -04:00
|
|
|
Print(ID + ": triple-barrier labels - stop " + DoubleToString(slMultLog, 2) + "*ATR, target " +
|
|
|
|
|
DoubleToString(tpMultLog, 2) + "*ATR, horizon " + IntegerToString(m_barrierHorizonBars) +
|
|
|
|
|
" bars (median confirmed ZigZag leg, snapped) | spread charged " +
|
fix: both-won bars were labelled "do not trade" - resolve by first touch
Removing the min-reward:risk raise let the MEASURED geometry come back with
the target NEARER than the stop (SP500 H1: target 1.62*ATR at q50 of
favourable, stop 3.33*ATR at q75 of adverse). That reopened a branch the
code called unreachable: price can reach +target and -target inside one
horizon, winning in BOTH directions, and those bars fell through to Neutral.
Neutral has only three producers, both-lost is unreachable (you cannot touch
-3.33 without crossing -1.62 first, which wins the short), and timeouts logged
at 1.0% of Neutral - so ~27% of ALL bars were being handed to the model as the
abstain class when a trade either way would have collected its target. The
cleanest positives in the sample, labelled "do not trade", while the fitted
confidence threshold was being asked to find selectivity in what was left.
Resolved by FIRST TOUCH: the target reached earlier is the trade that would
have closed first. Same forward window, no extra lookahead. Same-bar ties stay
Neutral - OHLC cannot order two touches, and unlike an intrabar stop tie there
is no pessimistic side to fall to, so a guess would inject a coin-flip
direction into the target.
Also:
- count both-won and its same-bar tie subset in the prebuild line, so the
share is measured rather than inferred from arithmetic on a log line
- scope the timeout counter to IS, matching the tally it is reported as a
percentage OF; it was incremented over the whole scan and divided by an
in-sample denominator
- clear m_lastBarrierTimedOut at the top of the walk with the excursions, not
at the bottom - the two early returns published the previous bar's verdict
- mark the pass-1 label line PROVISIONAL. It prints the enum fallback because
geometry can only be derived from excursions that do not exist yet, and it
reads exactly like a config change that failed to take effect
FORCES RETRAIN. Both build variants compile 0 errors, 0 warnings.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 15:35:42 -04:00
|
|
|
IntegerToString(m_symbol.Spread()) + " points | intrabar ties score as the STOP" + stage);
|
2026-08-15 04:44:10 -04:00
|
|
|
if(IsFractalTarget())
|
|
|
|
|
Print(ID + ": FRACTAL TARGET active - the training label is the direction to the next confirmed"
|
|
|
|
|
" 5-bar fractal extreme (min move max(2 spreads, 0.10 ATR), outside bars Neutral), NOT the"
|
|
|
|
|
" barrier verdict. The barrier geometry above still sizes the live orders and the win-rate"
|
|
|
|
|
" gate: deploy is decided on what a trade at that SL/TP actually collected.");
|
feat(ai): triple-barrier labels replace exact-pivot ZigZag targets
The 31:1 class imbalance was self-inflicted by the TARGET, not a property
of the market. Labelling only the exact bar where a ZigZag pivot confirms
gave Buy 1164 / Sell 1164 / Neutral 35841, and every correction mechanism
this codebase accumulated sits downstream of that one choice: the
logit-adjusted loss and its range cap, the prior EMA, the +-3.0 output-bias
seed, balanced-accuracy-then-precision selection with its coverage floor,
the recall floor and its catch-22, the alternation gate, NMS, and the four
oversampling designs that collapsed before them.
The reference this engine is built on (references/neuronetworksbook.pdf
ch. 3.1/3.3) also uses ZigZag, but targets the DIRECTION TO THE NEXT
EXTREMUM on every bar - ~50/50 by construction, with no imbalance to
correct at all. It never had this problem because it never asked "is this
the pivot bar".
Labels are now the triple barrier (Lopez de Prado ch. 3), using the EA's
OWN SL_Mode/TP_Mode: does a trade opened at this bar's close reach its
target before its stop, within a horizon. Buy = long resolves, Sell =
short resolves, Neutral = neither. Consequences:
- dir-precision in the era line stops being a proxy and becomes the win
rate of the strategy under its own exit rules.
- Expected balance ~25/25/50 at the shipped 1:3 (gambler's ruin), i.e.
~2:1 instead of 31:1. Measured and logged at the end of the prebuild.
- Spread is charged on both legs, so it is a NET win rate.
- Intrabar ambiguity resolves to the STOP. OHLC cannot order two touches
inside one bar and the optimistic reading is how a backtested edge
becomes a live loss.
ZigZag stays as input features (EnableSwingContext) and now also supplies
the vertical barrier: the horizon is the median confirmed leg length,
snapped to a coarse ladder. Derived, not configured, and deliberately kept
out of the filename fingerprint - a filename keyed on a measured quantity
orphans a trained model the moment the measurement moves.
Removed, because the premise died with the old target:
- the alternation gate. Correct for pivot labels (a ZigZag cannot emit two
same-type pivots in a row, so a repeat was provably a false fire), and
wrong for barrier labels, which answer each bar independently. It also
took its worst consequence with it: a one-sided model previously got ONE
trade per backtest, a hard blocker on marketplace validation.
- SignalClusterWindow now defaults off - it de-duplicated repeats that are
now real trades. Kept as an opt-in display control.
- LABEL_WINDOW_BARS, the pivot-widening pass, ConfirmedZigZagLabel.
- the era-0 output-bias seed now needs a genuinely dominant class (0.70)
rather than 0.40; at ~50% Neutral a +-3.0 seed is a distortion, not a
correction.
Also fixed, both found while wiring the above:
1. RefreshConvergedSignal sized its buffers from a date delta
(Bars(sym, period, dtStudied, TimeCurrent())). dtStudied is a training
watermark; in the tester it is loaded from a live-chart save AHEAD of
the simulated date, so the interval inverted, Bars() returned ~0, and
the buffer came out at exactly m_historyBars - deep enough for the OHLC
window and far too shallow for the Donchian-50 / 20-bar-return / SMA
extension behind it. Inference silently computed DIFFERENT features
from the ones training learned on, live as well as in the tester. Now
sized from what the feature builder actually needs.
2. The barrier horizon is resolved on the deployed path too. A deployed
model never enters Train(), so it never reached the prebuild, and
OnlineLearnStep reads the horizon as its confirmation delay - left at
the fallback it would have backpropped bars whose barriers had not
resolved. Silent lookahead in the one place that writes to a live model.
SL_Mode/TP_Mode join the weights fingerprint: they define the labels now,
so a model trained at 1:3 must never be silently reused at 1:1. This
re-keys every pre-existing model by design - none were trained on this task.
Inference census extended with the vote gate. LongCondition/ShortCondition
open with a readiness check the refresh counters never see; in the tester it
reduces to "the seeded _optcache.nnw must have LOADED", and if it did not,
every vote is hard-zeroed while the model still answers Buy. The old three
counters would have read that as "the model says Neutral" - false, and a
completely different fix. This is the leading candidate for the
zero-direction backtest and the census can now name it in one run.
Both builds compile 0 errors / 0 warnings. Forces a full retrain.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 20:39:49 -04:00
|
|
|
}
|
|
|
|
|
//+------------------------------------------------------------------+
|
|
|
|
|
//| Resolves the triple-barrier label for whichever candidate bar is |
|
|
|
|
|
//| exactly m_barrierHorizonBars behind the one being visited - i.e. |
|
|
|
|
|
//| the newest bar whose outcome is now fully knowable. Mirrors the |
|
|
|
|
|
//| shape of the ZigZag-confirmation scan this replaced, with the |
|
|
|
|
|
//| lookahead depth changed from "how long until a pivot stops |
|
|
|
|
|
//| repainting" to "how long until the trade resolves". |
|
|
|
|
|
//| |
|
|
|
|
|
//| Unlike the ZigZag version, a bar's verdict here is FINAL the |
|
|
|
|
|
//| moment it is computed: the barrier outcome depends only on price |
|
|
|
|
|
//| within a fixed forward window, so nothing later can revise it. |
|
|
|
|
|
//| That is what lets the widening/re-spreading pass this file used |
|
|
|
|
|
//| to need disappear entirely. |
|
|
|
|
|
//+------------------------------------------------------------------+
|
|
|
|
|
void CExpertSignalAIBase::AdvanceBarrierLabelState(int i, int bars)
|
|
|
|
|
{
|
|
|
|
|
int idx = i + MathMax(m_barrierHorizonBars, 1);
|
|
|
|
|
if(idx >= bars || m_labelCacheHasValue[idx])
|
|
|
|
|
return;
|
|
|
|
|
ENUM_SIGNAL verdict = TripleBarrierLabel(idx);
|
2026-08-15 04:44:10 -04:00
|
|
|
//--- FRACTAL TARGET: the barrier walk above still runs in full - it fills the excursion/ladder/win
|
|
|
|
|
//--- caches that the measured geometry, the expectancy scan and the era gate's realized-win scoring
|
|
|
|
|
//--- all read - but the TRAINING label it returned is replaced by the fractal-direction verdict.
|
|
|
|
|
//--- The model learns "which way is the next confirmed fractal extreme"; the deploy gate keeps
|
|
|
|
|
//--- scoring what a trade at the EA's own SL/TP actually collected from these bars.
|
|
|
|
|
if(IsFractalTarget())
|
|
|
|
|
verdict = FractalDirectionLabel(idx);
|
fix: both-won bars were labelled "do not trade" - resolve by first touch
Removing the min-reward:risk raise let the MEASURED geometry come back with
the target NEARER than the stop (SP500 H1: target 1.62*ATR at q50 of
favourable, stop 3.33*ATR at q75 of adverse). That reopened a branch the
code called unreachable: price can reach +target and -target inside one
horizon, winning in BOTH directions, and those bars fell through to Neutral.
Neutral has only three producers, both-lost is unreachable (you cannot touch
-3.33 without crossing -1.62 first, which wins the short), and timeouts logged
at 1.0% of Neutral - so ~27% of ALL bars were being handed to the model as the
abstain class when a trade either way would have collected its target. The
cleanest positives in the sample, labelled "do not trade", while the fitted
confidence threshold was being asked to find selectivity in what was left.
Resolved by FIRST TOUCH: the target reached earlier is the trade that would
have closed first. Same forward window, no extra lookahead. Same-bar ties stay
Neutral - OHLC cannot order two touches, and unlike an intrabar stop tie there
is no pessimistic side to fall to, so a guess would inject a coin-flip
direction into the target.
Also:
- count both-won and its same-bar tie subset in the prebuild line, so the
share is measured rather than inferred from arithmetic on a log line
- scope the timeout counter to IS, matching the tally it is reported as a
percentage OF; it was incremented over the whole scan and divided by an
in-sample denominator
- clear m_lastBarrierTimedOut at the top of the walk with the excursions, not
at the bottom - the two early returns published the previous bar's verdict
- mark the pass-1 label line PROVISIONAL. It prints the enum fallback because
geometry can only be derived from excursions that do not exist yet, and it
reads exactly like a config change that failed to take effect
FORCES RETRAIN. Both build variants compile 0 errors, 0 warnings.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 15:35:42 -04:00
|
|
|
//--- IS-ONLY, matching the final tally pass exactly. These counters are reported as percentages OF the
|
|
|
|
|
//--- Buy/Sell/Neutral tallies, and those are IS-only - the timeout count was previously incremented over
|
|
|
|
|
//--- the whole scan and then divided by an in-sample denominator, so its "% of Neutral" could read high
|
|
|
|
|
//--- for no reason but the split. A diagnostic that mixes two populations is worse than no diagnostic:
|
|
|
|
|
//--- it is the horizon check, and it has to be trustworthy to do its job.
|
|
|
|
|
bool countable = (idx >= MathMax(2, m_labelPrebuildOosCutoff)
|
|
|
|
|
&& idx <= bars - MathMax(m_historyBars, 0) - 1);
|
|
|
|
|
if(countable)
|
|
|
|
|
{
|
fix(labels): overlapping-label sample correction + horizon cap on the scale ladder
Three defects, all surfaced by the 2026-08-17 SP500 H4 run that shipped
stop 4.86 / target 9.71 (width 14.57*ATR, horizon 384).
1. EVERY STANDARD ERROR ASSUMED INDEPENDENT SAMPLES. Triple-barrier labels
started one per bar overlap by the label's lifespan, so n calls are worth
~n/L independent observations (Lopez de Prado, AFML ch. 4 - sample
uniqueness). All three sqrt(p(1-p)/n) sites divided by the RAW count.
The tell: the operating point's null-of-the-maximum gate is family-wise and
should fire on ~5% of eras under the null. Measured fire rates - PAI 47/73
(64%), ConvLSTM 9/24, LSTM 8/21 (38%), CONV 4/62 (6%). CONV, the only model
whose margin distribution admits few bins, sat on the null; the rest cleared
a bar that was too low by ~sqrt(L). PAI's deployed threshold consequently
alternated between the ENDS of its own range era to era (0.10 -> 0.88 ->
0.86 -> 0.66; coverage 16% <-> 73%).
TripleBarrierLabel now records when each label became KNOWABLE - the first
winning touch, or both stops, or the timeout - and the prebuild accumulates
the mean. EffectiveSampleSize() feeds the operating point, the member deploy
gate and the ensemble vote gate. Conservative by construction (n/L is an
upper bound on the damage); gates get harder, never easier.
2. THE SCALE LADDER RAN AWAY, again. Horizon scales as swingMedian*sl*tp, and
since 4d8cb08 reachability is measured OVER that horizon - so a wider rung
buys itself the time that makes it look reachable. Same target -> horizon ->
reach -> target loop the excursion window is kept short to avoid; fixing the
window confusion reopened it through the other door. It walked 128 -> 256 ->
384 bars and stopped at q90, the widest rung there is, with every rung
reading 39-48% against a 20% floor. A floor nothing fails selects nothing.
Rungs whose required horizon exceeds BARRIER_HORIZON_MAX are now rejected -
the same rule ReportGeometryExpectancyScan already applied. It was printing
the shipped pair as CLAMPED and disqualified ('h384!') two lines under the
deriver that chose it: two subsystems, one geometry, opposite verdicts.
3. THE RUNG SNAP DESTROYED THE RATIO IT WAS COMPARING. Both legs snapped
independently to the coarse first-passage grid, re-rating each candidate:
q90 4.86/9.71 -> 5.00/10.00 (2.00), q85 4.07/8.14 -> 5.00/10.00 (IDENTICAL
measurement), q75 3.07/6.13 -> 4.00/6.50 (1.63 - a nearer target). So the
ladder compared win shares taken at ratios from 1.63 to 2.17 and read the
differences as scale. It is why the reach column came out non-monotone in
width (q75 48.5% above q90 42.9%). The stop now snaps to its nearest rung in
log space and the target follows the ratio off it; the pair actually measured
is returned and logged, so a collision reads as a collision.
Also: LadderWinShare guarded against the conditional (fractal) geometry path,
which fills n from m_fracLegCount while leaving idxList empty - a latent
out-of-bounds on a currently-dead path.
New log lines: mean label lifespan and effective n on the label-cache line, the
required-vs-available horizon per rung, and the grid pair the reconciliation
actually measured (its tolerance now scales with the grid skew instead of a flat
5pp).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 01:12:05 -04:00
|
|
|
//--- MEAN LABEL LIFESPAN, accumulated on the IS population for the same reason the timeout share is:
|
|
|
|
|
//--- it deflates standard errors computed on that population. 0 = the walk could not resolve this
|
|
|
|
|
//--- bar at all, which is not a lifespan of zero - it is no measurement, so it is skipped.
|
|
|
|
|
//--- Deliberately taken from the BARRIER walk even under the fractal target: the overlap being
|
|
|
|
|
//--- corrected for is the barrier outcome window, which every deploy-gate win rate is scored on.
|
|
|
|
|
if(m_lastLabelLifespan > 0)
|
|
|
|
|
{
|
|
|
|
|
m_labelLifespanSum += (double)m_lastLabelLifespan;
|
|
|
|
|
m_labelLifespanCount++;
|
|
|
|
|
}
|
fix: both-won bars were labelled "do not trade" - resolve by first touch
Removing the min-reward:risk raise let the MEASURED geometry come back with
the target NEARER than the stop (SP500 H1: target 1.62*ATR at q50 of
favourable, stop 3.33*ATR at q75 of adverse). That reopened a branch the
code called unreachable: price can reach +target and -target inside one
horizon, winning in BOTH directions, and those bars fell through to Neutral.
Neutral has only three producers, both-lost is unreachable (you cannot touch
-3.33 without crossing -1.62 first, which wins the short), and timeouts logged
at 1.0% of Neutral - so ~27% of ALL bars were being handed to the model as the
abstain class when a trade either way would have collected its target. The
cleanest positives in the sample, labelled "do not trade", while the fitted
confidence threshold was being asked to find selectivity in what was left.
Resolved by FIRST TOUCH: the target reached earlier is the trade that would
have closed first. Same forward window, no extra lookahead. Same-bar ties stay
Neutral - OHLC cannot order two touches, and unlike an intrabar stop tie there
is no pessimistic side to fall to, so a guess would inject a coin-flip
direction into the target.
Also:
- count both-won and its same-bar tie subset in the prebuild line, so the
share is measured rather than inferred from arithmetic on a log line
- scope the timeout counter to IS, matching the tally it is reported as a
percentage OF; it was incremented over the whole scan and divided by an
in-sample denominator
- clear m_lastBarrierTimedOut at the top of the walk with the excursions, not
at the bottom - the two early returns published the previous bar's verdict
- mark the pass-1 label line PROVISIONAL. It prints the enum fallback because
geometry can only be derived from excursions that do not exist yet, and it
reads exactly like a config change that failed to take effect
FORCES RETRAIN. Both build variants compile 0 errors, 0 warnings.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 15:35:42 -04:00
|
|
|
if(verdict == Neutral && m_lastBarrierTimedOut)
|
|
|
|
|
m_labelPrebuildTimeoutCount++;
|
|
|
|
|
//--- Counted for EVERY verdict, not just Neutral: after first-touch resolution most both-won bars
|
|
|
|
|
//--- now carry a direction, and the interesting number is how much of the label set this class is -
|
|
|
|
|
//--- not how much of it stayed unresolved.
|
|
|
|
|
if(m_lastBarrierBothWon)
|
|
|
|
|
{
|
|
|
|
|
m_labelPrebuildBothWonCount++;
|
|
|
|
|
if(m_lastBarrierBothWonTied)
|
|
|
|
|
m_labelPrebuildBothWonTieCount++;
|
|
|
|
|
}
|
|
|
|
|
}
|
feat(ai): triple-barrier labels replace exact-pivot ZigZag targets
The 31:1 class imbalance was self-inflicted by the TARGET, not a property
of the market. Labelling only the exact bar where a ZigZag pivot confirms
gave Buy 1164 / Sell 1164 / Neutral 35841, and every correction mechanism
this codebase accumulated sits downstream of that one choice: the
logit-adjusted loss and its range cap, the prior EMA, the +-3.0 output-bias
seed, balanced-accuracy-then-precision selection with its coverage floor,
the recall floor and its catch-22, the alternation gate, NMS, and the four
oversampling designs that collapsed before them.
The reference this engine is built on (references/neuronetworksbook.pdf
ch. 3.1/3.3) also uses ZigZag, but targets the DIRECTION TO THE NEXT
EXTREMUM on every bar - ~50/50 by construction, with no imbalance to
correct at all. It never had this problem because it never asked "is this
the pivot bar".
Labels are now the triple barrier (Lopez de Prado ch. 3), using the EA's
OWN SL_Mode/TP_Mode: does a trade opened at this bar's close reach its
target before its stop, within a horizon. Buy = long resolves, Sell =
short resolves, Neutral = neither. Consequences:
- dir-precision in the era line stops being a proxy and becomes the win
rate of the strategy under its own exit rules.
- Expected balance ~25/25/50 at the shipped 1:3 (gambler's ruin), i.e.
~2:1 instead of 31:1. Measured and logged at the end of the prebuild.
- Spread is charged on both legs, so it is a NET win rate.
- Intrabar ambiguity resolves to the STOP. OHLC cannot order two touches
inside one bar and the optimistic reading is how a backtested edge
becomes a live loss.
ZigZag stays as input features (EnableSwingContext) and now also supplies
the vertical barrier: the horizon is the median confirmed leg length,
snapped to a coarse ladder. Derived, not configured, and deliberately kept
out of the filename fingerprint - a filename keyed on a measured quantity
orphans a trained model the moment the measurement moves.
Removed, because the premise died with the old target:
- the alternation gate. Correct for pivot labels (a ZigZag cannot emit two
same-type pivots in a row, so a repeat was provably a false fire), and
wrong for barrier labels, which answer each bar independently. It also
took its worst consequence with it: a one-sided model previously got ONE
trade per backtest, a hard blocker on marketplace validation.
- SignalClusterWindow now defaults off - it de-duplicated repeats that are
now real trades. Kept as an opt-in display control.
- LABEL_WINDOW_BARS, the pivot-widening pass, ConfirmedZigZagLabel.
- the era-0 output-bias seed now needs a genuinely dominant class (0.70)
rather than 0.40; at ~50% Neutral a +-3.0 seed is a distortion, not a
correction.
Also fixed, both found while wiring the above:
1. RefreshConvergedSignal sized its buffers from a date delta
(Bars(sym, period, dtStudied, TimeCurrent())). dtStudied is a training
watermark; in the tester it is loaded from a live-chart save AHEAD of
the simulated date, so the interval inverted, Bars() returned ~0, and
the buffer came out at exactly m_historyBars - deep enough for the OHLC
window and far too shallow for the Donchian-50 / 20-bar-return / SMA
extension behind it. Inference silently computed DIFFERENT features
from the ones training learned on, live as well as in the tester. Now
sized from what the feature builder actually needs.
2. The barrier horizon is resolved on the deployed path too. A deployed
model never enters Train(), so it never reached the prebuild, and
OnlineLearnStep reads the horizon as its confirmation delay - left at
the fallback it would have backpropped bars whose barriers had not
resolved. Silent lookahead in the one place that writes to a live model.
SL_Mode/TP_Mode join the weights fingerprint: they define the labels now,
so a model trained at 1:3 must never be silently reused at 1:1. This
re-keys every pre-existing model by design - none were trained on this task.
Inference census extended with the vote gate. LongCondition/ShortCondition
open with a readiness check the refresh counters never see; in the tester it
reduces to "the seeded _optcache.nnw must have LOADED", and if it did not,
every vote is hard-zeroed while the model still answers Buy. The old three
counters would have read that as "the model says Neutral" - false, and a
completely different fix. This is the leading candidate for the
zero-direction backtest and the census can now name it in one run.
Both builds compile 0 errors / 0 warnings. Forces a full retrain.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 20:39:49 -04:00
|
|
|
m_labelCacheBuy[idx] = (verdict == Buy);
|
|
|
|
|
m_labelCacheSell[idx] = (verdict == Sell);
|
diag: is "optimal SL/TP" learnable? Score the features against excursions
Proposed direction: train the net to predict entry/SL/TP that maximise return
and minimise drawdown, rather than to classify direction. Before rebuilding a
head, measure whether the target is learnable at all.
That question splits into two that behave nothing alike:
HOW FAR price travels (MFE/MAE) - essentially volatility, and volatility
clustering is about the most robust regularity in markets.
WHICH WAY it goes first (the asymmetry) - direction, which is what every
noise-floor verdict in this project has been about.
Expectancy comes ONLY from the second. The first buys position sizing and
drawdown control - worth having under prop-firm limits, but not an edge: exit
management on RANDOM entries already moved the payoff ratio 0.92 -> 5.72 with
expectancy FLAT.
Crucially this is NOT already answered. Every MI figure here scored the
triple-barrier label, i.e. one specific question at one fixed geometry. A
noise-floor result there says nothing about whether excursion MAGNITUDE is
learnable - different target, different answer.
Four targets, and the verdict is the CONTRAST, printed explicitly because the
dangerous misreading of "UP clears" is "we can predict profitable trades":
RANGE (up+dn) - realised volatility, included as a POSITIVE CONTROL that
SHOULD clear. Every prior verdict here lacked a control
expected to pass; a range target at the floor indicts the
measurement, not the market.
UP / DOWN - MFE / MAE.
ASYMMETRY - up-dn, the only one that can pay.
Collected inside the walk the label already does (one max, one min per bar).
The early-out when both barriers resolved is GONE: it would have truncated the
excursions at whichever bar tripped the last barrier, making the measurement a
function of the CURRENT SL/TP - the circularity this is trying to escape. The
loop was already bounded by the horizon, so only the average cost moves.
Discretised into 3 EQUAL-FREQUENCY bins, so every downstream piece (block
permutation, null, p-value) is reused unchanged. Equal-frequency because MFE is
fat-tailed and fixed-width bins would put nearly every row in bin 0; it also
pins H(Y) at ln(3)=1.099 for all four, making them comparable to each other and
to the barrier label's ~1.02 instead of confounded by class balance.
Two bugs fixed in this code before it ever ran, both of which would have
produced a plausible quiet wrong answer rather than an error:
- TripleBarrierLabel early-returns on invalid ATR/close BEFORE the point the
accumulators were reset, so one bar's excursions would be cached under
another bar's index. Cleared at the top now, ahead of every return.
- An unresolvable bar is still flagged as labelled but carries excursions of
exactly 0. Under equal-frequency binning a block of identical zeros drags
the lowest cut onto zero and a third of the sample lands in one
uninformative bin - a depressed score that reads as "not predictable", a
false negative in the direction that would wrongly kill the idea. Rows
where both excursions are zero are dropped; price cannot travel zero both
ways over a whole horizon.
Read-only diagnostic. No topology or label change: no retrain of its own.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-07 10:22:41 -04:00
|
|
|
//--- Stored under the SAME validity flag as the label, set last so no reader can see one without the
|
|
|
|
|
//--- other. TripleBarrierLabel() publishes these for the bar it just walked.
|
|
|
|
|
if(idx < ArraySize(m_excUpCache))
|
|
|
|
|
{
|
|
|
|
|
m_excUpCache[idx] = m_lastExcUp;
|
|
|
|
|
m_excDownCache[idx] = m_lastExcDown;
|
|
|
|
|
}
|
feat: first-passage ladder + expectancy scan - price every geometry, not just the chosen one
Corrects the premise of the previous plan. Break-even is NOT a ceiling.
If the model shifts the win probability on the bars it selects from
p0 = m/(m+k) to p0 + d, then
EV = (p0+d)*k - (1-p0-d)*m = d*(k+m)
because p0*k - (1-p0)*m is zero by construction. The stop:target RATIO
is expectancy-neutral - a punishing break-even is exactly repaid by the
payoff - and only the real edge d and the TOTAL WIDTH (k+m) move EV.
Width matters because the spread is charged once per trade however wide
the barriers are, so a narrow barrier spends much of its own range on
costs. DeriveBarrierGeometry's own comment already said the ratio buys
nothing; the objective just never followed from it.
Blocker this had to solve first: m_excUpCache/m_excDownCache hold only
MAXIMUM travel each way, and a maximum cannot say which side was
reached FIRST - so any geometry other than the walked one was
undecidable on precisely the bars where both barriers were touched,
~28% of the sample.
- BARRIER_LADDER: per bar, the first-touch AGE for 8 travel distances
in each direction, filled during the walk the labels already run.
Cursors keep it O(1) amortised per walked bar rather than 16
comparisons. Levels are travel FROM ENTRY, not barrier prices, so one
ladder serves both directions and the spread is applied analytically
when a level converts back to an SL/TP multiple - storing prices
would need four ladders and bake today's spread into the cache.
Sized, invalidated and validity-gated with the label caches.
- ReportGeometryExpectancyScan: every ladder pair priced exactly off
that cache - width in ATR and in SPREADS (cost efficiency, knowable
without knowing d), break-even, both base rates, the share of bars
resolved inside the horizon, and EV per unit of edge. Compares the
widest resolvable pair against the quantile rule's pick.
MEASUREMENT ONLY - the quantile rule still chooses. Nothing here can
measure d, and width buys nothing if the wider target is less
predictable. Base rates are printed beside each break-even because a
persistent gap is DRIFT and must not be credited to the model.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-10 12:59:18 -04:00
|
|
|
//--- Published under the SAME validity flag as the label and the excursions, for the same reason:
|
|
|
|
|
//--- a reader must never see one without the others (see BARRIER_LADDER).
|
|
|
|
|
int ladderBase = idx * BARRIER_LADDER_COUNT;
|
|
|
|
|
if(ladderBase + BARRIER_LADDER_COUNT <= ArraySize(m_ladderUpAt))
|
|
|
|
|
for(int L = 0; L < BARRIER_LADDER_COUNT; L++)
|
|
|
|
|
{
|
|
|
|
|
m_ladderUpAt[ladderBase + L] = m_lastLadderUpAt[L];
|
|
|
|
|
m_ladderDownAt[ladderBase + L] = m_lastLadderDownAt[L];
|
|
|
|
|
}
|
fix: the deploy gate was benchmarking a win rate against a label frequency
The gate rests on an invariant stated at ExpertSignalAIBase.mqh:199 - under a
driftless walk P(touch +k before -m) is m/(m+k), and break-even for a k:m trade
is ALSO m/(m+k), so "beats chance" and "is profitable" are the same test.
That invariant needs reward >= risk, and the measured geometry no longer
satisfies it. With target 1.62*ATR and stop 3.33*ATR, break-even is 67.3%, but
both-won bars were stripped out of Buy and Sell so the label base rate read
37.5%. chancePrecPct is max(BuyTotal,SellTotal)/bars, so the gate was clearing
models nearly 30pp short of break-even: 42% "directional precision" is +4 sigma
against 37.5% and loses money on every single trade. Live since 217b9bc.
Root cause is that label agreement stopped being the same question as trade
profitability. Buy implies winLong, but the converse fails on every both-won
bar, and the label can only name one of two directions that both pay.
So stop asking the model whether it matched a label and start asking whether
its trade paid:
- cache winLong/winShort per bar beside the label, under the same validity
flag; published from the barrier walk before the collapse to 3 classes
- dirPrecPct now counts wins on the side actually called
- chancePrecPct is max(P(winLong), P(winShort)), MEASURED - the textbook
m/(m+k) would credit SP500's drift to the model
- the NMS "what would I have made" pair, the live-fired precision, and the
IS/OOS cumulative win rates all move to the same test. IS and OOS are read
side by side as the overfitting signal, so measuring one in wins and the
other in agreement would put a fixed gap between them that has nothing to do
with generalization
- the confidence threshold is FITTED on wins too, so the operating point
maximises what the gate grades
- per-class label-agreement precision is still computed and logged; it is the
right diagnostic for class separation, just not for a deploy decision
- era line renamed dir-precision -> win-rate, chance -> chance=break-even
Both build variants compile 0 errors, 0 warnings.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 16:00:51 -04:00
|
|
|
if(idx < ArraySize(m_winLongCache))
|
|
|
|
|
{
|
|
|
|
|
m_winLongCache[idx] = m_lastWinLong;
|
|
|
|
|
m_winShortCache[idx] = m_lastWinShort;
|
|
|
|
|
}
|
refactor: split CExpertSignalAIBase implementation by responsibility
ExpertSignalAIBase.mqh was 8216 lines: the class declaration followed by 87
method bodies covering training, labelling, feature extraction, persistence,
chart drawing, online learning, the GA auto-tuner and inference, all in one
file. Train() alone is 1492 lines; a change to arrow drawing meant scrolling
past the era loop.
Moved the bodies into Expert\AIBase\, included at the bottom of the original
after the class declaration:
Training.mqh 1607 era loop, plateau ladder, checkpoint select, deploy
Features.mqh 1093 indicator creation + per-bar input feature vector
ChartUI.mqh 634 arrows, arrow persistence, status panel, cleanup
Persistence.mqh 492 .stats/.cfg sidecars, CPU-inference validation, copy
OnlineLearning.mqh 461 live continual learning, EMA shadow, OOS simulator
Labels.mqh 309 ZigZag pivot labels, async label-cache prebuild
AutoTune.mqh 275 genetic tuner (population, crossover, halving)
Inference.mqh 235 softmax, prior calibration, class priors
ExpertSignalAIBase.mqh 8216 -> 3131 (declaration + topology build only)
This is a pure relocation - verified mechanically, not by eye: HEAD's file
reconstructed from the eight partials plus the surviving remainder is
byte-identical to HEAD, span for span (scratchpad verify_split.py). No
declaration moved, no signature changed, no code rewritten, so behaviour is
unchanged by construction.
Compiles 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 00:42:45 -04:00
|
|
|
m_labelCacheHasValue[idx] = true;
|
|
|
|
|
}
|
|
|
|
|
//+------------------------------------------------------------------+
|
|
|
|
|
//| Nearest confirmed ZigZag pivot at fromIdx or older (now-relative |
|
|
|
|
|
//| index, so "older" means scanning with INCREASING p - see this |
|
|
|
|
|
//| file's now-relative-index convention, same as AdvanceZigZagLabel- |
|
|
|
|
|
//| State() above). Capped at SWING_SCAN_CAP_BARS so a long quiet |
|
|
|
|
|
//| stretch with no qualifying pivot can't turn this into an unbounded |
|
|
|
|
|
//| scan; returns false (no pivot found) rather than looping forever |
|
|
|
|
|
//| if the cap is hit or history runs out first. |
|
|
|
|
|
//| |
|
|
|
|
|
//| Caller's responsibility, not this method's: applying the |
|
|
|
|
|
//| m_swingConfirmationBars repainting embargo to fromIdx before |
|
|
|
|
|
//| calling. This method itself just finds the nearest nonzero |
|
|
|
|
|
//| ADZigZag buffer entry at/after whatever index it's given - it has |
|
|
|
|
|
//| no opinion on whether that index is safe to read yet. The ONE |
|
|
|
|
|
//| caller that needs the embargo (BufferTempDataCompute()'s |
|
|
|
|
|
//| m_useSwingContext block, looking up "the pivot as of THIS bar") |
|
|
|
|
|
//| applies it before the first call; the second call in that same |
|
|
|
|
|
//| block (finding the PRIOR completed leg, starting from pivotIdx+1) |
|
|
|
|
|
//| doesn't need to re-apply it - anything at or before an already- |
|
|
|
|
|
//| confirmed pivot is necessarily even older, hence already confirmed |
|
|
|
|
|
//| too. |
|
|
|
|
|
//+------------------------------------------------------------------+
|
|
|
|
|
bool CExpertSignalAIBase::FindConfirmedZigZagPivot(int fromIdx, int &pivotIdx, double &pivotPrice, bool &pivotIsLow)
|
|
|
|
|
{
|
|
|
|
|
for(int p = MathMax(fromIdx, 0); p < fromIdx + SWING_SCAN_CAP_BARS; p++)
|
|
|
|
|
{
|
|
|
|
|
if(m_Open.GetData(p) == EMPTY_VALUE)
|
|
|
|
|
return false; // ran off the end of available history
|
|
|
|
|
double zz = m_ADZigZag.GetData(0, p);
|
|
|
|
|
if(zz == 0.0)
|
|
|
|
|
continue;
|
|
|
|
|
pivotIdx = p;
|
|
|
|
|
pivotPrice = zz;
|
|
|
|
|
pivotIsLow = (zz <= m_Low.GetData(p) + _Point);
|
|
|
|
|
return true;
|
|
|
|
|
}
|
|
|
|
|
return false;
|
|
|
|
|
}
|
|
|
|
|
//+------------------------------------------------------------------+
|
2026-08-15 04:44:10 -04:00
|
|
|
//| FRACTAL-DIRECTION LABEL for one bar (TrainingTarget=TARGET_ |
|
|
|
|
|
//| FRACTAL). Direction of price from bar idx's close to the NEXT |
|
|
|
|
|
//| confirmed strict 5-bar fractal extreme - the reference library's |
|
|
|
|
|
//| per-bar extremum-direction target, ~balanced by construction. |
|
|
|
|
|
//| |
|
|
|
|
|
//| A bar j is an up-fractal when its high strictly exceeds the highs |
|
|
|
|
|
//| of the two bars on each side (down-fractal mirrored on lows), and |
|
|
|
|
|
//| it is CONFIRMED two bars later - deterministic, no repainting, so |
|
|
|
|
|
//| unlike the ZigZag there is no embargo to wait out and labels |
|
|
|
|
|
//| resolve to within ~2 bars of the present. |
|
|
|
|
|
//| |
|
|
|
|
|
//| The walk runs newer (decreasing index) from idx-1: the first bar |
|
|
|
|
|
//| that proves out as a fractal extreme is the next swing marker in |
|
|
|
|
|
//| time. Costs use the same bid-series convention as the barrier |
|
|
|
|
|
//| label: a long fills at close+spread and is marked against the |
|
|
|
|
|
//| extreme's bid price, mirrored for shorts. A move that cannot |
|
|
|
|
|
//| clear PIVOT_MIN_MOVE (noise floor: whichever is larger of |
|
|
|
|
|
//| 2 spreads or 0.10 ATR) labels Neutral - "the next turn is too |
|
|
|
|
|
//| close to pay for reaching it". An outside bar that is both an up- |
|
|
|
|
|
//| and a down-fractal is directionally unorderable within OHLC and |
|
|
|
|
|
//| labels Neutral for the same reason barrier ties score as the |
|
|
|
|
|
//| stop: the optimistic reading is how a backtest lies. |
|
|
|
|
|
//+------------------------------------------------------------------+
|
|
|
|
|
ENUM_SIGNAL CExpertSignalAIBase::FractalDirectionLabel(int idx)
|
|
|
|
|
{
|
|
|
|
|
double entry = m_Close.GetData(idx);
|
|
|
|
|
double atr = m_ATR.Main(idx);
|
|
|
|
|
if(!MathIsValidNumber(entry) || entry <= 0.0 || !MathIsValidNumber(atr) || atr <= 0.0)
|
|
|
|
|
return Neutral;
|
|
|
|
|
double spread = (double)m_symbol.Spread() * m_symbol.Point();
|
|
|
|
|
if(!MathIsValidNumber(spread) || spread < 0.0)
|
|
|
|
|
spread = 0.0;
|
|
|
|
|
double minMove = MathMax(2.0 * spread, 0.10 * atr);
|
|
|
|
|
//--- p walks FORWARD IN TIME (indices shrink toward now). A fractal at p needs the two newer
|
|
|
|
|
//--- neighbours p-1/p-2 to exist, so the scan stops at p == 2; a bar closer to now than that has an
|
|
|
|
|
//--- unconfirmable label and stays Neutral - same convention as the barrier's unresolved horizon.
|
|
|
|
|
int deepest = idx - 1;
|
|
|
|
|
int shallowest = MathMax(idx - SWING_SCAN_CAP_BARS, 2);
|
2026-08-15 19:00:40 -04:00
|
|
|
//--- Leg extremes over every bar visited (the extreme bar included): the conditional MFE/MAE the
|
|
|
|
|
//--- geometry derivation feeds on - travel measured over exactly the leg the label points at.
|
|
|
|
|
double legHi = -DBL_MAX, legLo = DBL_MAX;
|
2026-08-15 04:44:10 -04:00
|
|
|
for(int p = deepest; p >= shallowest; p--)
|
|
|
|
|
{
|
|
|
|
|
double h0 = m_High.GetData(p);
|
|
|
|
|
double l0 = m_Low.GetData(p);
|
|
|
|
|
if(h0 == EMPTY_VALUE || l0 == EMPTY_VALUE || !MathIsValidNumber(h0) || !MathIsValidNumber(l0))
|
|
|
|
|
return Neutral; // ran off loaded history before a marker confirmed
|
2026-08-15 19:00:40 -04:00
|
|
|
if(h0 > legHi)
|
|
|
|
|
legHi = h0;
|
|
|
|
|
if(l0 < legLo)
|
|
|
|
|
legLo = l0;
|
2026-08-15 04:44:10 -04:00
|
|
|
bool up = h0 > m_High.GetData(p + 1) && h0 > m_High.GetData(p + 2)
|
|
|
|
|
&& h0 > m_High.GetData(p - 1) && h0 > m_High.GetData(p - 2);
|
|
|
|
|
bool dn = l0 < m_Low.GetData(p + 1) && l0 < m_Low.GetData(p + 2)
|
|
|
|
|
&& l0 < m_Low.GetData(p - 1) && l0 < m_Low.GetData(p - 2);
|
|
|
|
|
if(!up && !dn)
|
|
|
|
|
continue;
|
|
|
|
|
if(up && dn)
|
|
|
|
|
return Neutral; // outside bar: both extremes, unorderable within OHLC
|
2026-08-15 19:00:40 -04:00
|
|
|
ENUM_SIGNAL verdict;
|
2026-08-15 04:44:10 -04:00
|
|
|
if(up)
|
2026-08-15 19:00:40 -04:00
|
|
|
verdict = (h0 - (entry + spread) >= minMove) ? Buy : Neutral;
|
|
|
|
|
else
|
|
|
|
|
verdict = ((entry - spread) - l0 >= minMove) ? Sell : Neutral;
|
|
|
|
|
//--- Record the labeled leg's conditional excursions - IS region, prebuild passes only, and only
|
|
|
|
|
//--- until the geometry is derived and pinned (see m_fracLegFav's declaration comment). Both
|
|
|
|
|
//--- prebuild passes would otherwise record each bar twice: pass 1 (provisional geometry) fills
|
|
|
|
|
//--- these, derivation runs off them, and pass 2's relabel walk finds m_geometryDerived set.
|
|
|
|
|
if(verdict != Neutral && !m_geometryDerived && m_labelPrebuildActive
|
|
|
|
|
&& idx >= MathMax(2, m_labelPrebuildOosCutoff))
|
|
|
|
|
{
|
|
|
|
|
if(verdict == Buy)
|
|
|
|
|
RecordFractalLegExcursion((legHi - entry) / atr, (entry - legLo) / atr);
|
|
|
|
|
else
|
|
|
|
|
RecordFractalLegExcursion((entry - legLo) / atr, (legHi - entry) / atr);
|
|
|
|
|
}
|
|
|
|
|
return verdict;
|
2026-08-15 04:44:10 -04:00
|
|
|
}
|
|
|
|
|
return Neutral; // no fractal inside the scan cap - dead-quiet stretch, nothing to aim at
|
|
|
|
|
}
|
|
|
|
|
//+------------------------------------------------------------------+
|
refactor: split CExpertSignalAIBase implementation by responsibility
ExpertSignalAIBase.mqh was 8216 lines: the class declaration followed by 87
method bodies covering training, labelling, feature extraction, persistence,
chart drawing, online learning, the GA auto-tuner and inference, all in one
file. Train() alone is 1492 lines; a change to arrow drawing meant scrolling
past the era loop.
Moved the bodies into Expert\AIBase\, included at the bottom of the original
after the class declaration:
Training.mqh 1607 era loop, plateau ladder, checkpoint select, deploy
Features.mqh 1093 indicator creation + per-bar input feature vector
ChartUI.mqh 634 arrows, arrow persistence, status panel, cleanup
Persistence.mqh 492 .stats/.cfg sidecars, CPU-inference validation, copy
OnlineLearning.mqh 461 live continual learning, EMA shadow, OOS simulator
Labels.mqh 309 ZigZag pivot labels, async label-cache prebuild
AutoTune.mqh 275 genetic tuner (population, crossover, halving)
Inference.mqh 235 softmax, prior calibration, class priors
ExpertSignalAIBase.mqh 8216 -> 3131 (declaration + topology build only)
This is a pure relocation - verified mechanically, not by eye: HEAD's file
reconstructed from the eight partials plus the surviving remainder is
byte-identical to HEAD, span for span (scratchpad verify_split.py). No
declaration moved, no signature changed, no code rewritten, so behaviour is
unchanged by construction.
Compiles 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 00:42:45 -04:00
|
|
|
//| Kicks off the one-time eager label-cache pre-build for a fresh |
|
|
|
|
|
//| start (see m_labelCachePrebuilt's declaration comment). Computes |
|
|
|
|
|
//| the bar count/OOS split exactly as Train()'s era-start block |
|
|
|
|
|
//| would, then arms AdvanceLabelCachePrebuild() to do the actual |
|
|
|
|
|
//| chunked scan on this and subsequent Train() calls. |
|
|
|
|
|
//+------------------------------------------------------------------+
|
|
|
|
|
void CExpertSignalAIBase::StartLabelCachePrebuild(void)
|
|
|
|
|
{
|
fix: a restart no longer loses the measured geometry or the training window
Terminal restart, 22:25: all four resumed models sat on empty windows with
enum 2:6 barriers. Three interlocking causes, all visible in one log excerpt:
1) THE PRE-SCAN WINDOW WAS SIZED BY THE SAVED WATERMARK. A resumed model's
dtStudied sits at its last studied bar, so Bars(dtStudied, now) ~ 0 and the
resumed-model MI pre-scan built a zero-bar "complete" label cache - logged as
"Buy: 0 | Sell: 0 | Neutral: 0". Train()'s own era start RESETS dtStudied to
the training-window rule before computing its window; the pre-scan did not.
The rule is now factored into TrainWindowStart() and both use it. The scan
also refuses to arm before SERIES_SYNCHRONIZED (it ran in the same second as
OnInit), and deployed models keep their watermark - for them it gates
inference recency, not a training window.
2) THE HORIZON LATCHED ON AN INDICATOR WARM-UP. ComputeBarrierHorizonBars ran
against a ZigZag with 0 calculated legs, fell back, and EnsureBarrierHorizon
latched fallback(32) x slMult x tpMult = 384 for the process lifetime. A
leg-starved horizon is now PROVISIONAL: re-resolved on the next rebuild, the
label cache wiped if it moved (labels from two horizons answer different
questions), and the geometry deriver refuses to run from it - a pair derived
over a warm-up window would get PINNED.
3) THE DERIVED GEOMETRY WAS NEVER PERSISTED. The .cfg is written at model
creation and at weights-reset - both BEFORE era 0 derives - so the measured
pair lived only in memory: every restart read back zeros, adopted nothing,
fell back to the enum barriers, and the era-0-only gate meant a resumed model
could NEVER re-derive. A full day of training on 3.33/1.62 resumed as 2:6.
Now: the settled pair is pinned to the .cfg the moment derivation completes
(one-shot, atomic write), and the derive gate accepts any model with no
pinned pair, not just era 0 - mid-run stability is carried by
m_geometryDerived itself, which never allows a second derivation.
Both build variants compile 0 errors, 0 warnings.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 22:40:43 -04:00
|
|
|
//--- Not armed until the history is synced - the caller retries on its next scheduled call. Without
|
|
|
|
|
//--- this, a terminal restart ran the resumed-model pre-scan in the same second as OnInit, against
|
|
|
|
|
//--- whatever the terminal had loaded so far.
|
|
|
|
|
if(!SeriesInfoInteger(m_symbol.Name(), PERIOD_CURRENT, SERIES_SYNCHRONIZED))
|
|
|
|
|
return;
|
fix(chart): a purge that reports "zero leftovers" was only ever checking its own list
2026-08-17 21:58: all three charts hit "Abnormal termination" ~5.3 s into
OnDeinit with NO cleanup-timings line - the teardown was starved again. The
22:00 init purge then removed 993 / 1373 / 1557 stranded objects and reported
ZERO by-name leftovers on every chart, and the charts still came up with
duplicated panels. "Nothing matching our prefixes remains" and "the chart is
clean" are different statements and only the first was being made.
Three changes, in the order they matter:
1. WHY the teardown starved, and it is a gap in ad80e0b. StartLabelCachePrebuild
runs ResizeBuffers + RefreshData over the FULL study window (33,984 bars on
XAUUSD), unchunked, and OnDeinit cannot begin until it returns. Normally a
once-per-run cost. That night SP500 and XAUUSD LSTM were wedged in the "cache
invalidated at era start" loop, which calls it on EVERY Train() call - two
members re-preparing tens of thousands of bars indefinitely. The terminal
closed into that. Guarded now, plus a resumable guard in the prebuild chunk
loop (the tally pass after it is not chunked).
2. Catch-all "Warrior" prefix in WarriorChartPrefixes. Every family this EA
creates is named Warrior* except the arrows (WarSig_), so one bare prefix
covers the three named entries AND anything a rename or a stale .ex5 left
under a name nobody remembers. Still a prefix delete, never
ObjectsDeleteAll(chart) - the user's own drawings are not ours to remove. Does
not defeat skipArrows: "WarSig_" does not start with "Warrior".
3. The init purge now REPORTS the residue it did not claim, by name (up to 12).
Not deleted - an unmatched object may belong to the user or another indicator.
If a Warrior panel is visible and appears in neither the removed count nor
this list, the prefix list has drifted a third time and the name is in the
journal instead of being inferred from a screenshot.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 22:07:09 -04:00
|
|
|
//--- THE GAP IN THE TEARDOWN GUARDS (ad80e0b), found by the 2026-08-17 21:58 shutdown. Everything below
|
|
|
|
|
//--- is ResizeBuffers + RefreshData over the FULL study window - 33,984 bars on XAUUSD - and it is
|
|
|
|
|
//--- unchunked, so OnDeinit cannot begin until it returns. Normally that is a once-per-run cost and it
|
|
|
|
|
//--- does not matter. It mattered that night because SP500 and XAUUSD LSTM were wedged in the
|
|
|
|
|
//--- "cache invalidated at era start" loop, which calls this on EVERY Train() call: two members
|
|
|
|
|
//--- re-preparing tens of thousands of bars, forever. The terminal closed into that, all three charts hit
|
|
|
|
|
//--- "Abnormal termination" ~5.3 s later with no cleanup-timings line, and 993/1373/1557 chart objects
|
|
|
|
|
//--- were stranded. The stall is a separate bug (see the ANCHOR MOVED / SIZE CHANGED diagnostic); this is
|
|
|
|
|
//--- the reason it took the CHARTS down with it, and it is worth closing on its own.
|
|
|
|
|
if(ShutdownRequested())
|
|
|
|
|
return;
|
fix: a restart no longer loses the measured geometry or the training window
Terminal restart, 22:25: all four resumed models sat on empty windows with
enum 2:6 barriers. Three interlocking causes, all visible in one log excerpt:
1) THE PRE-SCAN WINDOW WAS SIZED BY THE SAVED WATERMARK. A resumed model's
dtStudied sits at its last studied bar, so Bars(dtStudied, now) ~ 0 and the
resumed-model MI pre-scan built a zero-bar "complete" label cache - logged as
"Buy: 0 | Sell: 0 | Neutral: 0". Train()'s own era start RESETS dtStudied to
the training-window rule before computing its window; the pre-scan did not.
The rule is now factored into TrainWindowStart() and both use it. The scan
also refuses to arm before SERIES_SYNCHRONIZED (it ran in the same second as
OnInit), and deployed models keep their watermark - for them it gates
inference recency, not a training window.
2) THE HORIZON LATCHED ON AN INDICATOR WARM-UP. ComputeBarrierHorizonBars ran
against a ZigZag with 0 calculated legs, fell back, and EnsureBarrierHorizon
latched fallback(32) x slMult x tpMult = 384 for the process lifetime. A
leg-starved horizon is now PROVISIONAL: re-resolved on the next rebuild, the
label cache wiped if it moved (labels from two horizons answer different
questions), and the geometry deriver refuses to run from it - a pair derived
over a warm-up window would get PINNED.
3) THE DERIVED GEOMETRY WAS NEVER PERSISTED. The .cfg is written at model
creation and at weights-reset - both BEFORE era 0 derives - so the measured
pair lived only in memory: every restart read back zeros, adopted nothing,
fell back to the enum barriers, and the era-0-only gate meant a resumed model
could NEVER re-derive. A full day of training on 3.33/1.62 resumed as 2:6.
Now: the settled pair is pinned to the .cfg the moment derivation completes
(one-shot, atomic write), and the derive gate accepts any model with no
pinned pair, not just era 0 - mid-run stability is carried by
m_geometryDerived itself, which never allows a second derivation.
Both build variants compile 0 errors, 0 warnings.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 22:40:43 -04:00
|
|
|
//--- A model that is still TRAINING sizes its window by the training rule, not by the saved study
|
|
|
|
|
//--- watermark. The watermark of a caught-up model sits at its last studied bar, so on resume
|
|
|
|
|
//--- Bars(dtStudied, now) is ~0 and the whole pipeline downstream ran on an empty window: a zero-bar
|
|
|
|
|
//--- "complete" cache ("Buy: 0 | Sell: 0 | Neutral: 0"), a horizon from 0 ZigZag legs, no geometry.
|
|
|
|
|
//--- Train()'s own era start applies this exact reset (TrainWindowStart) - this makes the pre-scan
|
|
|
|
|
//--- and the era loop agree. Deployed (complete) models keep their watermark: for them dtStudied
|
|
|
|
|
//--- gates INFERENCE recency, and this scan must not touch it.
|
|
|
|
|
if(!m_trainingComplete)
|
|
|
|
|
dtStudied = TrainWindowStart(m_tuneStartTrainBar);
|
refactor: split CExpertSignalAIBase implementation by responsibility
ExpertSignalAIBase.mqh was 8216 lines: the class declaration followed by 87
method bodies covering training, labelling, feature extraction, persistence,
chart drawing, online learning, the GA auto-tuner and inference, all in one
file. Train() alone is 1492 lines; a change to arrow drawing meant scrolling
past the era loop.
Moved the bodies into Expert\AIBase\, included at the bottom of the original
after the class declaration:
Training.mqh 1607 era loop, plateau ladder, checkpoint select, deploy
Features.mqh 1093 indicator creation + per-bar input feature vector
ChartUI.mqh 634 arrows, arrow persistence, status panel, cleanup
Persistence.mqh 492 .stats/.cfg sidecars, CPU-inference validation, copy
OnlineLearning.mqh 461 live continual learning, EMA shadow, OOS simulator
Labels.mqh 309 ZigZag pivot labels, async label-cache prebuild
AutoTune.mqh 275 genetic tuner (population, crossover, halving)
Inference.mqh 235 softmax, prior calibration, class priors
ExpertSignalAIBase.mqh 8216 -> 3131 (declaration + topology build only)
This is a pure relocation - verified mechanically, not by eye: HEAD's file
reconstructed from the eight partials plus the surviving remainder is
byte-identical to HEAD, span for span (scratchpad verify_split.py). No
declaration moved, no signature changed, no code rewritten, so behaviour is
unchanged by construction.
Compiles 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 00:42:45 -04:00
|
|
|
int barsNow = (int)MathMin(Bars(m_symbol.Name(), PERIOD_CURRENT, dtStudied, TimeCurrent()) + m_historyBars, Bars(m_symbol.Name(), PERIOD_CURRENT));
|
fix(depth): route EVERY ResizeBuffers call site through one indicator-depth gate
1dda479 clamped the training sweep. It left five other paths asking the indicators
for a depth they cannot serve, and on a live account the quiet ones are worse than
the stall was - a stalled chart is visible, a chart trading on a degraded feature
window is not.
ServableBars(want, context) is now the single gate, and all six go through it:
training sweep clamp, floored at TRAIN_MIN_CLAMPED_BARS (below that a small
positive BarsCalculated is warm-up, which m_coldSweepTick owns)
label prebuild clamp - labels come from price/ADZigZag and would survive a
capped MA, but ResizeBuffers sizes EVERY buffer and a failed
CopyBuffer leaves m_MA EMPTY for the next reader, so this path
could silently re-break the block Train()'s clamp just fixed
live inference HOLD. Below `need` the swing block takes its degraded path and
inference runs on a different feature distribution than the model
was fitted on. This EA sizes real positions off that output, so
no signal beats a mismatched one
online learning HOLD, same reason and worse - this path WRITES to a live trading
model, so a mismatched (features, label) pair is not a wrong arrow,
it is a wrong weight update that compounds every bar
chart rescan clamp - SIGNAL_RESCAN_LOOKBACK_BARS is 5000 and MT5's smallest
"Max bars in chart" is also 5000, so this one is genuinely
reachable; uncapped it repaints the window all-Neutral
research export clamp before the emptiness test, so a capped symbol exports the
depth it has rather than writing a CSV with a dead feature block -
an artefact that looks complete and is silently wrong
Both HOLDs are insurance, not expected states: `need` tops out near 1,152 bars
(16 + 750 + 384 + 2) against a 5,000 floor on the terminal setting. They exist so
the failure mode is unreachable rather than merely unlikely.
Not changed: a genuinely SHORT price history still takes the old degraded path at
every site. That is pre-existing behaviour and narrowing it would mute charts that
trade today, so it stays a separate decision rather than a side effect of this fix.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 13:33:16 -04:00
|
|
|
//--- Clamped for TWO reasons, only one of which is about labels (see ServableBars()). The labels
|
|
|
|
|
//--- themselves come from price and ADZigZag and would survive a capped MA - but ResizeBuffers()
|
|
|
|
|
//--- sizes EVERY buffer, including m_MA, and a RefreshData() whose CopyBuffer fails leaves that
|
|
|
|
|
//--- buffer EMPTY for whoever reads it next. So an unclamped prebuild here would re-break the very
|
|
|
|
|
//--- feature block Train()'s clamp just repaired, from a path that looks unrelated to it.
|
|
|
|
|
//--- The second reason is grid agreement: a label cache deeper than the sweep that consumes it is
|
|
|
|
|
//--- history training can never reach anyway, and it made labelCacheBars (50,179 in the 2026-08-17
|
|
|
|
|
//--- stall reports) disagree with the depth actually studied.
|
feat(depth): prime -> settle -> sweep, and name which handle is short
"Max bars in chart" is set to Unlimited, so the static-terminal-limit reading in
1dda479 was wrong. Two other candidate causes are falsified too: the price series
is fully downloaded and flat (USDJPY 50,162 -> 50,163 over 80 minutes, i.e. one new
H4 bar), and the handles have been stable since 13:07 with zero windows for the 17
minutes after, so it is not download-in-progress and not handle churn. The MA period
tops out at 200 (ADIndicatorTuner MA_PERIOD_PRESETS) against ~50k bars, so it is not
indicator cost either.
What IS verified stays verified: CopyBuffer past the calculated depth fails outright
rather than short-reading, so the buffer holds nothing and every index reads
EMPTY_VALUE; m_MA is the only CiCustom whose block REJECTS on that (m_ADZigZag
neutral-fills, RSI/MACD/Ichimoku/ATR are built-ins); the wall is therefore feature 25
of every bar, exactly as the "24 of 832" stall lines said. And it is depth-correlated:
16k-bar charts train, 34k/50k get zero windows forever.
So the WHY is still open, and this fix does not depend on it. Per the user's protocol:
the request itself is the primer, so prime at full depth, then poll TunableBarsCalculated()
every 3s and hold the sweep until it stops changing (3 steady probes), then use whatever
it settled at. Bounded at 10 min, and a give-up is logged as a give-up so an abandoned
depth is never mistaken for a settled one. This supersedes 1dda479's clamp on the two
training paths, which snapshotted a value that may still have been climbing; the clamp
remains for the paths that cannot wait (inference/online/rescan/export, see 7e63a8b).
The load-bearing part is what does NOT happen while waiting: no sweep. A 50k-bar feature
scan starves the indicator threads the request just woke, which is how the failure
sustained itself for 40 minutes at a time - discard era, re-sweep, discard, which is the
0->100% oscillation on the panel.
Also adds IndicatorDepthReport(): per-handle BarsCalculated() on the priming, cap and
stall lines. The logs proved WHICH FEATURE died but never WHICH HANDLE was short, so the
cause had to be inferred - and was guessed wrong twice. The next occurrence reads it off.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 13:44:27 -04:00
|
|
|
//--- Prime first (this call IS the request that makes the terminal calculate that deep), then settle.
|
refactor: split CExpertSignalAIBase implementation by responsibility
ExpertSignalAIBase.mqh was 8216 lines: the class declaration followed by 87
method bodies covering training, labelling, feature extraction, persistence,
chart drawing, online learning, the GA auto-tuner and inference, all in one
file. Train() alone is 1492 lines; a change to arrow drawing meant scrolling
past the era loop.
Moved the bodies into Expert\AIBase\, included at the bottom of the original
after the class declaration:
Training.mqh 1607 era loop, plateau ladder, checkpoint select, deploy
Features.mqh 1093 indicator creation + per-bar input feature vector
ChartUI.mqh 634 arrows, arrow persistence, status panel, cleanup
Persistence.mqh 492 .stats/.cfg sidecars, CPU-inference validation, copy
OnlineLearning.mqh 461 live continual learning, EMA shadow, OOS simulator
Labels.mqh 309 ZigZag pivot labels, async label-cache prebuild
AutoTune.mqh 275 genetic tuner (population, crossover, halving)
Inference.mqh 235 softmax, prior calibration, class priors
ExpertSignalAIBase.mqh 8216 -> 3131 (declaration + topology build only)
This is a pure relocation - verified mechanically, not by eye: HEAD's file
reconstructed from the eight partials plus the surviving remainder is
byte-identical to HEAD, span for span (scratchpad verify_split.py). No
declaration moved, no signature changed, no code rewritten, so behaviour is
unchanged by construction.
Compiles 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 00:42:45 -04:00
|
|
|
if(!ResizeBuffers(barsNow) || !RefreshData())
|
fix(buffers): revert the MA +1 - it asked for a bar that does not exist and stopped every chart
REGRESSION I INTRODUCED IN 1dda479, live for ~15 minutes.
CSeries::BufferResize -> CheckLoadHistory -> CheckTerminalHistory succeeds only when
Bars() >= size. Train() calls ResizeBuffers with barIndex == Bars(), so sizing the MA
buffer to barIndex + 1 asks for one bar more than the symbol has and fails the WHOLE
ResizeBuffers call. The log named it exactly:
failed to get 50180 bars for USDJPY,PERIOD_H4 (Bars() = 50,179)
failed to get 33983 bars for XAUUSD,PERIOD_H4 (Bars() = 33,982)
StartLabelCachePrebuild() then bailed on the false return and stayed silent, so the
only symptom was Train() reporting "arming the first label-cache prebuild" forever
with labelCacheBars=0 - the panel's "getting ready".
The premise was wrong, not just the arithmetic. The MA block reads GetData(idx) AND
GetData(idx + 1), and at the OLDEST bar that second read is SUPPOSED to fail - there
is no older bar to difference against. Rejecting that one bar is correct behaviour;
buying it cost the entire history.
Two more things, since the same defect had a second instance and no alarm:
- The Ichimoku pair (closeBars and m_Ichimoku, both barIndex + ichiKijun) is the same
bug with a far larger constant, latent only because the feature is off. Both are now
clamped to Bars(). The oldest ichiKijun bars then have no cloud, which that block's
EMPTY_VALUE guard already handles per-bar - the right outcome.
- The prebuild's bare `return` on a false ResizeBuffers now says so once, naming the
depth and Bars(). MQL5's own "failed to get N bars" was in the log the entire time,
from a stack frame nothing connected to the prebuild.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 14:12:47 -04:00
|
|
|
{
|
|
|
|
|
//--- NEVER SILENT AGAIN. This return used to be bare, and on 2026-08-17 it swallowed a hard
|
|
|
|
|
//--- failure for as long as the chart was open: a buffer was being sized one bar past Bars(),
|
|
|
|
|
//--- CheckLoadHistory refused, ResizeBuffers returned false, and the ONLY visible symptom was
|
|
|
|
|
//--- Train() reporting "arming the first label-cache prebuild" forever with labelCacheBars=0.
|
|
|
|
|
//--- MQL5's own "failed to get N bars" line was in the log the whole time and belonged to a
|
|
|
|
|
//--- stack frame nothing connected to the prebuild. Say which depth, and say it is fatal here.
|
|
|
|
|
if(!m_prebuildBlockWarned)
|
|
|
|
|
{
|
|
|
|
|
m_prebuildBlockWarned = true;
|
|
|
|
|
PrintFormat("%s: label prebuild BLOCKED - buffers would not prepare for %d bars. If MQL5"
|
|
|
|
|
" printed 'failed to get %d bars' just above, a buffer is being sized beyond the"
|
|
|
|
|
" %d bars this symbol actually has, and no era can start until that is fixed.",
|
|
|
|
|
ID, barsNow, barsNow, Bars(m_symbol.Name(), PERIOD_CURRENT));
|
|
|
|
|
}
|
|
|
|
|
return; // m_labelCachePrebuilt stays false, retried next call
|
|
|
|
|
}
|
feat(depth): prime -> settle -> sweep, and name which handle is short
"Max bars in chart" is set to Unlimited, so the static-terminal-limit reading in
1dda479 was wrong. Two other candidate causes are falsified too: the price series
is fully downloaded and flat (USDJPY 50,162 -> 50,163 over 80 minutes, i.e. one new
H4 bar), and the handles have been stable since 13:07 with zero windows for the 17
minutes after, so it is not download-in-progress and not handle churn. The MA period
tops out at 200 (ADIndicatorTuner MA_PERIOD_PRESETS) against ~50k bars, so it is not
indicator cost either.
What IS verified stays verified: CopyBuffer past the calculated depth fails outright
rather than short-reading, so the buffer holds nothing and every index reads
EMPTY_VALUE; m_MA is the only CiCustom whose block REJECTS on that (m_ADZigZag
neutral-fills, RSI/MACD/Ichimoku/ATR are built-ins); the wall is therefore feature 25
of every bar, exactly as the "24 of 832" stall lines said. And it is depth-correlated:
16k-bar charts train, 34k/50k get zero windows forever.
So the WHY is still open, and this fix does not depend on it. Per the user's protocol:
the request itself is the primer, so prime at full depth, then poll TunableBarsCalculated()
every 3s and hold the sweep until it stops changing (3 steady probes), then use whatever
it settled at. Bounded at 10 min, and a give-up is logged as a give-up so an abandoned
depth is never mistaken for a settled one. This supersedes 1dda479's clamp on the two
training paths, which snapshotted a value that may still have been climbing; the clamp
remains for the paths that cannot wait (inference/online/rescan/export, see 7e63a8b).
The load-bearing part is what does NOT happen while waiting: no sweep. A 50k-bar feature
scan starves the indicator threads the request just woke, which is how the failure
sustained itself for 40 minutes at a time - discard era, re-sweep, discard, which is the
0->100% oscillation on the panel.
Also adds IndicatorDepthReport(): per-handle BarsCalculated() on the priming, cap and
stall lines. The logs proved WHICH FEATURE died but never WHICH HANDLE was short, so the
cause had to be inferred - and was guessed wrong twice. The next occurrence reads it off.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 13:44:27 -04:00
|
|
|
int settled = SettledBars(barsNow, "label prebuild");
|
|
|
|
|
if(settled <= 0)
|
|
|
|
|
return; // depth still moving - retried next call, same contract as the line above
|
|
|
|
|
if(settled < barsNow)
|
|
|
|
|
{
|
|
|
|
|
barsNow = settled;
|
|
|
|
|
if(!ResizeBuffers(barsNow) || !RefreshData())
|
|
|
|
|
return;
|
|
|
|
|
}
|
refactor: split CExpertSignalAIBase implementation by responsibility
ExpertSignalAIBase.mqh was 8216 lines: the class declaration followed by 87
method bodies covering training, labelling, feature extraction, persistence,
chart drawing, online learning, the GA auto-tuner and inference, all in one
file. Train() alone is 1492 lines; a change to arrow drawing meant scrolling
past the era loop.
Moved the bodies into Expert\AIBase\, included at the bottom of the original
after the class declaration:
Training.mqh 1607 era loop, plateau ladder, checkpoint select, deploy
Features.mqh 1093 indicator creation + per-bar input feature vector
ChartUI.mqh 634 arrows, arrow persistence, status panel, cleanup
Persistence.mqh 492 .stats/.cfg sidecars, CPU-inference validation, copy
OnlineLearning.mqh 461 live continual learning, EMA shadow, OOS simulator
Labels.mqh 309 ZigZag pivot labels, async label-cache prebuild
AutoTune.mqh 275 genetic tuner (population, crossover, halving)
Inference.mqh 235 softmax, prior calibration, class priors
ExpertSignalAIBase.mqh 8216 -> 3131 (declaration + topology build only)
This is a pure relocation - verified mechanically, not by eye: HEAD's file
reconstructed from the eight partials plus the surviving remainder is
byte-identical to HEAD, span for span (scratchpad verify_split.py). No
declaration moved, no signature changed, no code rewritten, so behaviour is
unchanged by construction.
Compiles 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 00:42:45 -04:00
|
|
|
EnsureBarCachesCapacity(barsNow);
|
feat(ai): triple-barrier labels replace exact-pivot ZigZag targets
The 31:1 class imbalance was self-inflicted by the TARGET, not a property
of the market. Labelling only the exact bar where a ZigZag pivot confirms
gave Buy 1164 / Sell 1164 / Neutral 35841, and every correction mechanism
this codebase accumulated sits downstream of that one choice: the
logit-adjusted loss and its range cap, the prior EMA, the +-3.0 output-bias
seed, balanced-accuracy-then-precision selection with its coverage floor,
the recall floor and its catch-22, the alternation gate, NMS, and the four
oversampling designs that collapsed before them.
The reference this engine is built on (references/neuronetworksbook.pdf
ch. 3.1/3.3) also uses ZigZag, but targets the DIRECTION TO THE NEXT
EXTREMUM on every bar - ~50/50 by construction, with no imbalance to
correct at all. It never had this problem because it never asked "is this
the pivot bar".
Labels are now the triple barrier (Lopez de Prado ch. 3), using the EA's
OWN SL_Mode/TP_Mode: does a trade opened at this bar's close reach its
target before its stop, within a horizon. Buy = long resolves, Sell =
short resolves, Neutral = neither. Consequences:
- dir-precision in the era line stops being a proxy and becomes the win
rate of the strategy under its own exit rules.
- Expected balance ~25/25/50 at the shipped 1:3 (gambler's ruin), i.e.
~2:1 instead of 31:1. Measured and logged at the end of the prebuild.
- Spread is charged on both legs, so it is a NET win rate.
- Intrabar ambiguity resolves to the STOP. OHLC cannot order two touches
inside one bar and the optimistic reading is how a backtested edge
becomes a live loss.
ZigZag stays as input features (EnableSwingContext) and now also supplies
the vertical barrier: the horizon is the median confirmed leg length,
snapped to a coarse ladder. Derived, not configured, and deliberately kept
out of the filename fingerprint - a filename keyed on a measured quantity
orphans a trained model the moment the measurement moves.
Removed, because the premise died with the old target:
- the alternation gate. Correct for pivot labels (a ZigZag cannot emit two
same-type pivots in a row, so a repeat was provably a false fire), and
wrong for barrier labels, which answer each bar independently. It also
took its worst consequence with it: a one-sided model previously got ONE
trade per backtest, a hard blocker on marketplace validation.
- SignalClusterWindow now defaults off - it de-duplicated repeats that are
now real trades. Kept as an opt-in display control.
- LABEL_WINDOW_BARS, the pivot-widening pass, ConfirmedZigZagLabel.
- the era-0 output-bias seed now needs a genuinely dominant class (0.70)
rather than 0.40; at ~50% Neutral a +-3.0 seed is a distortion, not a
correction.
Also fixed, both found while wiring the above:
1. RefreshConvergedSignal sized its buffers from a date delta
(Bars(sym, period, dtStudied, TimeCurrent())). dtStudied is a training
watermark; in the tester it is loaded from a live-chart save AHEAD of
the simulated date, so the interval inverted, Bars() returned ~0, and
the buffer came out at exactly m_historyBars - deep enough for the OHLC
window and far too shallow for the Donchian-50 / 20-bar-return / SMA
extension behind it. Inference silently computed DIFFERENT features
from the ones training learned on, live as well as in the tester. Now
sized from what the feature builder actually needs.
2. The barrier horizon is resolved on the deployed path too. A deployed
model never enters Train(), so it never reached the prebuild, and
OnlineLearnStep reads the horizon as its confirmation delay - left at
the fallback it would have backpropped bars whose barriers had not
resolved. Silent lookahead in the one place that writes to a live model.
SL_Mode/TP_Mode join the weights fingerprint: they define the labels now,
so a model trained at 1:3 must never be silently reused at 1:1. This
re-keys every pre-existing model by design - none were trained on this task.
Inference census extended with the vote gate. LongCondition/ShortCondition
open with a readiness check the refresh counters never see; in the tester it
reduces to "the seeded _optcache.nnw must have LOADED", and if it did not,
every vote is hard-zeroed while the model still answers Buy. The old three
counters would have read that as "the model says Neutral" - false, and a
completely different fix. This is the leading candidate for the
zero-direction backtest and the census can now name it in one run.
Both builds compile 0 errors / 0 warnings. Forces a full retrain.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 20:39:49 -04:00
|
|
|
//--- Settle the vertical barrier BEFORE the first label is computed. Derived once per process and then
|
|
|
|
|
//--- held: AdvanceBarrierLabelState() indexes off it, so a value that moved mid-scan would leave the
|
|
|
|
|
//--- cache holding labels from two different rules.
|
|
|
|
|
EnsureBarrierHorizon(barsNow);
|
refactor: split CExpertSignalAIBase implementation by responsibility
ExpertSignalAIBase.mqh was 8216 lines: the class declaration followed by 87
method bodies covering training, labelling, feature extraction, persistence,
chart drawing, online learning, the GA auto-tuner and inference, all in one
file. Train() alone is 1492 lines; a change to arrow drawing meant scrolling
past the era loop.
Moved the bodies into Expert\AIBase\, included at the bottom of the original
after the class declaration:
Training.mqh 1607 era loop, plateau ladder, checkpoint select, deploy
Features.mqh 1093 indicator creation + per-bar input feature vector
ChartUI.mqh 634 arrows, arrow persistence, status panel, cleanup
Persistence.mqh 492 .stats/.cfg sidecars, CPU-inference validation, copy
OnlineLearning.mqh 461 live continual learning, EMA shadow, OOS simulator
Labels.mqh 309 ZigZag pivot labels, async label-cache prebuild
AutoTune.mqh 275 genetic tuner (population, crossover, halving)
Inference.mqh 235 softmax, prior calibration, class priors
ExpertSignalAIBase.mqh 8216 -> 3131 (declaration + topology build only)
This is a pure relocation - verified mechanically, not by eye: HEAD's file
reconstructed from the eight partials plus the surviving remainder is
byte-identical to HEAD, span for span (scratchpad verify_split.py). No
declaration moved, no signature changed, no code rewritten, so behaviour is
unchanged by construction.
Compiles 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 00:42:45 -04:00
|
|
|
int totalIter = (int)MathMax(barsNow - MathMax(m_historyBars, 0), 0);
|
|
|
|
|
m_labelPrebuildBars = barsNow;
|
|
|
|
|
m_labelPrebuildOosCutoff = (int)(MathMax(0, MathMin(100, m_oosSplitPct)) / 100.0 * totalIter);
|
|
|
|
|
m_labelPrebuildIndex = (int)(barsNow - MathMax(m_historyBars, 0) - 1);
|
|
|
|
|
m_labelPrebuildBuyCount = 0;
|
|
|
|
|
m_labelPrebuildSellCount = 0;
|
|
|
|
|
m_labelPrebuildNeutralCount = 0;
|
fix: refuse invalid SL/TP, fix the unreachable deploy floor, scale the horizon
Three defects found by reading the 2026-08-01 training logs, all of which
only became visible because the relabel made the numbers mean something.
1. A STALE ENUM TRAINED FOUR MODELS ON THE WRONG TARGET.
`OnInit: trade settings snapshot - SL_Mode=1 TP_Mode=-101`
-101 was TP_PREV_SWING, deleted from TAKE_PROFIT_MODE on 2026-07-31 in
7eb48f5. MetaTrader does not validate a saved enum input against the
enum's current members, so charts saved before that kept the old
integer. BarrierMultiples()'s `if(tpMult <= 0.0) tpMult = slMult;`
then quietly turned it into a 1:1 barrier, and all four topologies
trained ~250 eras against a strategy nobody selected - while the log
reported "target 1.00*ATR" as though it were configured.
Since the relabel these two inputs ARE the label definition, so this
is not a bad trade setting, it is a wrong dataset. ValidateBarrier-
Inputs() now refuses to start (INIT_FAILED + Alert + an explicit fix)
on any value that is not an enum member. Members are enumerated rather
than range-checked because both enums are sparse and carry negative
sentinels, so no min/max test can tell a legal value from a deleted
one - which is the entire failure mode. The fallback survives as
belt-and-braces but now announces itself: a fallback that cannot say
it fired is indistinguishable from correct behaviour.
2. THE DEPLOYABILITY FLOOR BECAME MATHEMATICALLY UNREACHABLE.
`tradeableOK` required `dirPrecPct >= baseRatePct`, where baseRatePct
is Buy+Sell as a share of all bars. At the old exact-pivot target that
was ~6%, so "beat the base rate" read as "beat chance" and the test
looked sound. Triple-barrier labels put it at ~83%, so the gate now
demanded 83% directional precision - impossible by construction.
Observed live: all four topologies cycling "PLATEAU stage 3 ... nothing
safe to deploy" at a perfectly healthy 43-45% precision, with no
checkpoint able to ship however good it got.
Replaced with ZERO-SKILL precision, max(Buy,Sell)/allBars: exactly the
score of the degenerate always-call-one-direction model this floor
exists to reject. Correct at any base rate - ~43% on the current
labels, ~3% on the old rare-pivot ones. The era line now prints
"(chance N%, edge +Mpp)" beside the selection score, because 44%
precision is excellent against a 3% chance level and worthless against
a 43% one, and reading the first as the second is what made tonight's
run look better than it was.
3. THE HORIZON IGNORED THE BARRIER GEOMETRY.
ComputeBarrierHorizonBars() returned the median ZigZag leg, which
measures how long a ~1 ATR move takes and says nothing about how long
the CONFIGURED barrier needs. First-passage time out of [-m,+k] scales
with m*k, so a 1:3 barrier takes ~3x as long as 1:1; the unscaled
horizon would have timed out most 1:3 trades and pushed Neutral
straight back up, re-creating the imbalance the relabel removes.
Now multiplied by slMult*tpMult, calibrated against a real measurement
rather than assumed: the accidental 1:1 run resolved at horizon 12 with
only 16.7% timeouts, so the swing median is the right scale at m*k=1.
Verifiable, not just asserted: the prebuild now counts barriers that
ended on the VERTICAL barrier and reports them as a share of Neutral.
Neutral conflates "timed out" with "stopped out" and only the first
indicts the horizon.
Both builds compile 0 errors / 0 warnings. Forces a retrain - correcting
TP_Mode re-keys the fingerprint (|TB:1:-101 -> |TB:1:3), which is right:
no existing model was trained on the intended target.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 00:30:49 -04:00
|
|
|
m_labelPrebuildTimeoutCount = 0;
|
fix: both-won bars were labelled "do not trade" - resolve by first touch
Removing the min-reward:risk raise let the MEASURED geometry come back with
the target NEARER than the stop (SP500 H1: target 1.62*ATR at q50 of
favourable, stop 3.33*ATR at q75 of adverse). That reopened a branch the
code called unreachable: price can reach +target and -target inside one
horizon, winning in BOTH directions, and those bars fell through to Neutral.
Neutral has only three producers, both-lost is unreachable (you cannot touch
-3.33 without crossing -1.62 first, which wins the short), and timeouts logged
at 1.0% of Neutral - so ~27% of ALL bars were being handed to the model as the
abstain class when a trade either way would have collected its target. The
cleanest positives in the sample, labelled "do not trade", while the fitted
confidence threshold was being asked to find selectivity in what was left.
Resolved by FIRST TOUCH: the target reached earlier is the trade that would
have closed first. Same forward window, no extra lookahead. Same-bar ties stay
Neutral - OHLC cannot order two touches, and unlike an intrabar stop tie there
is no pessimistic side to fall to, so a guess would inject a coin-flip
direction into the target.
Also:
- count both-won and its same-bar tie subset in the prebuild line, so the
share is measured rather than inferred from arithmetic on a log line
- scope the timeout counter to IS, matching the tally it is reported as a
percentage OF; it was incremented over the whole scan and divided by an
in-sample denominator
- clear m_lastBarrierTimedOut at the top of the walk with the excursions, not
at the bottom - the two early returns published the previous bar's verdict
- mark the pass-1 label line PROVISIONAL. It prints the enum fallback because
geometry can only be derived from excursions that do not exist yet, and it
reads exactly like a config change that failed to take effect
FORCES RETRAIN. Both build variants compile 0 errors, 0 warnings.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 15:35:42 -04:00
|
|
|
m_labelPrebuildBothWonCount = 0;
|
|
|
|
|
m_labelPrebuildBothWonTieCount = 0;
|
fix(labels): overlapping-label sample correction + horizon cap on the scale ladder
Three defects, all surfaced by the 2026-08-17 SP500 H4 run that shipped
stop 4.86 / target 9.71 (width 14.57*ATR, horizon 384).
1. EVERY STANDARD ERROR ASSUMED INDEPENDENT SAMPLES. Triple-barrier labels
started one per bar overlap by the label's lifespan, so n calls are worth
~n/L independent observations (Lopez de Prado, AFML ch. 4 - sample
uniqueness). All three sqrt(p(1-p)/n) sites divided by the RAW count.
The tell: the operating point's null-of-the-maximum gate is family-wise and
should fire on ~5% of eras under the null. Measured fire rates - PAI 47/73
(64%), ConvLSTM 9/24, LSTM 8/21 (38%), CONV 4/62 (6%). CONV, the only model
whose margin distribution admits few bins, sat on the null; the rest cleared
a bar that was too low by ~sqrt(L). PAI's deployed threshold consequently
alternated between the ENDS of its own range era to era (0.10 -> 0.88 ->
0.86 -> 0.66; coverage 16% <-> 73%).
TripleBarrierLabel now records when each label became KNOWABLE - the first
winning touch, or both stops, or the timeout - and the prebuild accumulates
the mean. EffectiveSampleSize() feeds the operating point, the member deploy
gate and the ensemble vote gate. Conservative by construction (n/L is an
upper bound on the damage); gates get harder, never easier.
2. THE SCALE LADDER RAN AWAY, again. Horizon scales as swingMedian*sl*tp, and
since 4d8cb08 reachability is measured OVER that horizon - so a wider rung
buys itself the time that makes it look reachable. Same target -> horizon ->
reach -> target loop the excursion window is kept short to avoid; fixing the
window confusion reopened it through the other door. It walked 128 -> 256 ->
384 bars and stopped at q90, the widest rung there is, with every rung
reading 39-48% against a 20% floor. A floor nothing fails selects nothing.
Rungs whose required horizon exceeds BARRIER_HORIZON_MAX are now rejected -
the same rule ReportGeometryExpectancyScan already applied. It was printing
the shipped pair as CLAMPED and disqualified ('h384!') two lines under the
deriver that chose it: two subsystems, one geometry, opposite verdicts.
3. THE RUNG SNAP DESTROYED THE RATIO IT WAS COMPARING. Both legs snapped
independently to the coarse first-passage grid, re-rating each candidate:
q90 4.86/9.71 -> 5.00/10.00 (2.00), q85 4.07/8.14 -> 5.00/10.00 (IDENTICAL
measurement), q75 3.07/6.13 -> 4.00/6.50 (1.63 - a nearer target). So the
ladder compared win shares taken at ratios from 1.63 to 2.17 and read the
differences as scale. It is why the reach column came out non-monotone in
width (q75 48.5% above q90 42.9%). The stop now snaps to its nearest rung in
log space and the target follows the ratio off it; the pair actually measured
is returned and logged, so a collision reads as a collision.
Also: LadderWinShare guarded against the conditional (fractal) geometry path,
which fills n from m_fracLegCount while leaving idxList empty - a latent
out-of-bounds on a currently-dead path.
New log lines: mean label lifespan and effective n on the label-cache line, the
required-vs-available horizon per rung, and the grid pair the reconciliation
actually measured (its tolerance now scales with the grid skew instead of a flat
5pp).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 01:12:05 -04:00
|
|
|
//--- Reset WITH the cache, not once per process: a rebuild follows a geometry or horizon change, and
|
|
|
|
|
//--- lifespans measured under the old barrier answer a different question. Carrying them forward would
|
|
|
|
|
//--- deflate the new geometry's standard errors by the old geometry's overlap.
|
|
|
|
|
m_labelLifespanSum = 0.0;
|
|
|
|
|
m_labelLifespanCount = 0;
|
refactor: split CExpertSignalAIBase implementation by responsibility
ExpertSignalAIBase.mqh was 8216 lines: the class declaration followed by 87
method bodies covering training, labelling, feature extraction, persistence,
chart drawing, online learning, the GA auto-tuner and inference, all in one
file. Train() alone is 1492 lines; a change to arrow drawing meant scrolling
past the era loop.
Moved the bodies into Expert\AIBase\, included at the bottom of the original
after the class declaration:
Training.mqh 1607 era loop, plateau ladder, checkpoint select, deploy
Features.mqh 1093 indicator creation + per-bar input feature vector
ChartUI.mqh 634 arrows, arrow persistence, status panel, cleanup
Persistence.mqh 492 .stats/.cfg sidecars, CPU-inference validation, copy
OnlineLearning.mqh 461 live continual learning, EMA shadow, OOS simulator
Labels.mqh 309 ZigZag pivot labels, async label-cache prebuild
AutoTune.mqh 275 genetic tuner (population, crossover, halving)
Inference.mqh 235 softmax, prior calibration, class priors
ExpertSignalAIBase.mqh 8216 -> 3131 (declaration + topology build only)
This is a pure relocation - verified mechanically, not by eye: HEAD's file
reconstructed from the eight partials plus the surviving remainder is
byte-identical to HEAD, span for span (scratchpad verify_split.py). No
declaration moved, no signature changed, no code rewritten, so behaviour is
unchanged by construction.
Compiles 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 00:42:45 -04:00
|
|
|
m_labelPrebuildActive = true;
|
|
|
|
|
}
|
|
|
|
|
//+------------------------------------------------------------------+
|
|
|
|
|
//| Advances the eager label-cache pre-build by up to a time budget, |
|
|
|
|
|
//| then yields (same chunking pattern as the era loop). Mirrors the |
|
|
|
|
|
//| era loop's own labeling eligibility gate (minus the dPrevSignal |
|
|
|
|
|
//| check, meaningless pre-first-feedForward). On completion, seeds |
|
|
|
|
|
//| m_prevEraTrueBuyCount/Sell/Neutral from the upfront IS-only tally |
|
2026-08-01 11:27:28 -04:00
|
|
|
//| so era 0's class priors are measured, not empty. |
|
refactor: split CExpertSignalAIBase implementation by responsibility
ExpertSignalAIBase.mqh was 8216 lines: the class declaration followed by 87
method bodies covering training, labelling, feature extraction, persistence,
chart drawing, online learning, the GA auto-tuner and inference, all in one
file. Train() alone is 1492 lines; a change to arrow drawing meant scrolling
past the era loop.
Moved the bodies into Expert\AIBase\, included at the bottom of the original
after the class declaration:
Training.mqh 1607 era loop, plateau ladder, checkpoint select, deploy
Features.mqh 1093 indicator creation + per-bar input feature vector
ChartUI.mqh 634 arrows, arrow persistence, status panel, cleanup
Persistence.mqh 492 .stats/.cfg sidecars, CPU-inference validation, copy
OnlineLearning.mqh 461 live continual learning, EMA shadow, OOS simulator
Labels.mqh 309 ZigZag pivot labels, async label-cache prebuild
AutoTune.mqh 275 genetic tuner (population, crossover, halving)
Inference.mqh 235 softmax, prior calibration, class priors
ExpertSignalAIBase.mqh 8216 -> 3131 (declaration + topology build only)
This is a pure relocation - verified mechanically, not by eye: HEAD's file
reconstructed from the eight partials plus the surviving remainder is
byte-identical to HEAD, span for span (scratchpad verify_split.py). No
declaration moved, no signature changed, no code rewritten, so behaviour is
unchanged by construction.
Compiles 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 00:42:45 -04:00
|
|
|
//+------------------------------------------------------------------+
|
|
|
|
|
void CExpertSignalAIBase::AdvanceLabelCachePrebuild(void)
|
|
|
|
|
{
|
|
|
|
|
const uint PREBUILD_TIME_BUDGET_MS = 80;
|
|
|
|
|
uint chunkStartTick = GetTickCount();
|
|
|
|
|
int i;
|
|
|
|
|
for(i = m_labelPrebuildIndex; i >= 2; i--)
|
|
|
|
|
{
|
fix(chart): a purge that reports "zero leftovers" was only ever checking its own list
2026-08-17 21:58: all three charts hit "Abnormal termination" ~5.3 s into
OnDeinit with NO cleanup-timings line - the teardown was starved again. The
22:00 init purge then removed 993 / 1373 / 1557 stranded objects and reported
ZERO by-name leftovers on every chart, and the charts still came up with
duplicated panels. "Nothing matching our prefixes remains" and "the chart is
clean" are different statements and only the first was being made.
Three changes, in the order they matter:
1. WHY the teardown starved, and it is a gap in ad80e0b. StartLabelCachePrebuild
runs ResizeBuffers + RefreshData over the FULL study window (33,984 bars on
XAUUSD), unchunked, and OnDeinit cannot begin until it returns. Normally a
once-per-run cost. That night SP500 and XAUUSD LSTM were wedged in the "cache
invalidated at era start" loop, which calls it on EVERY Train() call - two
members re-preparing tens of thousands of bars indefinitely. The terminal
closed into that. Guarded now, plus a resumable guard in the prebuild chunk
loop (the tally pass after it is not chunked).
2. Catch-all "Warrior" prefix in WarriorChartPrefixes. Every family this EA
creates is named Warrior* except the arrows (WarSig_), so one bare prefix
covers the three named entries AND anything a rename or a stale .ex5 left
under a name nobody remembers. Still a prefix delete, never
ObjectsDeleteAll(chart) - the user's own drawings are not ours to remove. Does
not defeat skipArrows: "WarSig_" does not start with "Warrior".
3. The init purge now REPORTS the residue it did not claim, by name (up to 12).
Not deleted - an unmatched object may belong to the user or another indicator.
If a Warrior panel is visible and appears in neither the removed count nor
this list, the prefix list has drifted a third time and the name is in the
journal instead of being inferred from a screenshot.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 22:07:09 -04:00
|
|
|
//--- Already chunked at 80 ms, so this costs at most one chunk - but the tally pass below is NOT
|
|
|
|
|
//--- chunked, and on a stop there is no reason to walk the rest of the window to reach it.
|
|
|
|
|
//--- Resumable by construction: m_labelPrebuildIndex is written before returning either way.
|
|
|
|
|
if(ShutdownRequested())
|
|
|
|
|
{
|
|
|
|
|
m_labelPrebuildIndex = i;
|
|
|
|
|
return;
|
|
|
|
|
}
|
refactor: split CExpertSignalAIBase implementation by responsibility
ExpertSignalAIBase.mqh was 8216 lines: the class declaration followed by 87
method bodies covering training, labelling, feature extraction, persistence,
chart drawing, online learning, the GA auto-tuner and inference, all in one
file. Train() alone is 1492 lines; a change to arrow drawing meant scrolling
past the era loop.
Moved the bodies into Expert\AIBase\, included at the bottom of the original
after the class declaration:
Training.mqh 1607 era loop, plateau ladder, checkpoint select, deploy
Features.mqh 1093 indicator creation + per-bar input feature vector
ChartUI.mqh 634 arrows, arrow persistence, status panel, cleanup
Persistence.mqh 492 .stats/.cfg sidecars, CPU-inference validation, copy
OnlineLearning.mqh 461 live continual learning, EMA shadow, OOS simulator
Labels.mqh 309 ZigZag pivot labels, async label-cache prebuild
AutoTune.mqh 275 genetic tuner (population, crossover, halving)
Inference.mqh 235 softmax, prior calibration, class priors
ExpertSignalAIBase.mqh 8216 -> 3131 (declaration + topology build only)
This is a pure relocation - verified mechanically, not by eye: HEAD's file
reconstructed from the eight partials plus the surviving remainder is
byte-identical to HEAD, span for span (scratchpad verify_split.py). No
declaration moved, no signature changed, no code rewritten, so behaviour is
unchanged by construction.
Compiles 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 00:42:45 -04:00
|
|
|
if(GetTickCount() - chunkStartTick >= PREBUILD_TIME_BUDGET_MS)
|
|
|
|
|
{
|
|
|
|
|
m_labelPrebuildIndex = i;
|
|
|
|
|
return;
|
|
|
|
|
}
|
|
|
|
|
if(!(i < (int)(m_labelPrebuildBars - MathMax(m_historyBars, 0) - 1) && m_Time.GetData(i) > dtStudied))
|
|
|
|
|
continue;
|
feat(ai): triple-barrier labels replace exact-pivot ZigZag targets
The 31:1 class imbalance was self-inflicted by the TARGET, not a property
of the market. Labelling only the exact bar where a ZigZag pivot confirms
gave Buy 1164 / Sell 1164 / Neutral 35841, and every correction mechanism
this codebase accumulated sits downstream of that one choice: the
logit-adjusted loss and its range cap, the prior EMA, the +-3.0 output-bias
seed, balanced-accuracy-then-precision selection with its coverage floor,
the recall floor and its catch-22, the alternation gate, NMS, and the four
oversampling designs that collapsed before them.
The reference this engine is built on (references/neuronetworksbook.pdf
ch. 3.1/3.3) also uses ZigZag, but targets the DIRECTION TO THE NEXT
EXTREMUM on every bar - ~50/50 by construction, with no imbalance to
correct at all. It never had this problem because it never asked "is this
the pivot bar".
Labels are now the triple barrier (Lopez de Prado ch. 3), using the EA's
OWN SL_Mode/TP_Mode: does a trade opened at this bar's close reach its
target before its stop, within a horizon. Buy = long resolves, Sell =
short resolves, Neutral = neither. Consequences:
- dir-precision in the era line stops being a proxy and becomes the win
rate of the strategy under its own exit rules.
- Expected balance ~25/25/50 at the shipped 1:3 (gambler's ruin), i.e.
~2:1 instead of 31:1. Measured and logged at the end of the prebuild.
- Spread is charged on both legs, so it is a NET win rate.
- Intrabar ambiguity resolves to the STOP. OHLC cannot order two touches
inside one bar and the optimistic reading is how a backtested edge
becomes a live loss.
ZigZag stays as input features (EnableSwingContext) and now also supplies
the vertical barrier: the horizon is the median confirmed leg length,
snapped to a coarse ladder. Derived, not configured, and deliberately kept
out of the filename fingerprint - a filename keyed on a measured quantity
orphans a trained model the moment the measurement moves.
Removed, because the premise died with the old target:
- the alternation gate. Correct for pivot labels (a ZigZag cannot emit two
same-type pivots in a row, so a repeat was provably a false fire), and
wrong for barrier labels, which answer each bar independently. It also
took its worst consequence with it: a one-sided model previously got ONE
trade per backtest, a hard blocker on marketplace validation.
- SignalClusterWindow now defaults off - it de-duplicated repeats that are
now real trades. Kept as an opt-in display control.
- LABEL_WINDOW_BARS, the pivot-widening pass, ConfirmedZigZagLabel.
- the era-0 output-bias seed now needs a genuinely dominant class (0.70)
rather than 0.40; at ~50% Neutral a +-3.0 seed is a distortion, not a
correction.
Also fixed, both found while wiring the above:
1. RefreshConvergedSignal sized its buffers from a date delta
(Bars(sym, period, dtStudied, TimeCurrent())). dtStudied is a training
watermark; in the tester it is loaded from a live-chart save AHEAD of
the simulated date, so the interval inverted, Bars() returned ~0, and
the buffer came out at exactly m_historyBars - deep enough for the OHLC
window and far too shallow for the Donchian-50 / 20-bar-return / SMA
extension behind it. Inference silently computed DIFFERENT features
from the ones training learned on, live as well as in the tester. Now
sized from what the feature builder actually needs.
2. The barrier horizon is resolved on the deployed path too. A deployed
model never enters Train(), so it never reached the prebuild, and
OnlineLearnStep reads the horizon as its confirmation delay - left at
the fallback it would have backpropped bars whose barriers had not
resolved. Silent lookahead in the one place that writes to a live model.
SL_Mode/TP_Mode join the weights fingerprint: they define the labels now,
so a model trained at 1:3 must never be silently reused at 1:1. This
re-keys every pre-existing model by design - none were trained on this task.
Inference census extended with the vote gate. LongCondition/ShortCondition
open with a readiness check the refresh counters never see; in the tester it
reduces to "the seeded _optcache.nnw must have LOADED", and if it did not,
every vote is hard-zeroed while the model still answers Buy. The old three
counters would have read that as "the model says Neutral" - false, and a
completely different fix. This is the leading candidate for the
zero-direction backtest and the census can now name it in one run.
Both builds compile 0 errors / 0 warnings. Forces a full retrain.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 20:39:49 -04:00
|
|
|
// A barrier label needs m_barrierHorizonBars of FUTURE (lower-index) bars to resolve, so visiting
|
|
|
|
|
// bar i settles the label for bar i+horizon - see AdvanceBarrierLabelState(). Unlike the ZigZag
|
|
|
|
|
// scan this replaced, each verdict is final when written: the outcome depends only on price inside
|
|
|
|
|
// a fixed forward window, so no later iteration can revise it and there is no widening/re-spread
|
|
|
|
|
// pass to run afterwards. The tally still happens in one pass at the end, purely because the loop
|
|
|
|
|
// above is chunked across Train() calls and may resume mid-scan.
|
refactor: split CExpertSignalAIBase implementation by responsibility
ExpertSignalAIBase.mqh was 8216 lines: the class declaration followed by 87
method bodies covering training, labelling, feature extraction, persistence,
chart drawing, online learning, the GA auto-tuner and inference, all in one
file. Train() alone is 1492 lines; a change to arrow drawing meant scrolling
past the era loop.
Moved the bodies into Expert\AIBase\, included at the bottom of the original
after the class declaration:
Training.mqh 1607 era loop, plateau ladder, checkpoint select, deploy
Features.mqh 1093 indicator creation + per-bar input feature vector
ChartUI.mqh 634 arrows, arrow persistence, status panel, cleanup
Persistence.mqh 492 .stats/.cfg sidecars, CPU-inference validation, copy
OnlineLearning.mqh 461 live continual learning, EMA shadow, OOS simulator
Labels.mqh 309 ZigZag pivot labels, async label-cache prebuild
AutoTune.mqh 275 genetic tuner (population, crossover, halving)
Inference.mqh 235 softmax, prior calibration, class priors
ExpertSignalAIBase.mqh 8216 -> 3131 (declaration + topology build only)
This is a pure relocation - verified mechanically, not by eye: HEAD's file
reconstructed from the eight partials plus the surviving remainder is
byte-identical to HEAD, span for span (scratchpad verify_split.py). No
declaration moved, no signature changed, no code rewritten, so behaviour is
unchanged by construction.
Compiles 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 00:42:45 -04:00
|
|
|
if(!m_labelCacheHasValue[i])
|
feat(ai): triple-barrier labels replace exact-pivot ZigZag targets
The 31:1 class imbalance was self-inflicted by the TARGET, not a property
of the market. Labelling only the exact bar where a ZigZag pivot confirms
gave Buy 1164 / Sell 1164 / Neutral 35841, and every correction mechanism
this codebase accumulated sits downstream of that one choice: the
logit-adjusted loss and its range cap, the prior EMA, the +-3.0 output-bias
seed, balanced-accuracy-then-precision selection with its coverage floor,
the recall floor and its catch-22, the alternation gate, NMS, and the four
oversampling designs that collapsed before them.
The reference this engine is built on (references/neuronetworksbook.pdf
ch. 3.1/3.3) also uses ZigZag, but targets the DIRECTION TO THE NEXT
EXTREMUM on every bar - ~50/50 by construction, with no imbalance to
correct at all. It never had this problem because it never asked "is this
the pivot bar".
Labels are now the triple barrier (Lopez de Prado ch. 3), using the EA's
OWN SL_Mode/TP_Mode: does a trade opened at this bar's close reach its
target before its stop, within a horizon. Buy = long resolves, Sell =
short resolves, Neutral = neither. Consequences:
- dir-precision in the era line stops being a proxy and becomes the win
rate of the strategy under its own exit rules.
- Expected balance ~25/25/50 at the shipped 1:3 (gambler's ruin), i.e.
~2:1 instead of 31:1. Measured and logged at the end of the prebuild.
- Spread is charged on both legs, so it is a NET win rate.
- Intrabar ambiguity resolves to the STOP. OHLC cannot order two touches
inside one bar and the optimistic reading is how a backtested edge
becomes a live loss.
ZigZag stays as input features (EnableSwingContext) and now also supplies
the vertical barrier: the horizon is the median confirmed leg length,
snapped to a coarse ladder. Derived, not configured, and deliberately kept
out of the filename fingerprint - a filename keyed on a measured quantity
orphans a trained model the moment the measurement moves.
Removed, because the premise died with the old target:
- the alternation gate. Correct for pivot labels (a ZigZag cannot emit two
same-type pivots in a row, so a repeat was provably a false fire), and
wrong for barrier labels, which answer each bar independently. It also
took its worst consequence with it: a one-sided model previously got ONE
trade per backtest, a hard blocker on marketplace validation.
- SignalClusterWindow now defaults off - it de-duplicated repeats that are
now real trades. Kept as an opt-in display control.
- LABEL_WINDOW_BARS, the pivot-widening pass, ConfirmedZigZagLabel.
- the era-0 output-bias seed now needs a genuinely dominant class (0.70)
rather than 0.40; at ~50% Neutral a +-3.0 seed is a distortion, not a
correction.
Also fixed, both found while wiring the above:
1. RefreshConvergedSignal sized its buffers from a date delta
(Bars(sym, period, dtStudied, TimeCurrent())). dtStudied is a training
watermark; in the tester it is loaded from a live-chart save AHEAD of
the simulated date, so the interval inverted, Bars() returned ~0, and
the buffer came out at exactly m_historyBars - deep enough for the OHLC
window and far too shallow for the Donchian-50 / 20-bar-return / SMA
extension behind it. Inference silently computed DIFFERENT features
from the ones training learned on, live as well as in the tester. Now
sized from what the feature builder actually needs.
2. The barrier horizon is resolved on the deployed path too. A deployed
model never enters Train(), so it never reached the prebuild, and
OnlineLearnStep reads the horizon as its confirmation delay - left at
the fallback it would have backpropped bars whose barriers had not
resolved. Silent lookahead in the one place that writes to a live model.
SL_Mode/TP_Mode join the weights fingerprint: they define the labels now,
so a model trained at 1:3 must never be silently reused at 1:1. This
re-keys every pre-existing model by design - none were trained on this task.
Inference census extended with the vote gate. LongCondition/ShortCondition
open with a readiness check the refresh counters never see; in the tester it
reduces to "the seeded _optcache.nnw must have LOADED", and if it did not,
every vote is hard-zeroed while the model still answers Buy. The old three
counters would have read that as "the model says Neutral" - false, and a
completely different fix. This is the leading candidate for the
zero-direction backtest and the census can now name it in one run.
Both builds compile 0 errors / 0 warnings. Forces a full retrain.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 20:39:49 -04:00
|
|
|
AdvanceBarrierLabelState(i, m_labelPrebuildBars);
|
refactor: split CExpertSignalAIBase implementation by responsibility
ExpertSignalAIBase.mqh was 8216 lines: the class declaration followed by 87
method bodies covering training, labelling, feature extraction, persistence,
chart drawing, online learning, the GA auto-tuner and inference, all in one
file. Train() alone is 1492 lines; a change to arrow drawing meant scrolling
past the era loop.
Moved the bodies into Expert\AIBase\, included at the bottom of the original
after the class declaration:
Training.mqh 1607 era loop, plateau ladder, checkpoint select, deploy
Features.mqh 1093 indicator creation + per-bar input feature vector
ChartUI.mqh 634 arrows, arrow persistence, status panel, cleanup
Persistence.mqh 492 .stats/.cfg sidecars, CPU-inference validation, copy
OnlineLearning.mqh 461 live continual learning, EMA shadow, OOS simulator
Labels.mqh 309 ZigZag pivot labels, async label-cache prebuild
AutoTune.mqh 275 genetic tuner (population, crossover, halving)
Inference.mqh 235 softmax, prior calibration, class priors
ExpertSignalAIBase.mqh 8216 -> 3131 (declaration + topology build only)
This is a pure relocation - verified mechanically, not by eye: HEAD's file
reconstructed from the eight partials plus the surviving remainder is
byte-identical to HEAD, span for span (scratchpad verify_split.py). No
declaration moved, no signature changed, no code rewritten, so behaviour is
unchanged by construction.
Compiles 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 00:42:45 -04:00
|
|
|
}
|
feat(ai): triple-barrier labels replace exact-pivot ZigZag targets
The 31:1 class imbalance was self-inflicted by the TARGET, not a property
of the market. Labelling only the exact bar where a ZigZag pivot confirms
gave Buy 1164 / Sell 1164 / Neutral 35841, and every correction mechanism
this codebase accumulated sits downstream of that one choice: the
logit-adjusted loss and its range cap, the prior EMA, the +-3.0 output-bias
seed, balanced-accuracy-then-precision selection with its coverage floor,
the recall floor and its catch-22, the alternation gate, NMS, and the four
oversampling designs that collapsed before them.
The reference this engine is built on (references/neuronetworksbook.pdf
ch. 3.1/3.3) also uses ZigZag, but targets the DIRECTION TO THE NEXT
EXTREMUM on every bar - ~50/50 by construction, with no imbalance to
correct at all. It never had this problem because it never asked "is this
the pivot bar".
Labels are now the triple barrier (Lopez de Prado ch. 3), using the EA's
OWN SL_Mode/TP_Mode: does a trade opened at this bar's close reach its
target before its stop, within a horizon. Buy = long resolves, Sell =
short resolves, Neutral = neither. Consequences:
- dir-precision in the era line stops being a proxy and becomes the win
rate of the strategy under its own exit rules.
- Expected balance ~25/25/50 at the shipped 1:3 (gambler's ruin), i.e.
~2:1 instead of 31:1. Measured and logged at the end of the prebuild.
- Spread is charged on both legs, so it is a NET win rate.
- Intrabar ambiguity resolves to the STOP. OHLC cannot order two touches
inside one bar and the optimistic reading is how a backtested edge
becomes a live loss.
ZigZag stays as input features (EnableSwingContext) and now also supplies
the vertical barrier: the horizon is the median confirmed leg length,
snapped to a coarse ladder. Derived, not configured, and deliberately kept
out of the filename fingerprint - a filename keyed on a measured quantity
orphans a trained model the moment the measurement moves.
Removed, because the premise died with the old target:
- the alternation gate. Correct for pivot labels (a ZigZag cannot emit two
same-type pivots in a row, so a repeat was provably a false fire), and
wrong for barrier labels, which answer each bar independently. It also
took its worst consequence with it: a one-sided model previously got ONE
trade per backtest, a hard blocker on marketplace validation.
- SignalClusterWindow now defaults off - it de-duplicated repeats that are
now real trades. Kept as an opt-in display control.
- LABEL_WINDOW_BARS, the pivot-widening pass, ConfirmedZigZagLabel.
- the era-0 output-bias seed now needs a genuinely dominant class (0.70)
rather than 0.40; at ~50% Neutral a +-3.0 seed is a distortion, not a
correction.
Also fixed, both found while wiring the above:
1. RefreshConvergedSignal sized its buffers from a date delta
(Bars(sym, period, dtStudied, TimeCurrent())). dtStudied is a training
watermark; in the tester it is loaded from a live-chart save AHEAD of
the simulated date, so the interval inverted, Bars() returned ~0, and
the buffer came out at exactly m_historyBars - deep enough for the OHLC
window and far too shallow for the Donchian-50 / 20-bar-return / SMA
extension behind it. Inference silently computed DIFFERENT features
from the ones training learned on, live as well as in the tester. Now
sized from what the feature builder actually needs.
2. The barrier horizon is resolved on the deployed path too. A deployed
model never enters Train(), so it never reached the prebuild, and
OnlineLearnStep reads the horizon as its confirmation delay - left at
the fallback it would have backpropped bars whose barriers had not
resolved. Silent lookahead in the one place that writes to a live model.
SL_Mode/TP_Mode join the weights fingerprint: they define the labels now,
so a model trained at 1:3 must never be silently reused at 1:1. This
re-keys every pre-existing model by design - none were trained on this task.
Inference census extended with the vote gate. LongCondition/ShortCondition
open with a readiness check the refresh counters never see; in the tester it
reduces to "the seeded _optcache.nnw must have LOADED", and if it did not,
every vote is hard-zeroed while the model still answers Buy. The old three
counters would have read that as "the model says Neutral" - false, and a
completely different fix. This is the leading candidate for the
zero-direction backtest and the census can now name it in one run.
Both builds compile 0 errors / 0 warnings. Forces a full retrain.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 20:39:49 -04:00
|
|
|
//--- Final tally pass (IS-only, matches isOOS = (i < oosCutoff) used by the era loop).
|
refactor: split CExpertSignalAIBase implementation by responsibility
ExpertSignalAIBase.mqh was 8216 lines: the class declaration followed by 87
method bodies covering training, labelling, feature extraction, persistence,
chart drawing, online learning, the GA auto-tuner and inference, all in one
file. Train() alone is 1492 lines; a change to arrow drawing meant scrolling
past the era loop.
Moved the bodies into Expert\AIBase\, included at the bottom of the original
after the class declaration:
Training.mqh 1607 era loop, plateau ladder, checkpoint select, deploy
Features.mqh 1093 indicator creation + per-bar input feature vector
ChartUI.mqh 634 arrows, arrow persistence, status panel, cleanup
Persistence.mqh 492 .stats/.cfg sidecars, CPU-inference validation, copy
OnlineLearning.mqh 461 live continual learning, EMA shadow, OOS simulator
Labels.mqh 309 ZigZag pivot labels, async label-cache prebuild
AutoTune.mqh 275 genetic tuner (population, crossover, halving)
Inference.mqh 235 softmax, prior calibration, class priors
ExpertSignalAIBase.mqh 8216 -> 3131 (declaration + topology build only)
This is a pure relocation - verified mechanically, not by eye: HEAD's file
reconstructed from the eight partials plus the surviving remainder is
byte-identical to HEAD, span for span (scratchpad verify_split.py). No
declaration moved, no signature changed, no code rewritten, so behaviour is
unchanged by construction.
Compiles 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 00:42:45 -04:00
|
|
|
for(i = m_labelPrebuildBars - MathMax(m_historyBars, 0) - 1; i >= MathMax(2, m_labelPrebuildOosCutoff); i--)
|
|
|
|
|
{
|
|
|
|
|
if(!m_labelCacheHasValue[i])
|
|
|
|
|
continue; // e.g. bar was outside the dtStudied/window-edge eligibility gate above
|
|
|
|
|
if(m_labelCacheBuy[i])
|
|
|
|
|
m_labelPrebuildBuyCount++;
|
|
|
|
|
else
|
|
|
|
|
if(m_labelCacheSell[i])
|
|
|
|
|
m_labelPrebuildSellCount++;
|
|
|
|
|
else
|
|
|
|
|
m_labelPrebuildNeutralCount++;
|
|
|
|
|
}
|
2026-08-01 11:27:28 -04:00
|
|
|
//--- Prebuild complete - seed era 0's class base rates from the real upfront tally instead of leaving
|
|
|
|
|
//--- UpdateClassPriors() nothing to measure (see m_prevEraTrueBuyCount's declaration comment).
|
|
|
|
|
//--- Consumed (and cleared) by Train()'s era-start block on era 0 specifically - m_prebuildSeedPending.
|
refactor: split CExpertSignalAIBase implementation by responsibility
ExpertSignalAIBase.mqh was 8216 lines: the class declaration followed by 87
method bodies covering training, labelling, feature extraction, persistence,
chart drawing, online learning, the GA auto-tuner and inference, all in one
file. Train() alone is 1492 lines; a change to arrow drawing meant scrolling
past the era loop.
Moved the bodies into Expert\AIBase\, included at the bottom of the original
after the class declaration:
Training.mqh 1607 era loop, plateau ladder, checkpoint select, deploy
Features.mqh 1093 indicator creation + per-bar input feature vector
ChartUI.mqh 634 arrows, arrow persistence, status panel, cleanup
Persistence.mqh 492 .stats/.cfg sidecars, CPU-inference validation, copy
OnlineLearning.mqh 461 live continual learning, EMA shadow, OOS simulator
Labels.mqh 309 ZigZag pivot labels, async label-cache prebuild
AutoTune.mqh 275 genetic tuner (population, crossover, halving)
Inference.mqh 235 softmax, prior calibration, class priors
ExpertSignalAIBase.mqh 8216 -> 3131 (declaration + topology build only)
This is a pure relocation - verified mechanically, not by eye: HEAD's file
reconstructed from the eight partials plus the surviving remainder is
byte-identical to HEAD, span for span (scratchpad verify_split.py). No
declaration moved, no signature changed, no code rewritten, so behaviour is
unchanged by construction.
Compiles 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 00:42:45 -04:00
|
|
|
m_prevEraTrueBuyCount = m_labelPrebuildBuyCount;
|
|
|
|
|
m_prevEraTrueSellCount = m_labelPrebuildSellCount;
|
|
|
|
|
m_prevEraTrueNeutralCount = m_labelPrebuildNeutralCount;
|
|
|
|
|
m_prebuildSeedPending = true;
|
|
|
|
|
m_labelCachePrebuilt = true;
|
|
|
|
|
m_labelPrebuildActive = false;
|
refactor(ai): nine class-imbalance inputs down to two
The imbalance section offered nine controls for one job. Audited against the
code, five of them did not do what their names said at the shipped defaults:
AILogitPriorStrength DEAD - Inference.mqh's post-hoc prior early-returns
whenever the adjusted loss is on, which is default.
OversampleParity DEAD in training - Training.mqh gated the replay loop
on !useLogitAdjustedLoss (correctly, citing Buda et
al. 2018). Live only in the online-learning path.
EnableMinorityReplay DEAD as replay. It survived ONLY as a focal-gamma
damper - "replay minority bars through pass-2
oversampling" was a focal-loss switch.
ConstrainReplay DEAD as a cap; it only chose damper 0.125 vs 0.25.
UseStaticPrior An exact duplicate of FreezePriorCalibration - the two
were OR'd together in the single place either is read.
So they were not five mechanisms fighting; they were one mechanism plus eight
knobs that mostly described machinery that no longer ran. That is worse than
a real conflict, because the log agreed with the names: the label-cache line
printed "reps up to 28x (90% parity) (seeding era 0's class-balance
oversampling)" on every run, describing an oversampling pass that had been
switched off. It is fixed here too - it cost this session a wrong diagnosis.
The one genuine redundancy was focal loss, running at gamma*0.125 alongside
the adjusted loss: two corrections on the same axis, the exact stacking
failure this file already cited Buda et al. for in two other places, damped
by a replay flag whose replay path was itself dead. Removed rather than
re-tuned. The plateau ladder is unaffected - its escape is the learning-rate
warm restart; the gamma anneal beside it only ever stepped toward zero.
WHAT REMAINS is logit-adjusted loss (Menon et al. 2021) plus a prior freeze:
LogitAdjustTau 0 = off; replaces the separate EnableLogitAdjusted-
Loss boolean, since a strength dial where 0 already
means off does not need an on/off switch beside it.
FreezePriorCalibration unchanged.
It is the only one of the six corrections with a consistency guarantee, and
it is consistent for exactly the balanced-error metric checkpoint selection
already ranks on - so the loss and the deploy decision optimize one thing.
The online continual-learning path keeps its own alpha-balanced focal weight,
now as constants pinned to the removed inputs' shipped defaults, so its
behaviour is unchanged. It legitimately needs its own correction:
ApplyLogitAdjustment() only runs inside a training run, so a deployed model
that was reloaded carries no logit offsets and would otherwise stream 31:1
data into itself uncorrected.
The weights-filename fingerprint is BYTE-IDENTICAL. The focal slot was a
double fed to a %d conversion and had always emitted a literal 0; the |MR:
segment is written as the constant its shipped defaults produced. Dropping
either would have re-keyed every model and forced a from-scratch retrain of
the one topology currently converged and trading.
Also removed as orphans: FOCAL_GAMMA_PRESET, MAX_OVERSAMPLE_REPLICAS,
OVERSAMPLE_PARITY_FRACTION, PLATEAU_GAMMA_STEP, and the now-unreachable
"neutralized by prior correction" diagnostic.
Both builds compile 0 errors, 0 warnings. No retrain forced.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 11:46:57 -04:00
|
|
|
//--- Measured-imbalance visibility. This line used to also report "reps up to Nx (M% parity)" and
|
|
|
|
|
//--- "(seeding era 0's class-balance oversampling)" - describing an oversampling pass that the
|
|
|
|
|
//--- logit-adjusted loss had already disabled, and which no longer exists at all since 2026-07-31.
|
|
|
|
|
//--- It was pure fiction in every shipped run, and convincing enough to send a diagnosis down the
|
|
|
|
|
//--- wrong path. A log line must describe what the code DID, not what some earlier version would
|
|
|
|
|
//--- have done: report the measured distribution, which is real and useful, and nothing else.
|
refactor: split CExpertSignalAIBase implementation by responsibility
ExpertSignalAIBase.mqh was 8216 lines: the class declaration followed by 87
method bodies covering training, labelling, feature extraction, persistence,
chart drawing, online learning, the GA auto-tuner and inference, all in one
file. Train() alone is 1492 lines; a change to arrow drawing meant scrolling
past the era loop.
Moved the bodies into Expert\AIBase\, included at the bottom of the original
after the class declaration:
Training.mqh 1607 era loop, plateau ladder, checkpoint select, deploy
Features.mqh 1093 indicator creation + per-bar input feature vector
ChartUI.mqh 634 arrows, arrow persistence, status panel, cleanup
Persistence.mqh 492 .stats/.cfg sidecars, CPU-inference validation, copy
OnlineLearning.mqh 461 live continual learning, EMA shadow, OOS simulator
Labels.mqh 309 ZigZag pivot labels, async label-cache prebuild
AutoTune.mqh 275 genetic tuner (population, crossover, halving)
Inference.mqh 235 softmax, prior calibration, class priors
ExpertSignalAIBase.mqh 8216 -> 3131 (declaration + topology build only)
This is a pure relocation - verified mechanically, not by eye: HEAD's file
reconstructed from the eight partials plus the surviving remainder is
byte-identical to HEAD, span for span (scratchpad verify_split.py). No
declaration moved, no signature changed, no code rewritten, so behaviour is
unchanged by construction.
Compiles 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 00:42:45 -04:00
|
|
|
int prebuildMinDir = (int)MathMin(m_labelPrebuildBuyCount, m_labelPrebuildSellCount);
|
|
|
|
|
int prebuildMaxCls = (int)MathMax(m_labelPrebuildNeutralCount, MathMax(m_labelPrebuildBuyCount, m_labelPrebuildSellCount));
|
|
|
|
|
string prebuildRatioInfo = (prebuildMinDir > 0 && prebuildMaxCls > 0)
|
refactor(ai): nine class-imbalance inputs down to two
The imbalance section offered nine controls for one job. Audited against the
code, five of them did not do what their names said at the shipped defaults:
AILogitPriorStrength DEAD - Inference.mqh's post-hoc prior early-returns
whenever the adjusted loss is on, which is default.
OversampleParity DEAD in training - Training.mqh gated the replay loop
on !useLogitAdjustedLoss (correctly, citing Buda et
al. 2018). Live only in the online-learning path.
EnableMinorityReplay DEAD as replay. It survived ONLY as a focal-gamma
damper - "replay minority bars through pass-2
oversampling" was a focal-loss switch.
ConstrainReplay DEAD as a cap; it only chose damper 0.125 vs 0.25.
UseStaticPrior An exact duplicate of FreezePriorCalibration - the two
were OR'd together in the single place either is read.
So they were not five mechanisms fighting; they were one mechanism plus eight
knobs that mostly described machinery that no longer ran. That is worse than
a real conflict, because the log agreed with the names: the label-cache line
printed "reps up to 28x (90% parity) (seeding era 0's class-balance
oversampling)" on every run, describing an oversampling pass that had been
switched off. It is fixed here too - it cost this session a wrong diagnosis.
The one genuine redundancy was focal loss, running at gamma*0.125 alongside
the adjusted loss: two corrections on the same axis, the exact stacking
failure this file already cited Buda et al. for in two other places, damped
by a replay flag whose replay path was itself dead. Removed rather than
re-tuned. The plateau ladder is unaffected - its escape is the learning-rate
warm restart; the gamma anneal beside it only ever stepped toward zero.
WHAT REMAINS is logit-adjusted loss (Menon et al. 2021) plus a prior freeze:
LogitAdjustTau 0 = off; replaces the separate EnableLogitAdjusted-
Loss boolean, since a strength dial where 0 already
means off does not need an on/off switch beside it.
FreezePriorCalibration unchanged.
It is the only one of the six corrections with a consistency guarantee, and
it is consistent for exactly the balanced-error metric checkpoint selection
already ranks on - so the loss and the deploy decision optimize one thing.
The online continual-learning path keeps its own alpha-balanced focal weight,
now as constants pinned to the removed inputs' shipped defaults, so its
behaviour is unchanged. It legitimately needs its own correction:
ApplyLogitAdjustment() only runs inside a training run, so a deployed model
that was reloaded carries no logit offsets and would otherwise stream 31:1
data into itself uncorrected.
The weights-filename fingerprint is BYTE-IDENTICAL. The focal slot was a
double fed to a %d conversion and had always emitted a literal 0; the |MR:
segment is written as the constant its shipped defaults produced. Dropping
either would have re-keyed every model and forced a from-scratch retrain of
the one topology currently converged and trading.
Also removed as orphans: FOCAL_GAMMA_PRESET, MAX_OVERSAMPLE_REPLICAS,
OVERSAMPLE_PARITY_FRACTION, PLATEAU_GAMMA_STEP, and the now-unreachable
"neutralized by prior correction" diagnostic.
Both builds compile 0 errors, 0 warnings. No retrain forced.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 11:46:57 -04:00
|
|
|
? " | measured imbalance ~" + DoubleToString((double)prebuildMaxCls / prebuildMinDir, 1) + ":1"
|
refactor: split CExpertSignalAIBase implementation by responsibility
ExpertSignalAIBase.mqh was 8216 lines: the class declaration followed by 87
method bodies covering training, labelling, feature extraction, persistence,
chart drawing, online learning, the GA auto-tuner and inference, all in one
file. Train() alone is 1492 lines; a change to arrow drawing meant scrolling
past the era loop.
Moved the bodies into Expert\AIBase\, included at the bottom of the original
after the class declaration:
Training.mqh 1607 era loop, plateau ladder, checkpoint select, deploy
Features.mqh 1093 indicator creation + per-bar input feature vector
ChartUI.mqh 634 arrows, arrow persistence, status panel, cleanup
Persistence.mqh 492 .stats/.cfg sidecars, CPU-inference validation, copy
OnlineLearning.mqh 461 live continual learning, EMA shadow, OOS simulator
Labels.mqh 309 ZigZag pivot labels, async label-cache prebuild
AutoTune.mqh 275 genetic tuner (population, crossover, halving)
Inference.mqh 235 softmax, prior calibration, class priors
ExpertSignalAIBase.mqh 8216 -> 3131 (declaration + topology build only)
This is a pure relocation - verified mechanically, not by eye: HEAD's file
reconstructed from the eight partials plus the surviving remainder is
byte-identical to HEAD, span for span (scratchpad verify_split.py). No
declaration moved, no signature changed, no code rewritten, so behaviour is
unchanged by construction.
Compiles 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 00:42:45 -04:00
|
|
|
: " | measured imbalance n/a (a directional class has no labeled bars in this window)";
|
feat(ai): triple-barrier labels replace exact-pivot ZigZag targets
The 31:1 class imbalance was self-inflicted by the TARGET, not a property
of the market. Labelling only the exact bar where a ZigZag pivot confirms
gave Buy 1164 / Sell 1164 / Neutral 35841, and every correction mechanism
this codebase accumulated sits downstream of that one choice: the
logit-adjusted loss and its range cap, the prior EMA, the +-3.0 output-bias
seed, balanced-accuracy-then-precision selection with its coverage floor,
the recall floor and its catch-22, the alternation gate, NMS, and the four
oversampling designs that collapsed before them.
The reference this engine is built on (references/neuronetworksbook.pdf
ch. 3.1/3.3) also uses ZigZag, but targets the DIRECTION TO THE NEXT
EXTREMUM on every bar - ~50/50 by construction, with no imbalance to
correct at all. It never had this problem because it never asked "is this
the pivot bar".
Labels are now the triple barrier (Lopez de Prado ch. 3), using the EA's
OWN SL_Mode/TP_Mode: does a trade opened at this bar's close reach its
target before its stop, within a horizon. Buy = long resolves, Sell =
short resolves, Neutral = neither. Consequences:
- dir-precision in the era line stops being a proxy and becomes the win
rate of the strategy under its own exit rules.
- Expected balance ~25/25/50 at the shipped 1:3 (gambler's ruin), i.e.
~2:1 instead of 31:1. Measured and logged at the end of the prebuild.
- Spread is charged on both legs, so it is a NET win rate.
- Intrabar ambiguity resolves to the STOP. OHLC cannot order two touches
inside one bar and the optimistic reading is how a backtested edge
becomes a live loss.
ZigZag stays as input features (EnableSwingContext) and now also supplies
the vertical barrier: the horizon is the median confirmed leg length,
snapped to a coarse ladder. Derived, not configured, and deliberately kept
out of the filename fingerprint - a filename keyed on a measured quantity
orphans a trained model the moment the measurement moves.
Removed, because the premise died with the old target:
- the alternation gate. Correct for pivot labels (a ZigZag cannot emit two
same-type pivots in a row, so a repeat was provably a false fire), and
wrong for barrier labels, which answer each bar independently. It also
took its worst consequence with it: a one-sided model previously got ONE
trade per backtest, a hard blocker on marketplace validation.
- SignalClusterWindow now defaults off - it de-duplicated repeats that are
now real trades. Kept as an opt-in display control.
- LABEL_WINDOW_BARS, the pivot-widening pass, ConfirmedZigZagLabel.
- the era-0 output-bias seed now needs a genuinely dominant class (0.70)
rather than 0.40; at ~50% Neutral a +-3.0 seed is a distortion, not a
correction.
Also fixed, both found while wiring the above:
1. RefreshConvergedSignal sized its buffers from a date delta
(Bars(sym, period, dtStudied, TimeCurrent())). dtStudied is a training
watermark; in the tester it is loaded from a live-chart save AHEAD of
the simulated date, so the interval inverted, Bars() returned ~0, and
the buffer came out at exactly m_historyBars - deep enough for the OHLC
window and far too shallow for the Donchian-50 / 20-bar-return / SMA
extension behind it. Inference silently computed DIFFERENT features
from the ones training learned on, live as well as in the tester. Now
sized from what the feature builder actually needs.
2. The barrier horizon is resolved on the deployed path too. A deployed
model never enters Train(), so it never reached the prebuild, and
OnlineLearnStep reads the horizon as its confirmation delay - left at
the fallback it would have backpropped bars whose barriers had not
resolved. Silent lookahead in the one place that writes to a live model.
SL_Mode/TP_Mode join the weights fingerprint: they define the labels now,
so a model trained at 1:3 must never be silently reused at 1:1. This
re-keys every pre-existing model by design - none were trained on this task.
Inference census extended with the vote gate. LongCondition/ShortCondition
open with a readiness check the refresh counters never see; in the tester it
reduces to "the seeded _optcache.nnw must have LOADED", and if it did not,
every vote is hard-zeroed while the model still answers Buy. The old three
counters would have read that as "the model says Neutral" - false, and a
completely different fix. This is the leading candidate for the
zero-direction backtest and the census can now name it in one run.
Both builds compile 0 errors / 0 warnings. Forces a full retrain.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 20:39:49 -04:00
|
|
|
//--- These three counts are now WIN / LOSS-or-timeout counts under the EA's real stop and target, not
|
|
|
|
|
//--- pivot-spotting counts, so the Buy+Sell share here IS the fraction of bars offering a tradeable
|
|
|
|
|
//--- setup - and the era line's dir-precision against it is a win rate. This is the number that
|
|
|
|
|
//--- decides whether LogitAdjustTau still has a job: at a near-balanced split the log-prior spread
|
|
|
|
|
//--- collapses and the correction (plus its range cap, and the SoftMax port behind it) is redundant.
|
|
|
|
|
int prebuildTotal = m_labelPrebuildBuyCount + m_labelPrebuildSellCount + m_labelPrebuildNeutralCount;
|
fix: both-won bars were labelled "do not trade" - resolve by first touch
Removing the min-reward:risk raise let the MEASURED geometry come back with
the target NEARER than the stop (SP500 H1: target 1.62*ATR at q50 of
favourable, stop 3.33*ATR at q75 of adverse). That reopened a branch the
code called unreachable: price can reach +target and -target inside one
horizon, winning in BOTH directions, and those bars fell through to Neutral.
Neutral has only three producers, both-lost is unreachable (you cannot touch
-3.33 without crossing -1.62 first, which wins the short), and timeouts logged
at 1.0% of Neutral - so ~27% of ALL bars were being handed to the model as the
abstain class when a trade either way would have collected its target. The
cleanest positives in the sample, labelled "do not trade", while the fitted
confidence threshold was being asked to find selectivity in what was left.
Resolved by FIRST TOUCH: the target reached earlier is the trade that would
have closed first. Same forward window, no extra lookahead. Same-bar ties stay
Neutral - OHLC cannot order two touches, and unlike an intrabar stop tie there
is no pessimistic side to fall to, so a guess would inject a coin-flip
direction into the target.
Also:
- count both-won and its same-bar tie subset in the prebuild line, so the
share is measured rather than inferred from arithmetic on a log line
- scope the timeout counter to IS, matching the tally it is reported as a
percentage OF; it was incremented over the whole scan and divided by an
in-sample denominator
- clear m_lastBarrierTimedOut at the top of the walk with the excursions, not
at the bottom - the two early returns published the previous bar's verdict
- mark the pass-1 label line PROVISIONAL. It prints the enum fallback because
geometry can only be derived from excursions that do not exist yet, and it
reads exactly like a config change that failed to take effect
FORCES RETRAIN. Both build variants compile 0 errors, 0 warnings.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 15:35:42 -04:00
|
|
|
//--- BOTH-WON composition. Reported unconditionally rather than only when non-zero, because zero is
|
|
|
|
|
//--- itself the answer to "is the target closer than the stop" and a line that vanishes cannot say so.
|
|
|
|
|
//--- The tie share is the one to watch: it is the only part of this class still landing in Neutral, and
|
|
|
|
|
//--- if it is not small then first-touch resolution is not actually recovering these bars.
|
|
|
|
|
string prebuildBothWon = (prebuildTotal > 0)
|
|
|
|
|
? " | both targets reached (target nearer than stop) " + IntegerToString(m_labelPrebuildBothWonCount) +
|
|
|
|
|
" = " + DoubleToString(100.0 * m_labelPrebuildBothWonCount / prebuildTotal, 1) +
|
|
|
|
|
"% of bars, resolved by first touch; " + IntegerToString(m_labelPrebuildBothWonTieCount) +
|
|
|
|
|
" same-bar tie" + (m_labelPrebuildBothWonTieCount == 1 ? "" : "s") + " left Neutral"
|
|
|
|
|
: "";
|
feat(ai): triple-barrier labels replace exact-pivot ZigZag targets
The 31:1 class imbalance was self-inflicted by the TARGET, not a property
of the market. Labelling only the exact bar where a ZigZag pivot confirms
gave Buy 1164 / Sell 1164 / Neutral 35841, and every correction mechanism
this codebase accumulated sits downstream of that one choice: the
logit-adjusted loss and its range cap, the prior EMA, the +-3.0 output-bias
seed, balanced-accuracy-then-precision selection with its coverage floor,
the recall floor and its catch-22, the alternation gate, NMS, and the four
oversampling designs that collapsed before them.
The reference this engine is built on (references/neuronetworksbook.pdf
ch. 3.1/3.3) also uses ZigZag, but targets the DIRECTION TO THE NEXT
EXTREMUM on every bar - ~50/50 by construction, with no imbalance to
correct at all. It never had this problem because it never asked "is this
the pivot bar".
Labels are now the triple barrier (Lopez de Prado ch. 3), using the EA's
OWN SL_Mode/TP_Mode: does a trade opened at this bar's close reach its
target before its stop, within a horizon. Buy = long resolves, Sell =
short resolves, Neutral = neither. Consequences:
- dir-precision in the era line stops being a proxy and becomes the win
rate of the strategy under its own exit rules.
- Expected balance ~25/25/50 at the shipped 1:3 (gambler's ruin), i.e.
~2:1 instead of 31:1. Measured and logged at the end of the prebuild.
- Spread is charged on both legs, so it is a NET win rate.
- Intrabar ambiguity resolves to the STOP. OHLC cannot order two touches
inside one bar and the optimistic reading is how a backtested edge
becomes a live loss.
ZigZag stays as input features (EnableSwingContext) and now also supplies
the vertical barrier: the horizon is the median confirmed leg length,
snapped to a coarse ladder. Derived, not configured, and deliberately kept
out of the filename fingerprint - a filename keyed on a measured quantity
orphans a trained model the moment the measurement moves.
Removed, because the premise died with the old target:
- the alternation gate. Correct for pivot labels (a ZigZag cannot emit two
same-type pivots in a row, so a repeat was provably a false fire), and
wrong for barrier labels, which answer each bar independently. It also
took its worst consequence with it: a one-sided model previously got ONE
trade per backtest, a hard blocker on marketplace validation.
- SignalClusterWindow now defaults off - it de-duplicated repeats that are
now real trades. Kept as an opt-in display control.
- LABEL_WINDOW_BARS, the pivot-widening pass, ConfirmedZigZagLabel.
- the era-0 output-bias seed now needs a genuinely dominant class (0.70)
rather than 0.40; at ~50% Neutral a +-3.0 seed is a distortion, not a
correction.
Also fixed, both found while wiring the above:
1. RefreshConvergedSignal sized its buffers from a date delta
(Bars(sym, period, dtStudied, TimeCurrent())). dtStudied is a training
watermark; in the tester it is loaded from a live-chart save AHEAD of
the simulated date, so the interval inverted, Bars() returned ~0, and
the buffer came out at exactly m_historyBars - deep enough for the OHLC
window and far too shallow for the Donchian-50 / 20-bar-return / SMA
extension behind it. Inference silently computed DIFFERENT features
from the ones training learned on, live as well as in the tester. Now
sized from what the feature builder actually needs.
2. The barrier horizon is resolved on the deployed path too. A deployed
model never enters Train(), so it never reached the prebuild, and
OnlineLearnStep reads the horizon as its confirmation delay - left at
the fallback it would have backpropped bars whose barriers had not
resolved. Silent lookahead in the one place that writes to a live model.
SL_Mode/TP_Mode join the weights fingerprint: they define the labels now,
so a model trained at 1:3 must never be silently reused at 1:1. This
re-keys every pre-existing model by design - none were trained on this task.
Inference census extended with the vote gate. LongCondition/ShortCondition
open with a readiness check the refresh counters never see; in the tester it
reduces to "the seeded _optcache.nnw must have LOADED", and if it did not,
every vote is hard-zeroed while the model still answers Buy. The old three
counters would have read that as "the model says Neutral" - false, and a
completely different fix. This is the leading candidate for the
zero-direction backtest and the census can now name it in one run.
Both builds compile 0 errors / 0 warnings. Forces a full retrain.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 20:39:49 -04:00
|
|
|
string prebuildShare = (prebuildTotal > 0)
|
|
|
|
|
? " | share Buy " + DoubleToString(100.0 * m_labelPrebuildBuyCount / prebuildTotal, 1) +
|
|
|
|
|
"% Sell " + DoubleToString(100.0 * m_labelPrebuildSellCount / prebuildTotal, 1) +
|
|
|
|
|
"% Neutral " + DoubleToString(100.0 * m_labelPrebuildNeutralCount / prebuildTotal, 1) + "%"
|
|
|
|
|
: "";
|
refactor: split CExpertSignalAIBase implementation by responsibility
ExpertSignalAIBase.mqh was 8216 lines: the class declaration followed by 87
method bodies covering training, labelling, feature extraction, persistence,
chart drawing, online learning, the GA auto-tuner and inference, all in one
file. Train() alone is 1492 lines; a change to arrow drawing meant scrolling
past the era loop.
Moved the bodies into Expert\AIBase\, included at the bottom of the original
after the class declaration:
Training.mqh 1607 era loop, plateau ladder, checkpoint select, deploy
Features.mqh 1093 indicator creation + per-bar input feature vector
ChartUI.mqh 634 arrows, arrow persistence, status panel, cleanup
Persistence.mqh 492 .stats/.cfg sidecars, CPU-inference validation, copy
OnlineLearning.mqh 461 live continual learning, EMA shadow, OOS simulator
Labels.mqh 309 ZigZag pivot labels, async label-cache prebuild
AutoTune.mqh 275 genetic tuner (population, crossover, halving)
Inference.mqh 235 softmax, prior calibration, class priors
ExpertSignalAIBase.mqh 8216 -> 3131 (declaration + topology build only)
This is a pure relocation - verified mechanically, not by eye: HEAD's file
reconstructed from the eight partials plus the surviving remainder is
byte-identical to HEAD, span for span (scratchpad verify_split.py). No
declaration moved, no signature changed, no code rewritten, so behaviour is
unchanged by construction.
Compiles 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 00:42:45 -04:00
|
|
|
Print(ID + ": label cache pre-built - IS true-label distribution -> Buy: " + IntegerToString(m_labelPrebuildBuyCount) +
|
|
|
|
|
" | Sell: " + IntegerToString(m_labelPrebuildSellCount) + " | Neutral: " + IntegerToString(m_labelPrebuildNeutralCount) +
|
fix: both-won bars were labelled "do not trade" - resolve by first touch
Removing the min-reward:risk raise let the MEASURED geometry come back with
the target NEARER than the stop (SP500 H1: target 1.62*ATR at q50 of
favourable, stop 3.33*ATR at q75 of adverse). That reopened a branch the
code called unreachable: price can reach +target and -target inside one
horizon, winning in BOTH directions, and those bars fell through to Neutral.
Neutral has only three producers, both-lost is unreachable (you cannot touch
-3.33 without crossing -1.62 first, which wins the short), and timeouts logged
at 1.0% of Neutral - so ~27% of ALL bars were being handed to the model as the
abstain class when a trade either way would have collected its target. The
cleanest positives in the sample, labelled "do not trade", while the fitted
confidence threshold was being asked to find selectivity in what was left.
Resolved by FIRST TOUCH: the target reached earlier is the trade that would
have closed first. Same forward window, no extra lookahead. Same-bar ties stay
Neutral - OHLC cannot order two touches, and unlike an intrabar stop tie there
is no pessimistic side to fall to, so a guess would inject a coin-flip
direction into the target.
Also:
- count both-won and its same-bar tie subset in the prebuild line, so the
share is measured rather than inferred from arithmetic on a log line
- scope the timeout counter to IS, matching the tally it is reported as a
percentage OF; it was incremented over the whole scan and divided by an
in-sample denominator
- clear m_lastBarrierTimedOut at the top of the walk with the excursions, not
at the bottom - the two early returns published the previous bar's verdict
- mark the pass-1 label line PROVISIONAL. It prints the enum fallback because
geometry can only be derived from excursions that do not exist yet, and it
reads exactly like a config change that failed to take effect
FORCES RETRAIN. Both build variants compile 0 errors, 0 warnings.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 15:35:42 -04:00
|
|
|
prebuildRatioInfo + prebuildShare + prebuildBothWon +
|
fix: refuse invalid SL/TP, fix the unreachable deploy floor, scale the horizon
Three defects found by reading the 2026-08-01 training logs, all of which
only became visible because the relabel made the numbers mean something.
1. A STALE ENUM TRAINED FOUR MODELS ON THE WRONG TARGET.
`OnInit: trade settings snapshot - SL_Mode=1 TP_Mode=-101`
-101 was TP_PREV_SWING, deleted from TAKE_PROFIT_MODE on 2026-07-31 in
7eb48f5. MetaTrader does not validate a saved enum input against the
enum's current members, so charts saved before that kept the old
integer. BarrierMultiples()'s `if(tpMult <= 0.0) tpMult = slMult;`
then quietly turned it into a 1:1 barrier, and all four topologies
trained ~250 eras against a strategy nobody selected - while the log
reported "target 1.00*ATR" as though it were configured.
Since the relabel these two inputs ARE the label definition, so this
is not a bad trade setting, it is a wrong dataset. ValidateBarrier-
Inputs() now refuses to start (INIT_FAILED + Alert + an explicit fix)
on any value that is not an enum member. Members are enumerated rather
than range-checked because both enums are sparse and carry negative
sentinels, so no min/max test can tell a legal value from a deleted
one - which is the entire failure mode. The fallback survives as
belt-and-braces but now announces itself: a fallback that cannot say
it fired is indistinguishable from correct behaviour.
2. THE DEPLOYABILITY FLOOR BECAME MATHEMATICALLY UNREACHABLE.
`tradeableOK` required `dirPrecPct >= baseRatePct`, where baseRatePct
is Buy+Sell as a share of all bars. At the old exact-pivot target that
was ~6%, so "beat the base rate" read as "beat chance" and the test
looked sound. Triple-barrier labels put it at ~83%, so the gate now
demanded 83% directional precision - impossible by construction.
Observed live: all four topologies cycling "PLATEAU stage 3 ... nothing
safe to deploy" at a perfectly healthy 43-45% precision, with no
checkpoint able to ship however good it got.
Replaced with ZERO-SKILL precision, max(Buy,Sell)/allBars: exactly the
score of the degenerate always-call-one-direction model this floor
exists to reject. Correct at any base rate - ~43% on the current
labels, ~3% on the old rare-pivot ones. The era line now prints
"(chance N%, edge +Mpp)" beside the selection score, because 44%
precision is excellent against a 3% chance level and worthless against
a 43% one, and reading the first as the second is what made tonight's
run look better than it was.
3. THE HORIZON IGNORED THE BARRIER GEOMETRY.
ComputeBarrierHorizonBars() returned the median ZigZag leg, which
measures how long a ~1 ATR move takes and says nothing about how long
the CONFIGURED barrier needs. First-passage time out of [-m,+k] scales
with m*k, so a 1:3 barrier takes ~3x as long as 1:1; the unscaled
horizon would have timed out most 1:3 trades and pushed Neutral
straight back up, re-creating the imbalance the relabel removes.
Now multiplied by slMult*tpMult, calibrated against a real measurement
rather than assumed: the accidental 1:1 run resolved at horizon 12 with
only 16.7% timeouts, so the swing median is the right scale at m*k=1.
Verifiable, not just asserted: the prebuild now counts barriers that
ended on the VERTICAL barrier and reports them as a share of Neutral.
Neutral conflates "timed out" with "stopped out" and only the first
indicts the horizon.
Both builds compile 0 errors / 0 warnings. Forces a retrain - correcting
TP_Mode re-keys the fingerprint (|TB:1:-101 -> |TB:1:3), which is right:
no existing model was trained on the intended target.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 00:30:49 -04:00
|
|
|
" | of which timed out (horizon too short?) " + IntegerToString(m_labelPrebuildTimeoutCount) +
|
|
|
|
|
(m_labelPrebuildNeutralCount > 0
|
|
|
|
|
? " = " + DoubleToString(100.0 * m_labelPrebuildTimeoutCount / m_labelPrebuildNeutralCount, 1) + "% of Neutral"
|
|
|
|
|
: "") +
|
fix(labels): overlapping-label sample correction + horizon cap on the scale ladder
Three defects, all surfaced by the 2026-08-17 SP500 H4 run that shipped
stop 4.86 / target 9.71 (width 14.57*ATR, horizon 384).
1. EVERY STANDARD ERROR ASSUMED INDEPENDENT SAMPLES. Triple-barrier labels
started one per bar overlap by the label's lifespan, so n calls are worth
~n/L independent observations (Lopez de Prado, AFML ch. 4 - sample
uniqueness). All three sqrt(p(1-p)/n) sites divided by the RAW count.
The tell: the operating point's null-of-the-maximum gate is family-wise and
should fire on ~5% of eras under the null. Measured fire rates - PAI 47/73
(64%), ConvLSTM 9/24, LSTM 8/21 (38%), CONV 4/62 (6%). CONV, the only model
whose margin distribution admits few bins, sat on the null; the rest cleared
a bar that was too low by ~sqrt(L). PAI's deployed threshold consequently
alternated between the ENDS of its own range era to era (0.10 -> 0.88 ->
0.86 -> 0.66; coverage 16% <-> 73%).
TripleBarrierLabel now records when each label became KNOWABLE - the first
winning touch, or both stops, or the timeout - and the prebuild accumulates
the mean. EffectiveSampleSize() feeds the operating point, the member deploy
gate and the ensemble vote gate. Conservative by construction (n/L is an
upper bound on the damage); gates get harder, never easier.
2. THE SCALE LADDER RAN AWAY, again. Horizon scales as swingMedian*sl*tp, and
since 4d8cb08 reachability is measured OVER that horizon - so a wider rung
buys itself the time that makes it look reachable. Same target -> horizon ->
reach -> target loop the excursion window is kept short to avoid; fixing the
window confusion reopened it through the other door. It walked 128 -> 256 ->
384 bars and stopped at q90, the widest rung there is, with every rung
reading 39-48% against a 20% floor. A floor nothing fails selects nothing.
Rungs whose required horizon exceeds BARRIER_HORIZON_MAX are now rejected -
the same rule ReportGeometryExpectancyScan already applied. It was printing
the shipped pair as CLAMPED and disqualified ('h384!') two lines under the
deriver that chose it: two subsystems, one geometry, opposite verdicts.
3. THE RUNG SNAP DESTROYED THE RATIO IT WAS COMPARING. Both legs snapped
independently to the coarse first-passage grid, re-rating each candidate:
q90 4.86/9.71 -> 5.00/10.00 (2.00), q85 4.07/8.14 -> 5.00/10.00 (IDENTICAL
measurement), q75 3.07/6.13 -> 4.00/6.50 (1.63 - a nearer target). So the
ladder compared win shares taken at ratios from 1.63 to 2.17 and read the
differences as scale. It is why the reach column came out non-monotone in
width (q75 48.5% above q90 42.9%). The stop now snaps to its nearest rung in
log space and the target follows the ratio off it; the pair actually measured
is returned and logged, so a collision reads as a collision.
Also: LadderWinShare guarded against the conditional (fractal) geometry path,
which fills n from m_fracLegCount while leaving idxList empty - a latent
out-of-bounds on a currently-dead path.
New log lines: mean label lifespan and effective n on the label-cache line, the
required-vs-available horizon per rung, and the grid pair the reconciliation
actually measured (its tolerance now scales with the grid skew instead of a flat
5pp).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 01:12:05 -04:00
|
|
|
//--- LABEL OVERLAP, printed with the distribution because it is a property of the same
|
|
|
|
|
//--- measurement and because every standard error downstream is divided by it. A mean lifespan
|
|
|
|
|
//--- of L bars means consecutive labels share L-1 bars of outcome window, so these N labels are
|
|
|
|
|
//--- worth about N/L independent observations - see EffectiveSampleSize(). Read it beside the
|
|
|
|
|
//--- horizon: a lifespan approaching the horizon means most labels are running to the vertical
|
|
|
|
|
//--- barrier, which is the timeout share saying the same thing a different way.
|
|
|
|
|
(m_labelLifespanCount > 0
|
|
|
|
|
? StringFormat(" | mean label lifespan %.1f bars of a %d-bar horizon -> %d overlapping labels "
|
|
|
|
|
"are worth ~%d independent ones (every SE below is sized on that)",
|
|
|
|
|
MeanLabelLifespan(), m_barrierHorizonBars, (int)m_labelLifespanCount,
|
|
|
|
|
(int)EffectiveSampleSize((double)m_labelLifespanCount))
|
|
|
|
|
: "") +
|
feat(ai): triple-barrier labels replace exact-pivot ZigZag targets
The 31:1 class imbalance was self-inflicted by the TARGET, not a property
of the market. Labelling only the exact bar where a ZigZag pivot confirms
gave Buy 1164 / Sell 1164 / Neutral 35841, and every correction mechanism
this codebase accumulated sits downstream of that one choice: the
logit-adjusted loss and its range cap, the prior EMA, the +-3.0 output-bias
seed, balanced-accuracy-then-precision selection with its coverage floor,
the recall floor and its catch-22, the alternation gate, NMS, and the four
oversampling designs that collapsed before them.
The reference this engine is built on (references/neuronetworksbook.pdf
ch. 3.1/3.3) also uses ZigZag, but targets the DIRECTION TO THE NEXT
EXTREMUM on every bar - ~50/50 by construction, with no imbalance to
correct at all. It never had this problem because it never asked "is this
the pivot bar".
Labels are now the triple barrier (Lopez de Prado ch. 3), using the EA's
OWN SL_Mode/TP_Mode: does a trade opened at this bar's close reach its
target before its stop, within a horizon. Buy = long resolves, Sell =
short resolves, Neutral = neither. Consequences:
- dir-precision in the era line stops being a proxy and becomes the win
rate of the strategy under its own exit rules.
- Expected balance ~25/25/50 at the shipped 1:3 (gambler's ruin), i.e.
~2:1 instead of 31:1. Measured and logged at the end of the prebuild.
- Spread is charged on both legs, so it is a NET win rate.
- Intrabar ambiguity resolves to the STOP. OHLC cannot order two touches
inside one bar and the optimistic reading is how a backtested edge
becomes a live loss.
ZigZag stays as input features (EnableSwingContext) and now also supplies
the vertical barrier: the horizon is the median confirmed leg length,
snapped to a coarse ladder. Derived, not configured, and deliberately kept
out of the filename fingerprint - a filename keyed on a measured quantity
orphans a trained model the moment the measurement moves.
Removed, because the premise died with the old target:
- the alternation gate. Correct for pivot labels (a ZigZag cannot emit two
same-type pivots in a row, so a repeat was provably a false fire), and
wrong for barrier labels, which answer each bar independently. It also
took its worst consequence with it: a one-sided model previously got ONE
trade per backtest, a hard blocker on marketplace validation.
- SignalClusterWindow now defaults off - it de-duplicated repeats that are
now real trades. Kept as an opt-in display control.
- LABEL_WINDOW_BARS, the pivot-widening pass, ConfirmedZigZagLabel.
- the era-0 output-bias seed now needs a genuinely dominant class (0.70)
rather than 0.40; at ~50% Neutral a +-3.0 seed is a distortion, not a
correction.
Also fixed, both found while wiring the above:
1. RefreshConvergedSignal sized its buffers from a date delta
(Bars(sym, period, dtStudied, TimeCurrent())). dtStudied is a training
watermark; in the tester it is loaded from a live-chart save AHEAD of
the simulated date, so the interval inverted, Bars() returned ~0, and
the buffer came out at exactly m_historyBars - deep enough for the OHLC
window and far too shallow for the Donchian-50 / 20-bar-return / SMA
extension behind it. Inference silently computed DIFFERENT features
from the ones training learned on, live as well as in the tester. Now
sized from what the feature builder actually needs.
2. The barrier horizon is resolved on the deployed path too. A deployed
model never enters Train(), so it never reached the prebuild, and
OnlineLearnStep reads the horizon as its confirmation delay - left at
the fallback it would have backpropped bars whose barriers had not
resolved. Silent lookahead in the one place that writes to a live model.
SL_Mode/TP_Mode join the weights fingerprint: they define the labels now,
so a model trained at 1:3 must never be silently reused at 1:1. This
re-keys every pre-existing model by design - none were trained on this task.
Inference census extended with the vote gate. LongCondition/ShortCondition
open with a readiness check the refresh counters never see; in the tester it
reduces to "the seeded _optcache.nnw must have LOADED", and if it did not,
every vote is hard-zeroed while the model still answers Buy. The old three
counters would have read that as "the model says Neutral" - false, and a
completely different fix. This is the leading candidate for the
zero-direction backtest and the census can now name it in one run.
Both builds compile 0 errors / 0 warnings. Forces a full retrain.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 20:39:49 -04:00
|
|
|
(m_eraCount == 0 ? " (seeding era 0 - triple-barrier targets, so Buy/Sell mean 'target hit before stop')"
|
refactor: split CExpertSignalAIBase implementation by responsibility
ExpertSignalAIBase.mqh was 8216 lines: the class declaration followed by 87
method bodies covering training, labelling, feature extraction, persistence,
chart drawing, online learning, the GA auto-tuner and inference, all in one
file. Train() alone is 1492 lines; a change to arrow drawing meant scrolling
past the era loop.
Moved the bodies into Expert\AIBase\, included at the bottom of the original
after the class declaration:
Training.mqh 1607 era loop, plateau ladder, checkpoint select, deploy
Features.mqh 1093 indicator creation + per-bar input feature vector
ChartUI.mqh 634 arrows, arrow persistence, status panel, cleanup
Persistence.mqh 492 .stats/.cfg sidecars, CPU-inference validation, copy
OnlineLearning.mqh 461 live continual learning, EMA shadow, OOS simulator
Labels.mqh 309 ZigZag pivot labels, async label-cache prebuild
AutoTune.mqh 275 genetic tuner (population, crossover, halving)
Inference.mqh 235 softmax, prior calibration, class priors
ExpertSignalAIBase.mqh 8216 -> 3131 (declaration + topology build only)
This is a pure relocation - verified mechanically, not by eye: HEAD's file
reconstructed from the eight partials plus the surviving remainder is
byte-identical to HEAD, span for span (scratchpad verify_split.py). No
declaration moved, no signature changed, no code rewritten, so behaviour is
unchanged by construction.
Compiles 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 00:42:45 -04:00
|
|
|
: " (mid-run rebuild after new-bar cache invalidation - era " + IntegerToString(m_eraCount) + " resumes on the relabeled window)"));
|
fix: a restart no longer loses the measured geometry or the training window
Terminal restart, 22:25: all four resumed models sat on empty windows with
enum 2:6 barriers. Three interlocking causes, all visible in one log excerpt:
1) THE PRE-SCAN WINDOW WAS SIZED BY THE SAVED WATERMARK. A resumed model's
dtStudied sits at its last studied bar, so Bars(dtStudied, now) ~ 0 and the
resumed-model MI pre-scan built a zero-bar "complete" label cache - logged as
"Buy: 0 | Sell: 0 | Neutral: 0". Train()'s own era start RESETS dtStudied to
the training-window rule before computing its window; the pre-scan did not.
The rule is now factored into TrainWindowStart() and both use it. The scan
also refuses to arm before SERIES_SYNCHRONIZED (it ran in the same second as
OnInit), and deployed models keep their watermark - for them it gates
inference recency, not a training window.
2) THE HORIZON LATCHED ON AN INDICATOR WARM-UP. ComputeBarrierHorizonBars ran
against a ZigZag with 0 calculated legs, fell back, and EnsureBarrierHorizon
latched fallback(32) x slMult x tpMult = 384 for the process lifetime. A
leg-starved horizon is now PROVISIONAL: re-resolved on the next rebuild, the
label cache wiped if it moved (labels from two horizons answer different
questions), and the geometry deriver refuses to run from it - a pair derived
over a warm-up window would get PINNED.
3) THE DERIVED GEOMETRY WAS NEVER PERSISTED. The .cfg is written at model
creation and at weights-reset - both BEFORE era 0 derives - so the measured
pair lived only in memory: every restart read back zeros, adopted nothing,
fell back to the enum barriers, and the era-0-only gate meant a resumed model
could NEVER re-derive. A full day of training on 3.33/1.62 resumed as 2:6.
Now: the settled pair is pinned to the .cfg the moment derivation completes
(one-shot, atomic write), and the derive gate accepts any model with no
pinned pair, not just era 0 - mid-run stability is carried by
m_geometryDerived itself, which never allows a second derivation.
Both build variants compile 0 errors, 0 warnings.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 22:40:43 -04:00
|
|
|
//--- DERIVE THE GEOMETRY FROM WHAT WAS JUST MEASURED, then relabel under it. At era 0, OR whenever
|
|
|
|
|
//--- no pair is pinned yet (m_geometryDerived false): the "only at era 0" form of this gate meant a
|
|
|
|
|
//--- RESUMED model whose .cfg carried no derived pair could never derive - it fell back to the enum
|
|
|
|
|
//--- barriers permanently, relabelling weights that had been trained on the measured pair. The
|
|
|
|
|
//--- mid-run stability the era gate was protecting is carried by m_geometryDerived itself: once a
|
|
|
|
|
//--- pair is derived or adopted it is never re-derived, so a mid-run rebuild still cannot move the
|
|
|
|
|
//--- target under a fitted model.
|
|
|
|
|
//--- NOT while the horizon is leg-starved: the excursion window would be the indicator's warm-up
|
|
|
|
|
//--- fallback, and a pair derived from it gets PINNED (in the .cfg, below) - pinning an artifact.
|
feat: derive the ATR multiples from measured excursions - no hardcoded geometry
The barrier was still two constants. SL_Mode/TP_Mode left the Inputs tab in
3482b6c, but the fallback was a hardcoded 2:6 and the geometry scan only ever
chose from a hardcoded grid {2,3} x {2,3,4,6,8,10}. Picking the least-bad of
eleven guesses is not deriving anything.
WHY THE SCAN WAS THE WRONG INSTRUMENT, now measurable rather than argued. It
ranks pairings by how predictable their OUTCOME is - a question about direction.
The excursion test (2c78f3b) ran on SP500 H1 and direction is the one thing
absent: ASYMMETRY p=0.0846, against RANGE/UP/DOWN all at p=0.0050, with RANGE
scoring 0.01345 vs a 0.00343 null - 4x, where the barrier label sits at 1.01x.
Hence the scan failing its own gate on every run, and its "winner" wandering
2:8 -> 3:8 -> 2:8 -> 2:4 across four runs of the same data. Excursion SIZE is
strongly measurable, so derive the geometry from that instead.
stop = q25 of measured ADVERSE travel (ordinary noise does not reach it)
target = q50 of measured FAVOURABLE travel (reached ~half the time, by
construction, inside the horizon)
Continuous, in ATR units, superseding the enum multiples. Reachability ("target
on X% of bars, stop on Y%") and the implied break-even are printed so the choice
is auditable rather than trusted.
FIXED-POINT ITERATION, not one-shot. ComputeBarrierHorizonBars scales the
horizon with the target (first-passage time grows with the band) and the
excursions are measured OVER the horizon, so target -> horizon -> excursions ->
target is a real loop - deriving once sizes the target from travel measured
under the PREVIOUS horizon. Re-measures until the multiples move <5%, capped at
3 passes, and says so if it does not settle.
Does NOT create expectancy, and the log says as much: chance precision equals
break-even at every geometry (m/(m+k) on both sides). It buys a target the
market reaches and a stop that survives noise. Where Min_Risk_Reward_Ratio
forces a target the market rarely reaches, it WARNS rather than overriding -
the ratio is the user's risk policy, so the honest move is to state its cost.
That is the collision that once rejected 100% of setups.
Pinned in the .cfg as doubles appended AFTER this morning's two ints, so .cfg
files written earlier today still load (their length guard finds no doubles) and
a model that carries them was trained on them and never re-derives.
Also fixes a message from e5ceed6 that claimed "this model resumed from disk"
unconditionally - it printed above a "seeding era 0" line on a brand-new model,
because the branch fires whenever the cache is not built, which is equally true
before a fresh model's first prebuild. A diagnostic that misreports its own
trigger is worse than one that says nothing: it gets quoted back as evidence.
FORCES A FULL RETRAIN (labels change).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-07 12:06:25 -04:00
|
|
|
//--- Iterated because the horizon scales with the target and the excursions are measured over the
|
|
|
|
|
//--- horizon (see BARRIER_DERIVE_MAX_PASSES) - one pass would size the target from travel measured
|
|
|
|
|
//--- under the previous horizon.
|
fix(consistency): one geometry authority, one exit authority, and the MI screen finally gets a veto
Consistency pass before a fresh deployment. Three places where two systems were
choosing the same thing and one of them silently lost.
1. THE GEOMETRY SCAN'S DECISION WAS INERT - measured, not suspected.
USDJPY, 2026-08-17:
14:24:12.844 adopting barrier geometry 2:8 ... Relabelling and training on it.
14:24:12.979 triple-barrier labels - stop 1.61*ATR, target 3.21*ATR ... this is what trains
It adopted 2:8 and trained on 1.61:3.21. ReportBarrierGeometryScan wrote only
m_sl_mode/m_tp_mode, and BarrierMultiples ranks the DERIVED pair ABOVE those ints -
so on any model carrying a derived pair (every model with a .cfg, including a fresh
one whose weights are gone but whose sidecar survived) the adoption changed nothing.
Worse, had it changed something it would have been undone immediately: the adoption
sets m_labelCachePrebuilt = false, and that prebuild re-runs DeriveBarrierGeometry at
era 0, which overwrites m_derivedSl/TpMult from the excursion quantiles.
ONE AUTHORITY: the derived pair, because it is what the labels read, what the deploy
gate certifies, what g_Derived*AtrMult places on the live order, and what the .cfg
pins across restarts. The scan now writes THAT (floored by MIN_SL_ATR_MULTIPLIER, the
same floor DeriveBarrierGeometry applies so the live stop can never be wider than the
labelled one), republishes to the bridge immediately rather than at the next era end,
and forces the sidecar to be rewritten. m_geometryAdopted latches it so the derive
pass the adoption itself triggers cannot overwrite it.
The scan outranks the derive for an evidential reason, not an architectural one: its
winner cleared a permutation test against the null of the MAXIMUM over every eligible
pairing, and it scores the incumbent derived pair as a peer in that same field. The
derive is a descriptive quantile read with no significance test attached.
BEHAVIOURAL CHANGE, and the reason to flag it before a fresh test: barrier geometry
will now actually move when the scan says so. Until today it never did.
2. ONE EXIT AUTHORITY, tied to whose certificate the trade was placed under.
CheckClosePosition had two routes. The AI early-exit reads the AI vote undiluted and
is exactly what the new exit replay reproduces. The blended route thresholds
m_direction, the average over EVERY filter including classic ones whose live votes
pass 3 never computes - so it can close a position the certificate never modelled,
and no replay can ever check it.
When g_DerivedSlAtrMult > 0 the AI's measured geometry is on the order, which means
the deploy gate's certificate is the reason the trade exists. In that state the AI now
governs the exit and the blended route is suppressed. Classic-only configurations are
untouched: there the blended route is the only exit opinion and stays exactly as it
was. Nothing moves at the shipped defaults either way (Min_Vote_Close = Disabled).
3. EDGEFINDER, SECOND HALF: THE MEASUREMENT NOW STEERS.
The MI suite has always printed its verdicts and then trained the direction target
regardless of what they said. That gap IS the difference between this and the
EdgeFinder discipline: measure what the market offers, THEN aim.
m_dirEvidence is set when EITHER the feature/label mutual information OR the
normalised excursion asymmetry clears its block-permuted null - an OR, because the two
look for the same thing by different routes and requiring both would reject on the
weaker of two independent measurements. Normalised asymmetry specifically, never the
raw one, which is the volatility confound.
Deploy - solo AND ensemble - now requires it. A run without it still trains, and keeps
its checkpoint: the research value is real and the measurement can be wrong. It simply
may not go live. Reported separately from the statistical gate because the remedy is
different: a failed selection test says train differently, this says look somewhere
else. Excursion SIZE keeps clearing where direction does not, and that is a
risk-control head rather than an entry signal.
For the ensemble the check is per-chart by construction - the MI suite runs once and
shares its outcome across members - which is the honest treatment: four models finding
nothing between them is not four chances at an edge, it is four fits to the same
absent information.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 16:43:20 -04:00
|
|
|
//--- !m_geometryAdopted: once ReportBarrierGeometryScan has crowned a pairing that cleared its
|
|
|
|
|
//--- family-wise null, that pair is the geometry - re-measuring the scale here would overwrite a
|
|
|
|
|
//--- significance-tested choice with an untested quantile read, on the very pass the adoption itself
|
|
|
|
|
//--- triggers. See m_geometryAdopted for why the scan outranks this function rather than the reverse.
|
|
|
|
|
if((m_eraCount == 0 || !m_geometryDerived) && !m_geometryAdopted && !m_barrierHorizonLegStarved &&
|
fix: a restart no longer loses the measured geometry or the training window
Terminal restart, 22:25: all four resumed models sat on empty windows with
enum 2:6 barriers. Three interlocking causes, all visible in one log excerpt:
1) THE PRE-SCAN WINDOW WAS SIZED BY THE SAVED WATERMARK. A resumed model's
dtStudied sits at its last studied bar, so Bars(dtStudied, now) ~ 0 and the
resumed-model MI pre-scan built a zero-bar "complete" label cache - logged as
"Buy: 0 | Sell: 0 | Neutral: 0". Train()'s own era start RESETS dtStudied to
the training-window rule before computing its window; the pre-scan did not.
The rule is now factored into TrainWindowStart() and both use it. The scan
also refuses to arm before SERIES_SYNCHRONIZED (it ran in the same second as
OnInit), and deployed models keep their watermark - for them it gates
inference recency, not a training window.
2) THE HORIZON LATCHED ON AN INDICATOR WARM-UP. ComputeBarrierHorizonBars ran
against a ZigZag with 0 calculated legs, fell back, and EnsureBarrierHorizon
latched fallback(32) x slMult x tpMult = 384 for the process lifetime. A
leg-starved horizon is now PROVISIONAL: re-resolved on the next rebuild, the
label cache wiped if it moved (labels from two horizons answer different
questions), and the geometry deriver refuses to run from it - a pair derived
over a warm-up window would get PINNED.
3) THE DERIVED GEOMETRY WAS NEVER PERSISTED. The .cfg is written at model
creation and at weights-reset - both BEFORE era 0 derives - so the measured
pair lived only in memory: every restart read back zeros, adopted nothing,
fell back to the enum barriers, and the era-0-only gate meant a resumed model
could NEVER re-derive. A full day of training on 3.33/1.62 resumed as 2:6.
Now: the settled pair is pinned to the .cfg the moment derivation completes
(one-shot, atomic write), and the derive gate accepts any model with no
pinned pair, not just era 0 - mid-run stability is carried by
m_geometryDerived itself, which never allows a second derivation.
Both build variants compile 0 errors, 0 warnings.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 22:40:43 -04:00
|
|
|
m_geometryDerivePasses < BARRIER_DERIVE_MAX_PASSES)
|
feat: derive the ATR multiples from measured excursions - no hardcoded geometry
The barrier was still two constants. SL_Mode/TP_Mode left the Inputs tab in
3482b6c, but the fallback was a hardcoded 2:6 and the geometry scan only ever
chose from a hardcoded grid {2,3} x {2,3,4,6,8,10}. Picking the least-bad of
eleven guesses is not deriving anything.
WHY THE SCAN WAS THE WRONG INSTRUMENT, now measurable rather than argued. It
ranks pairings by how predictable their OUTCOME is - a question about direction.
The excursion test (2c78f3b) ran on SP500 H1 and direction is the one thing
absent: ASYMMETRY p=0.0846, against RANGE/UP/DOWN all at p=0.0050, with RANGE
scoring 0.01345 vs a 0.00343 null - 4x, where the barrier label sits at 1.01x.
Hence the scan failing its own gate on every run, and its "winner" wandering
2:8 -> 3:8 -> 2:8 -> 2:4 across four runs of the same data. Excursion SIZE is
strongly measurable, so derive the geometry from that instead.
stop = q25 of measured ADVERSE travel (ordinary noise does not reach it)
target = q50 of measured FAVOURABLE travel (reached ~half the time, by
construction, inside the horizon)
Continuous, in ATR units, superseding the enum multiples. Reachability ("target
on X% of bars, stop on Y%") and the implied break-even are printed so the choice
is auditable rather than trusted.
FIXED-POINT ITERATION, not one-shot. ComputeBarrierHorizonBars scales the
horizon with the target (first-passage time grows with the band) and the
excursions are measured OVER the horizon, so target -> horizon -> excursions ->
target is a real loop - deriving once sizes the target from travel measured
under the PREVIOUS horizon. Re-measures until the multiples move <5%, capped at
3 passes, and says so if it does not settle.
Does NOT create expectancy, and the log says as much: chance precision equals
break-even at every geometry (m/(m+k) on both sides). It buys a target the
market reaches and a stop that survives noise. Where Min_Risk_Reward_Ratio
forces a target the market rarely reaches, it WARNS rather than overriding -
the ratio is the user's risk policy, so the honest move is to state its cost.
That is the collision that once rejected 100% of setups.
Pinned in the .cfg as doubles appended AFTER this morning's two ints, so .cfg
files written earlier today still load (their length guard finds no doubles) and
a model that carries them was trained on them and never re-derives.
Also fixes a message from e5ceed6 that claimed "this model resumed from disk"
unconditionally - it printed above a "seeding era 0" line on a brand-new model,
because the branch fires whenever the cache is not built, which is equally true
before a fresh model's first prebuild. A diagnostic that misreports its own
trigger is worse than one that says nothing: it gets quoted back as evidence.
FORCES A FULL RETRAIN (labels change).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-07 12:06:25 -04:00
|
|
|
{
|
|
|
|
|
double prevSl = m_derivedSlMult, prevTp = m_derivedTpMult;
|
|
|
|
|
m_geometryDerivePasses++;
|
|
|
|
|
if(DeriveBarrierGeometry())
|
|
|
|
|
{
|
|
|
|
|
bool settled = (prevSl > 0.0 && prevTp > 0.0
|
|
|
|
|
&& MathAbs(m_derivedSlMult - prevSl) <= BARRIER_DERIVE_TOLERANCE * prevSl
|
|
|
|
|
&& MathAbs(m_derivedTpMult - prevTp) <= BARRIER_DERIVE_TOLERANCE * prevTp);
|
|
|
|
|
if(!settled)
|
|
|
|
|
{
|
|
|
|
|
if(m_geometryDerivePasses >= BARRIER_DERIVE_MAX_PASSES)
|
|
|
|
|
Print(ID + StringFormat(": barrier geometry did NOT settle within %d passes (last move "
|
|
|
|
|
"%.2f->%.2f stop, %.2f->%.2f target). Using the latest pair; the "
|
|
|
|
|
"reachability figures above are the ones to check.",
|
|
|
|
|
BARRIER_DERIVE_MAX_PASSES, prevSl, m_derivedSlMult, prevTp,
|
|
|
|
|
m_derivedTpMult));
|
|
|
|
|
else
|
|
|
|
|
{
|
|
|
|
|
//--- Re-derive the horizon for the NEW target and relabel the whole window under it.
|
|
|
|
|
//--- Train()'s !m_labelCachePrebuilt gate restarts the scan on the next call.
|
|
|
|
|
m_barrierHorizonResolved = false;
|
|
|
|
|
m_labelCachePrebuilt = false;
|
|
|
|
|
ArrayInitialize(m_labelCacheHasValue, false);
|
|
|
|
|
return;
|
|
|
|
|
}
|
|
|
|
|
}
|
|
|
|
|
}
|
|
|
|
|
}
|
fix: a restart no longer loses the measured geometry or the training window
Terminal restart, 22:25: all four resumed models sat on empty windows with
enum 2:6 barriers. Three interlocking causes, all visible in one log excerpt:
1) THE PRE-SCAN WINDOW WAS SIZED BY THE SAVED WATERMARK. A resumed model's
dtStudied sits at its last studied bar, so Bars(dtStudied, now) ~ 0 and the
resumed-model MI pre-scan built a zero-bar "complete" label cache - logged as
"Buy: 0 | Sell: 0 | Neutral: 0". Train()'s own era start RESETS dtStudied to
the training-window rule before computing its window; the pre-scan did not.
The rule is now factored into TrainWindowStart() and both use it. The scan
also refuses to arm before SERIES_SYNCHRONIZED (it ran in the same second as
OnInit), and deployed models keep their watermark - for them it gates
inference recency, not a training window.
2) THE HORIZON LATCHED ON AN INDICATOR WARM-UP. ComputeBarrierHorizonBars ran
against a ZigZag with 0 calculated legs, fell back, and EnsureBarrierHorizon
latched fallback(32) x slMult x tpMult = 384 for the process lifetime. A
leg-starved horizon is now PROVISIONAL: re-resolved on the next rebuild, the
label cache wiped if it moved (labels from two horizons answer different
questions), and the geometry deriver refuses to run from it - a pair derived
over a warm-up window would get PINNED.
3) THE DERIVED GEOMETRY WAS NEVER PERSISTED. The .cfg is written at model
creation and at weights-reset - both BEFORE era 0 derives - so the measured
pair lived only in memory: every restart read back zeros, adopted nothing,
fell back to the enum barriers, and the era-0-only gate meant a resumed model
could NEVER re-derive. A full day of training on 3.33/1.62 resumed as 2:6.
Now: the settled pair is pinned to the .cfg the moment derivation completes
(one-shot, atomic write), and the derive gate accepts any model with no
pinned pair, not just era 0 - mid-run stability is carried by
m_geometryDerived itself, which never allows a second derivation.
Both build variants compile 0 errors, 0 warnings.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 22:40:43 -04:00
|
|
|
//--- PIN THE SETTLED PAIR TO DISK. The .cfg was only ever written at model creation and at
|
|
|
|
|
//--- weights-reset - both BEFORE era 0's derivation - so the derived pair lived exclusively in
|
|
|
|
|
//--- memory: every restart read back zeros, adopted nothing, fell back to the enum barriers, and
|
|
|
|
|
//--- (with the era gate above) could never re-derive. A full day of training on the measured
|
|
|
|
|
//--- 3.33/1.62 pair resumed as 2:6 the moment the terminal restarted. One-shot per process; the
|
|
|
|
|
//--- adoption path sets the flag too, since there the pair is already on disk.
|
|
|
|
|
if(m_geometryDerived && !m_geometryCfgSaved)
|
|
|
|
|
{
|
|
|
|
|
m_geometryCfgSaved = true;
|
|
|
|
|
if(SaveTopologyConfiguration(m_activeFileName, m_initialNeuronsCount, m_hiddenLayersCount,
|
|
|
|
|
m_neuronsReduction, m_minNeuronsCount, m_optimizationAlgo,
|
|
|
|
|
m_historyBars, m_outputNeuronsCount, m_neuronsCount,
|
|
|
|
|
LEGACY_STUDY_PERIOD_SLOT, m_minTrainYear, m_isInitialized,
|
|
|
|
|
LEGACY_CONVERGE_WR_SLOT, m_fractalPeriods, m_convFilterCount,
|
|
|
|
|
m_lstmHiddenSize, m_activeFileCommon))
|
|
|
|
|
Print(ID + StringFormat(": derived geometry PINNED to the .cfg - stop %.2f*ATR, target "
|
|
|
|
|
"%.2f*ATR. A restart now adopts this pair instead of falling back "
|
|
|
|
|
"to the enum barriers.", m_derivedSlMult, m_derivedTpMult));
|
|
|
|
|
else
|
|
|
|
|
Print(ID + ": WARNING - failed to pin the derived geometry to the .cfg; a restart will "
|
|
|
|
|
"re-derive it from the same data instead of adopting it.");
|
|
|
|
|
}
|
refactor: split CExpertSignalAIBase implementation by responsibility
ExpertSignalAIBase.mqh was 8216 lines: the class declaration followed by 87
method bodies covering training, labelling, feature extraction, persistence,
chart drawing, online learning, the GA auto-tuner and inference, all in one
file. Train() alone is 1492 lines; a change to arrow drawing meant scrolling
past the era loop.
Moved the bodies into Expert\AIBase\, included at the bottom of the original
after the class declaration:
Training.mqh 1607 era loop, plateau ladder, checkpoint select, deploy
Features.mqh 1093 indicator creation + per-bar input feature vector
ChartUI.mqh 634 arrows, arrow persistence, status panel, cleanup
Persistence.mqh 492 .stats/.cfg sidecars, CPU-inference validation, copy
OnlineLearning.mqh 461 live continual learning, EMA shadow, OOS simulator
Labels.mqh 309 ZigZag pivot labels, async label-cache prebuild
AutoTune.mqh 275 genetic tuner (population, crossover, halving)
Inference.mqh 235 softmax, prior calibration, class priors
ExpertSignalAIBase.mqh 8216 -> 3131 (declaration + topology build only)
This is a pure relocation - verified mechanically, not by eye: HEAD's file
reconstructed from the eight partials plus the surviving remainder is
byte-identical to HEAD, span for span (scratchpad verify_split.py). No
declaration moved, no signature changed, no code rewritten, so behaviour is
unchanged by construction.
Compiles 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 00:42:45 -04:00
|
|
|
//--- Cold-start fix: a freshly-initialized (random-weight) network's argmax is close to uniform
|
|
|
|
|
//--- noise across the 3 classes, so on this typically heavily-skewed label distribution it fires
|
|
|
|
|
//--- far more non-majority-class calls at the very start of era 0 than the true base rate warrants,
|
|
|
|
|
//--- until enough backProp steps correct it. Now that the real prior is known, push the output
|
|
|
|
|
//--- layer's bias toward whichever class actually dominates - only the bias term moves, the
|
|
|
|
|
//--- per-input weights stay randomly initialized and still carry the real learning signal. Only
|
|
|
|
|
//--- meaningful for the 3-output classification head, and only for a fresh net (this whole prebuild
|
|
|
|
|
//--- path is skipped entirely when a trained net was loaded from disk - see m_labelCachePrebuilt).
|
|
|
|
|
//--- m_eraCount==0 gate: the prebuild can also re-run MID-run now (new-bar cache invalidation -
|
|
|
|
|
//--- see Train()'s era-start wipe check); stomping a partially-trained net's output biases with
|
|
|
|
|
//--- +-3.0 cold-start values there would erase real learned calibration, so fresh runs only.
|
|
|
|
|
if(m_outputNeuronsCount == 3 && m_eraCount == 0)
|
|
|
|
|
{
|
|
|
|
|
int dominant = 2; // Neutral
|
|
|
|
|
int dominantCount = m_labelPrebuildNeutralCount;
|
|
|
|
|
if(m_labelPrebuildBuyCount > dominantCount)
|
|
|
|
|
{
|
|
|
|
|
dominant = 0;
|
|
|
|
|
dominantCount = m_labelPrebuildBuyCount;
|
|
|
|
|
}
|
|
|
|
|
if(m_labelPrebuildSellCount > dominantCount)
|
|
|
|
|
{
|
|
|
|
|
dominant = 1;
|
|
|
|
|
dominantCount = m_labelPrebuildSellCount;
|
|
|
|
|
}
|
|
|
|
|
int totalLabeled = m_labelPrebuildBuyCount + m_labelPrebuildSellCount + m_labelPrebuildNeutralCount;
|
feat(ai): triple-barrier labels replace exact-pivot ZigZag targets
The 31:1 class imbalance was self-inflicted by the TARGET, not a property
of the market. Labelling only the exact bar where a ZigZag pivot confirms
gave Buy 1164 / Sell 1164 / Neutral 35841, and every correction mechanism
this codebase accumulated sits downstream of that one choice: the
logit-adjusted loss and its range cap, the prior EMA, the +-3.0 output-bias
seed, balanced-accuracy-then-precision selection with its coverage floor,
the recall floor and its catch-22, the alternation gate, NMS, and the four
oversampling designs that collapsed before them.
The reference this engine is built on (references/neuronetworksbook.pdf
ch. 3.1/3.3) also uses ZigZag, but targets the DIRECTION TO THE NEXT
EXTREMUM on every bar - ~50/50 by construction, with no imbalance to
correct at all. It never had this problem because it never asked "is this
the pivot bar".
Labels are now the triple barrier (Lopez de Prado ch. 3), using the EA's
OWN SL_Mode/TP_Mode: does a trade opened at this bar's close reach its
target before its stop, within a horizon. Buy = long resolves, Sell =
short resolves, Neutral = neither. Consequences:
- dir-precision in the era line stops being a proxy and becomes the win
rate of the strategy under its own exit rules.
- Expected balance ~25/25/50 at the shipped 1:3 (gambler's ruin), i.e.
~2:1 instead of 31:1. Measured and logged at the end of the prebuild.
- Spread is charged on both legs, so it is a NET win rate.
- Intrabar ambiguity resolves to the STOP. OHLC cannot order two touches
inside one bar and the optimistic reading is how a backtested edge
becomes a live loss.
ZigZag stays as input features (EnableSwingContext) and now also supplies
the vertical barrier: the horizon is the median confirmed leg length,
snapped to a coarse ladder. Derived, not configured, and deliberately kept
out of the filename fingerprint - a filename keyed on a measured quantity
orphans a trained model the moment the measurement moves.
Removed, because the premise died with the old target:
- the alternation gate. Correct for pivot labels (a ZigZag cannot emit two
same-type pivots in a row, so a repeat was provably a false fire), and
wrong for barrier labels, which answer each bar independently. It also
took its worst consequence with it: a one-sided model previously got ONE
trade per backtest, a hard blocker on marketplace validation.
- SignalClusterWindow now defaults off - it de-duplicated repeats that are
now real trades. Kept as an opt-in display control.
- LABEL_WINDOW_BARS, the pivot-widening pass, ConfirmedZigZagLabel.
- the era-0 output-bias seed now needs a genuinely dominant class (0.70)
rather than 0.40; at ~50% Neutral a +-3.0 seed is a distortion, not a
correction.
Also fixed, both found while wiring the above:
1. RefreshConvergedSignal sized its buffers from a date delta
(Bars(sym, period, dtStudied, TimeCurrent())). dtStudied is a training
watermark; in the tester it is loaded from a live-chart save AHEAD of
the simulated date, so the interval inverted, Bars() returned ~0, and
the buffer came out at exactly m_historyBars - deep enough for the OHLC
window and far too shallow for the Donchian-50 / 20-bar-return / SMA
extension behind it. Inference silently computed DIFFERENT features
from the ones training learned on, live as well as in the tester. Now
sized from what the feature builder actually needs.
2. The barrier horizon is resolved on the deployed path too. A deployed
model never enters Train(), so it never reached the prebuild, and
OnlineLearnStep reads the horizon as its confirmation delay - left at
the fallback it would have backpropped bars whose barriers had not
resolved. Silent lookahead in the one place that writes to a live model.
SL_Mode/TP_Mode join the weights fingerprint: they define the labels now,
so a model trained at 1:3 must never be silently reused at 1:1. This
re-keys every pre-existing model by design - none were trained on this task.
Inference census extended with the vote gate. LongCondition/ShortCondition
open with a readiness check the refresh counters never see; in the tester it
reduces to "the seeded _optcache.nnw must have LOADED", and if it did not,
every vote is hard-zeroed while the model still answers Buy. The old three
counters would have read that as "the model says Neutral" - false, and a
completely different fix. This is the leading candidate for the
zero-direction backtest and the census can now name it in one run.
Both builds compile 0 errors / 0 warnings. Forces a full retrain.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 20:39:49 -04:00
|
|
|
//--- Trigger raised 0.40 -> COLD_START_SEED_MIN_DOMINANCE with the triple-barrier relabel. This seed
|
|
|
|
|
//--- is an antidote to an EXTREME prior: under the old exact-pivot target Neutral held ~94% of bars
|
|
|
|
|
//--- and a uniform-ish random argmax over-called wildly for the first few thousand steps. Barrier
|
|
|
|
|
//--- labels land near 25/25/50, where sigmoid(+-3) ~ 0.95/0.05 is no longer a correction but a
|
|
|
|
|
//--- distortion - it would start the net further from the truth than random init does. Keeping the
|
|
|
|
|
//--- mechanism behind a genuinely-dominant threshold means it stays available for a skewed symbol
|
|
|
|
|
//--- (or a tight-target configuration that pushes Neutral back up) and self-disables otherwise.
|
|
|
|
|
if(totalLabeled > 0 && (double)dominantCount / totalLabeled > COLD_START_SEED_MIN_DOMINANCE)
|
refactor: split CExpertSignalAIBase implementation by responsibility
ExpertSignalAIBase.mqh was 8216 lines: the class declaration followed by 87
method bodies covering training, labelling, feature extraction, persistence,
chart drawing, online learning, the GA auto-tuner and inference, all in one
file. Train() alone is 1492 lines; a change to arrow drawing meant scrolling
past the era loop.
Moved the bodies into Expert\AIBase\, included at the bottom of the original
after the class declaration:
Training.mqh 1607 era loop, plateau ladder, checkpoint select, deploy
Features.mqh 1093 indicator creation + per-bar input feature vector
ChartUI.mqh 634 arrows, arrow persistence, status panel, cleanup
Persistence.mqh 492 .stats/.cfg sidecars, CPU-inference validation, copy
OnlineLearning.mqh 461 live continual learning, EMA shadow, OOS simulator
Labels.mqh 309 ZigZag pivot labels, async label-cache prebuild
AutoTune.mqh 275 genetic tuner (population, crossover, halving)
Inference.mqh 235 softmax, prior calibration, class priors
ExpertSignalAIBase.mqh 8216 -> 3131 (declaration + topology build only)
This is a pure relocation - verified mechanically, not by eye: HEAD's file
reconstructed from the eight partials plus the surviving remainder is
byte-identical to HEAD, span for span (scratchpad verify_split.py). No
declaration moved, no signature changed, no code rewritten, so behaviour is
unchanged by construction.
Compiles 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 00:42:45 -04:00
|
|
|
{
|
|
|
|
|
const double BIAS_MAGNITUDE = 3.0; // sigmoid(+-3) ~= 0.95/0.05 - comfortably outweighs a
|
|
|
|
|
// fresh network's random per-input weighted-sum noise
|
|
|
|
|
double biasValues[3] = { -BIAS_MAGNITUDE, -BIAS_MAGNITUDE, -BIAS_MAGNITUDE };
|
|
|
|
|
biasValues[dominant] = BIAS_MAGNITUDE;
|
|
|
|
|
if(Net.SeedOutputLayerBias(biasValues))
|
|
|
|
|
PrintVerbose(ID + ": seeded output layer bias toward " + EnumToString((ENUM_SIGNAL)(dominant == 0 ? Buy : dominant == 1 ? Sell : Neutral)) +
|
|
|
|
|
" (era 0 cold-start fix)");
|
|
|
|
|
}
|
|
|
|
|
}
|
|
|
|
|
}
|
feat(ai): triple-barrier labels replace exact-pivot ZigZag targets
The 31:1 class imbalance was self-inflicted by the TARGET, not a property
of the market. Labelling only the exact bar where a ZigZag pivot confirms
gave Buy 1164 / Sell 1164 / Neutral 35841, and every correction mechanism
this codebase accumulated sits downstream of that one choice: the
logit-adjusted loss and its range cap, the prior EMA, the +-3.0 output-bias
seed, balanced-accuracy-then-precision selection with its coverage floor,
the recall floor and its catch-22, the alternation gate, NMS, and the four
oversampling designs that collapsed before them.
The reference this engine is built on (references/neuronetworksbook.pdf
ch. 3.1/3.3) also uses ZigZag, but targets the DIRECTION TO THE NEXT
EXTREMUM on every bar - ~50/50 by construction, with no imbalance to
correct at all. It never had this problem because it never asked "is this
the pivot bar".
Labels are now the triple barrier (Lopez de Prado ch. 3), using the EA's
OWN SL_Mode/TP_Mode: does a trade opened at this bar's close reach its
target before its stop, within a horizon. Buy = long resolves, Sell =
short resolves, Neutral = neither. Consequences:
- dir-precision in the era line stops being a proxy and becomes the win
rate of the strategy under its own exit rules.
- Expected balance ~25/25/50 at the shipped 1:3 (gambler's ruin), i.e.
~2:1 instead of 31:1. Measured and logged at the end of the prebuild.
- Spread is charged on both legs, so it is a NET win rate.
- Intrabar ambiguity resolves to the STOP. OHLC cannot order two touches
inside one bar and the optimistic reading is how a backtested edge
becomes a live loss.
ZigZag stays as input features (EnableSwingContext) and now also supplies
the vertical barrier: the horizon is the median confirmed leg length,
snapped to a coarse ladder. Derived, not configured, and deliberately kept
out of the filename fingerprint - a filename keyed on a measured quantity
orphans a trained model the moment the measurement moves.
Removed, because the premise died with the old target:
- the alternation gate. Correct for pivot labels (a ZigZag cannot emit two
same-type pivots in a row, so a repeat was provably a false fire), and
wrong for barrier labels, which answer each bar independently. It also
took its worst consequence with it: a one-sided model previously got ONE
trade per backtest, a hard blocker on marketplace validation.
- SignalClusterWindow now defaults off - it de-duplicated repeats that are
now real trades. Kept as an opt-in display control.
- LABEL_WINDOW_BARS, the pivot-widening pass, ConfirmedZigZagLabel.
- the era-0 output-bias seed now needs a genuinely dominant class (0.70)
rather than 0.40; at ~50% Neutral a +-3.0 seed is a distortion, not a
correction.
Also fixed, both found while wiring the above:
1. RefreshConvergedSignal sized its buffers from a date delta
(Bars(sym, period, dtStudied, TimeCurrent())). dtStudied is a training
watermark; in the tester it is loaded from a live-chart save AHEAD of
the simulated date, so the interval inverted, Bars() returned ~0, and
the buffer came out at exactly m_historyBars - deep enough for the OHLC
window and far too shallow for the Donchian-50 / 20-bar-return / SMA
extension behind it. Inference silently computed DIFFERENT features
from the ones training learned on, live as well as in the tester. Now
sized from what the feature builder actually needs.
2. The barrier horizon is resolved on the deployed path too. A deployed
model never enters Train(), so it never reached the prebuild, and
OnlineLearnStep reads the horizon as its confirmation delay - left at
the fallback it would have backpropped bars whose barriers had not
resolved. Silent lookahead in the one place that writes to a live model.
SL_Mode/TP_Mode join the weights fingerprint: they define the labels now,
so a model trained at 1:3 must never be silently reused at 1:1. This
re-keys every pre-existing model by design - none were trained on this task.
Inference census extended with the vote gate. LongCondition/ShortCondition
open with a readiness check the refresh counters never see; in the tester it
reduces to "the seeded _optcache.nnw must have LOADED", and if it did not,
every vote is hard-zeroed while the model still answers Buy. The old three
counters would have read that as "the model says Neutral" - false, and a
completely different fix. This is the leading candidate for the
zero-direction backtest and the census can now name it in one run.
Both builds compile 0 errors / 0 warnings. Forces a full retrain.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 20:39:49 -04:00
|
|
|
//--- ConfirmedZigZagLabel() REMOVED 2026-08-01. It was the online-learning path's copy of the exact-pivot
|
|
|
|
|
//--- target; that target is gone, and its one caller now asks TripleBarrierLabel() the same question
|
|
|
|
|
//--- training asks. Keeping a second label rule alive is how the live and trained tasks drift apart.
|
refactor: split CExpertSignalAIBase implementation by responsibility
ExpertSignalAIBase.mqh was 8216 lines: the class declaration followed by 87
method bodies covering training, labelling, feature extraction, persistence,
chart drawing, online learning, the GA auto-tuner and inference, all in one
file. Train() alone is 1492 lines; a change to arrow drawing meant scrolling
past the era loop.
Moved the bodies into Expert\AIBase\, included at the bottom of the original
after the class declaration:
Training.mqh 1607 era loop, plateau ladder, checkpoint select, deploy
Features.mqh 1093 indicator creation + per-bar input feature vector
ChartUI.mqh 634 arrows, arrow persistence, status panel, cleanup
Persistence.mqh 492 .stats/.cfg sidecars, CPU-inference validation, copy
OnlineLearning.mqh 461 live continual learning, EMA shadow, OOS simulator
Labels.mqh 309 ZigZag pivot labels, async label-cache prebuild
AutoTune.mqh 275 genetic tuner (population, crossover, halving)
Inference.mqh 235 softmax, prior calibration, class priors
ExpertSignalAIBase.mqh 8216 -> 3131 (declaration + topology build only)
This is a pure relocation - verified mechanically, not by eye: HEAD's file
reconstructed from the eight partials plus the surviving remainder is
byte-identical to HEAD, span for span (scratchpad verify_split.py). No
declaration moved, no signature changed, no code rewritten, so behaviour is
unchanged by construction.
Compiles 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 00:42:45 -04:00
|
|
|
#endif // WARRIOR_AIBASE_LABELS_MQH
|