User request: attaching a chart must need zero Inputs-tab edits. Private
(non-Market) build now defaults to AIType=META, all four classic families
ON (they are the sweep's candidate sources), Meta_ExportDataset=true.
Market-build defaults unchanged (AI_NONE, MA/RSI only, no export);
UseDatabaseRanking=true applies to both per the earlier request.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The user should not need a tester corpus run per symbol. Every pattern
condition in Signals\Signal{MA,RSI,MACD,Ichimoku}.mqh anchors its reads on
`int idx = StartIndex()` with zero hardcoded indices (verified), so a
name-hiding StartIndex override + EvalShift(i) on CExpertSignalCustom makes
the EXACT live ladder code answer "what would you have fired at bar i" -
the silent-divergence trap that justified the DB corpus does not exist on
this path, and neither do the GMT-offset ambiguity, the DB row caps, or
the wipe procedure.
- CExpertSignalCustom: m_evalShift + StartIndex()/EvalShift() +
SweepPrepare(bars) (deep-resizes the shared price series); the four
classic signal classes override SweepPrepare to deep-resize their own
indicator buffers.
- CSignalMETA::BuildCorpusBySweep: per bar x per source filter, run
Direction() shifted, harvest the per-side pattern slots + netVote into
the same corpus arrays the DB loader fills; entry=bar open so
MetaPrepareEra's resolution matches at offset +0 with zero price error.
DB corpus remains the fallback when classic filters are disabled.
- Warrior_EA.mq5: META gets the enabled classic filters as candidate
sources (family ids match the descriptor one-hot).
- UseDatabaseRanking default false -> true (user request): a META chart
journals + ranks out of the box.
Workflow per symbol is now: attach ONE chart with AIType=META (optionally
Meta_ExportDataset=true for the offline pool) - candidates, labels,
training and export all happen in place, ~10 seconds of sweep instead of a
tester run.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Loads the EA's MetaExport .f32 datasets (UTF-16 sidecars), applies the EA's
own discipline offline: chronological 55/15/30 split with horizon-length
purges, operating point fitted on the calibration slice only via
coverage x (precision - BE) with the 25% floor, test slice touched once,
deployability at the 2-sigma edge floor. Small leaky-ReLU MLP + Adam in
numpy; `stats` / `eval <tag>` / `pool` commands.
First run on XAUUSD_16388 validated the plumbing and exposed the data:
the 2.5h gold tester run only covered 2004-07..2006-10 (3,214 candidates)
because gold tick volume is huge - and corpus builds do not need ticks at
all (journaling is bar-open-keyed, labels come from bar history later), so
"Open prices only" modeling builds the same corpus ~100x faster.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Meta_ExportDataset input: with AIType=META the chart writes its complete
training set once per attach - every resolved+labeled candidate as
[barTime|family|pattern|side|won|NetInputWidth floats] using the SAME
window builder, descriptor and label caches pass 2 trains on, so offline
examples are byte-equivalent to the EA's own. Sidecar .meta.csv carries
layout + the geometry/BE the labels were computed at. Files land in
Common\Files\Warrior_EA\MetaExport\<sym>_<period>.f32.
This is the pooling architecture decision: multi-symbol training INSIDE the
per-chart God-class would be the riskiest surgery this codebase has seen;
instead each chart exports, the pooled head trains offline (small dense+BN
net, minutes on this box), is validated per-symbol under the same
chronological splits and coverage x (p - BE) gate, and only a WINNER gets
written back into a .nnw for the EA to load natively (format fully mapped).
Also turns every future meta experiment from a 20-minute tester cycle into
minutes of offline iteration.
Cost-model note for the record (user challenge, verified): spread is 0.099
ATR = ~2% of the 4.74 ATR trade width - tiny per bar, but expressed in
win-rate points it is 0.099/4.74 = 2.1pp, which is the measured base-vs-BE
gap and the size of the entire observed skill lift. Zero-spread relabeling
would put base == BE by construction. Multi-day holds additionally pay swap,
which the label does NOT charge - the true bar is higher, not lower.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
999 eras on 13,436 H4 candidates. The cost gap halved exactly as computed
(BE 64.1 vs base 63.0 = 1.1pp) and the lift did not come with it: max
cov x (p-BE) = +0.07 in 1 era of 999, selective-threshold skill +0.47pp
mean (noise), OOS ranking slightly inverted (skips won 67% vs 62% for
trades) so the expectancy fitter correctly pinned coverage at 100%. The
lift shrank faster than the cost - the tick-flow decay shape, now measured
at the setup-conditional level. SP500 is closed at both accessible cost
points under pre-registered hypotheses.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The S2 verdict localized precisely: the meta head's edge x width (0.02 x
4.74 ATR = 0.095 ATR/trade) equals the measured spread (0.099 ATR/trade) -
real signal, consumed exactly by cost. The breakdown line adds: the lift is
LONG-ONLY (shorts anti-selected) and MA-family-strongest (67-70% traded win,
<1 sigma over BE on ~350 trades, best-of-32 cells - not family-wise
evidence).
Next experiment, pre-registered in Meta_Labeling_Design.md before any H4
data exists: SP500 H4 doubles ATR against a fixed spread, halving the cost
drag (~1.3pp) that the ~+2pp lift must clear. Same pipeline end to end;
deployability still decided by the unchanged 2-sigma gate. H3 (the honest
risk) is that the lift decays with timeframe as fast as cost does - the
tick-flow failure shape - which would close the single-instrument well and
leave cross-sectional pooling as the only lever.
Enabler fixed here: LoadMetaCorpus picked the LARGEST .db on disk, so an H4
chart would have adopted the (bigger) H1 corpus and resolved candidates onto
wrong bars - and a chart could even adopt another SYMBOL's corpus. The
loader now requires a <symbol>_<period>_ filename match and says so when
nothing matches.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
A resumed META model hot-looped pass 1 (0->100% scan oscillation, silent for
3 minutes until the stall reporter fired) because EVERY window failed at the
first AD/Wyckoff feature: the init-time param adoption called
ReInitADIndicators unconditionally, destroying five freshly-calculating
indicator instances to recreate them with BYTE-IDENTICAL params (verified by
parsing the .nnw header - the MI tuner had kept the configured settings), at
process start, on a box with 1 GB free of 31. The replacements sat cold for
6+ minutes while full-history resweeps starved the indicator threads harder.
- AdoptIndicatorParams: installs a loaded param set into the tuner and
rebuilds handles ONLY when the set actually differs from what the live
indicators run. Both call sites (resume init + panel reload) use it.
- Resumed models get the same 3 warm-up passes as fresh ones. The skip was
the shared root cause of the cold-ATR (ba13eef), cold-AD (2026-08-11) and
this incident - custom indicators recompute from scratch every process
start regardless of what the .nnw proves.
- Cold-sweep backoff: a pass-1 sweep in which every window failed on a
TRANSIENT cause arms a 5s era-start pause instead of an immediate
full-history resweep, so the retry loop stops consuming the CPU/memory the
warming indicators need. The stall reporter names the backoff branch.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
350-era S2 verdict on SP500 H1: the meta head carries REAL ranking skill
(+1.0-1.3pp mean over base, 101/350 eras clear their own 2-sigma bar, traded
subset wins 66.1% at <30% coverage vs 64.5% base) but 0/350 eras produced a
positive cov x (p - BE): the candidate stream sits 3pp under the derived
geometry's 67.5% break-even and ~2.6pp of recovered skill cannot bridge it.
Skill plateaued by mid-run (1.28pp -> 1.05pp), so more eras only buy
multiplicity, and the deploy gate correctly shipped nothing.
The aggregate can hide a deployable subset (one family/side clearing BE
blended with junk), so the META era line now decomposes the SAME traded
population into MA/RSI/MACD/Ichimoku x LONG/SHORT cells, each as
traded/candidates base->traded win rate. 32 cells is a best-of-N search by
construction - any candidate cell faces the family-wise rule before belief.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The NN now has a target that is not per-bar direction (closed, best-of-999
p=1.0000): P(win | this journaled candidate, at the EA's own SL/TP, net of
cost). One net for all 52 pattern-sides, AIType=AI_META.
- NetForward.mqh: the host-side softmax+CE gradient generalized total==3 ->
2||3 on both backprop paths; a 2-class softmax IS a logistic head, and no
compute backend changes.
- SignalMETA.mqh (new): corpus loaded read-only from the LARGEST signal DB on
disk (decoupled from the config fingerprint that burned four S1 runs); the
GMT->server offset is measured PER ROW against entryPrice vs bar open
(DST-immune, histogram logged); a window-span regime filter drops the
pre-2017 daily-backfill rows; 31-feature setup descriptor appended at the
input (26 one-hot + side + tanh netVote + SL/TP ATR + spread/ATR).
- Training.mqh: candidate-queued pass 1, binary-target pass 2, per-candidate
calibration (2.5) and OOS (3) walks. Counter mapping win->Buy / loss->Sell
lets checkpoint selection, the edge floor, the plateau ladder and the
family-wise deploy gate run UNCHANGED: precision reads as win rate among
traded candidates, chance as the base win rate, recalls as sensitivity/
specificity. Era-end META line: coverage x (p - break-even) vs the null.
- Labels are the side-conditional triple-barrier win caches - never the DB's
stop-and-reverse outcome. Logit adjustment deliberately skipped (~40% base
rate). Live inference + online learning guarded off until S3.
- Fingerprint: conditional |TGT:META1; State\META\ folder + 2-output filename
slot keep meta models fully separate from direction models.
Compiles clean (0 errors, 0 warnings). S2 run = attach a chart with
AIType=AI_META; S3 wires the votes via the per-side hooks.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The warning lived inside the VerboseMode-gated corpus report, so a
forgotten wipe silently voided an entire 18-year corpus run - the
outdated-row guard rejected the whole replay against leftover rows
and the run appended 35 rows instead of building a corpus. The check
now runs unconditionally at tester OnInit (MetaCorpusStaleCheck): 52
quiet one-row newest-key probes vs the test start, with a loud stop-
wipe-rerun instruction when the DB is newer than the test. Absent
tables probe quietly via FetchNewestTimeKey''s new quiet flag.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
6819bb4 called dbm.FetchRecordCount() from ProcessSignal, but the
method only existed on CDatabaseOperationsManager - CDatabaseManager
never exposed it (nothing outside the DB layer had needed it before).
The 12 compile errors were the usual MQL cascade from one unknown
member.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The historical 1000-row cap existed for a real reason: ProcessSignal
pulled BOTH full tables into MQL struct arrays on every buffered
signal, and UpdateSignalsWeights pulled all 52 per cycle -
materializing thousands of string-bearing structs per event is the
practical limit the cap protected against (SQLite itself has none).
Raising the cap for an 18-year meta-label corpus build would have
made runs crawl; sharding across databases would re-read the same
rows and inherit the same cost.
Every question is now answered inside SQLite, one row or one number
per query, flat in table size:
- FetchOpenTradeEntry: the open (NA) trade''s entryPrice for
pattern+direction, LIMIT 1
- FetchNewestTimeKey: newest row''s yyyymmddhhmm via max ROWID
(rows insert chronologically) - the duplicate/outdated guard
- FetchWinLossCounts: COALESCE''d SUM aggregates with the
before-now bound applied in SQL, replacing the tester-only array
trim (now also active live, where it is harmless by construction)
ProcessSignal semantics preserved exactly: prune -> close opposite
(stop-and-reverse still registers its own row) -> duplicate/outdated
-> one-open-trade -> register. CalculatePatternWinRate''s array walk
becomes WinRateFromCounts; the private FetchTradeRecords wrapper and
ShouldDeleteOldestEntry are gone. DB_MaxRowsPerTable=20000 is now
cheap at any table size.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The first corpus build produced 3,681 rows, all 2026, from an 18-year
backtest: ProcessSignal''s outdated-row guard rejects any registration
older than a row its table already holds (correct for a live stream),
so a tester run starting before the leftover rows'' dates silently
registers nothing for the overlap. The corpus report now prints a
loud WARNING when running in the tester with DB rows newer than the
test''s start: corpus builds start from an empty DB.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Implements stage S1 of Meta_Labeling_Design.md, superseding the
original "training-time ladder sweep": the per-side journaling from
652bf81/195be20 already produces the exact candidate stream a sweep
would compute - every pattern instance the live ladders fire, both
sides, uncensored, with netVote and touchable entry price - so the
corpus is READ from the DB instead of re-implementing 26 ladder
conditions in training code. That eliminates the silent-divergence
trap outright: the corpus is by construction identical to live
behaviour. Accepted costs are documented in the module and the doc:
coverage equals the populating backtest, and sampling is one
candidate per fire-stretch (the right dedup for training anyway).
- Expert\AIBase\MetaCorpus.mqh: CMetaCorpus reader (52 tables ->
SMetaCandidate rows) + VerboseMode OnInit report: volume/closed/
S&R-win-rate per family, span, and the GMT->server bar-offset
match table (offsets +0..+3h) that S2''s label plumbing pins to -
measured, not assumed.
- DB_MaxRowsPerTable input (default 1000 = old MAX_TABLE_ROWS): a
corpus build raises it (e.g. 20000) so a 15-20 year backtest
isn''t pruned; wired through CExpertSignalCustom::MaxTableRows().
- Report-only stage: nothing downstream consumes the corpus yet.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Agreed architecture pivot: classic patterns become candidate
generators (WHEN), the excursion head keeps geometry (SHAPE), and a
new meta-head predicts per-instance P(win at the EA''s own geometry)
(WHETHER), replacing the AI direction vote whose question is measured
closed. One net for all patterns, setup descriptor appended to the
existing feature window, triple-barrier side-conditional labels,
existing optimizer/selection/gate machinery reused. Honest floor
stated: if primaries carry zero structure the head degenerates to a
vol/cost/session timer - bounded, and the family-wise gate decides.
Four compile-alone stages; cross-sectional pooling is the next lever
after S4. No SQX parsing anywhere in this path.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
seasonal.py gets a frame-based entry (analyse_frame) and a reusable
report() so the identical statistics - circular-rotation family-wise
null, max-|t| bar, split-half - can run on instruments whose book is
synthesised from M1 bars. breadth_seasonal.py runs it on the five
SQX-decoded instruments (FTSE100, UK100, WTI x2 feeds, USDCAD) that
share no data path with the four originals; the duplicate-market pairs
(FTSE100/UK100, WTI_d/WTI_5) double as replication checks. Caveats
stated in the module docstring: synthesised flat spread (move/spread
is approximate, no intraday spread shape) and file-time clock labels;
drift/t columns are spread-free and unaffected.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Every INSERT/UPDATE a backtest journaled printed a phantom "Failed to
execute bound query (error 5126)" + "Failed to insert/update" pair -
11.7k error lines in one tester run - while every row landed
correctly (verified: v5 DB complete and identical in totals to v4,
results populated, zero non-5126 database errors in the whole log).
5126 is ERR_DATABASE_NO_MORE_DATA, SQLite''s DONE: DatabaseRead()
stepped the statement to completion and there is nothing to read back,
which for DML IS the success outcome. The tester agent reports 5126
where the live terminal reports 0 for the same completed step, and
PrepareAndExecuteBound() treated any nonzero code as failure. Success
is now 0 or 5126; genuine failures (busy, locked, constraint, misuse)
surface as other codes and still fail.
No schema or semantics change - the v5 database and its data are
valid as-is; this only stops the misreporting that would bury a real
error.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The DB system logs objectively; the decision layer reads it to compute
win rates and adjust weights. The journaling path still had one
decision-layer tendril: rows were only written when the root''s
OpenLongParams()/OpenShortParams() succeeded. Those calls validate
ORDER PLACEMENT (broker stops-level, ATR warm-up, entry-mode
rejection) and their failures cluster in volatility/spread conditions,
so the gate non-randomly censored exactly those bars out of every
pattern''s win-rate sample - the same censoring class 652bf81/c8ef478
removed, one layer down. The ledger never needed placement to be
possible: entries are marked at the touchable side of the spread and
exits are same-pattern reversals, not broker fills.
Also documents netVote for what it is: a record of the decision
layer''s state at log time (per-pattern weights inside it drift as
ranking updates land), not an objective measure - the objective part
of a row is pattern/direction/price/result.
SIGNAL_DB_SEMANTICS_VERSION 4 -> 5: row populations gain the
previously censored bars, so the database re-keys.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Verification of 652bf81 on a fresh 7-month backtest DB surfaced the
last one-sided mechanism: ProcessSignal absorbed a reversing signal as
the exit of the opposite trade and skipped registering it. For pure
EVENT patterns that strictly alternate (MACD model 3, the zero-line
cross), every reversal was consumed and the whole ledger landed on
whichever side fired first - 60 Buy rows, 0 Sell rows - so the silent
side never earned a win rate and UpdateSignalsWeights() weighted the
pattern from one side only. State patterns escaped by re-firing one
bar later.
The reversal now closes the opposite trade AND registers its own row;
the existing duplicate/outdated/open-trade checks still bound the
table at one open trade per pattern+side. Row populations change
meaning, so SIGNAL_DB_SEMANTICS_VERSION 3 -> 4 re-keys the database
(the v3 file is orphaned, not wiped - schema is unchanged).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The labelMatchesVote gate compared a single last-writer-wins label
(LongCondition then ShortCondition) against the net vote sign, which
structurally censored the pattern tables: a long event co-occurring
with any short-side state model lost its label to the later writer and
was dropped, while the mirrored short event journaled fine. Ichimoku
models 0/3 and MA model 1 could not produce a row at all by
construction (MA model 1 was "revived" in 8710240 yet still could
never journal - its weight-10 vote is exactly cancelled by the
opposing Pattern_0 state), and every pattern's win rate was measured
on a with-trend-only subset - the exact statistic
UpdateSignalsWeights() feeds back into the weights, self-sealing:
no rows -> no win rate -> default weight -> still censored.
- Direction() now evaluates the two ladders separately and snapshots
each ladder's matched pattern into its own side slot; each side that
matched journals its own row. The flat-vote poisoning the old gate
fixed stays fixed: a label can no longer contradict its side.
- The filter's net vote (raw pattern-weight units) is stored as a new
netVote column - data, never a drop filter. Snapshot is keyed on the
ladder setting a label, not on its weight, so a 0%-win-rate pattern
keeps journaling and can recover.
- SIGNAL_DB_SEMANTICS_VERSION is folded unconditionally into the DB
filename fingerprint: pattern-definition changes (b2069bc, 8710240)
re-key the database instead of blending incompatible Pattern_N
populations under one key, which the input-hash fingerprint cannot
see. 7 months of mixed-semantics rows shared one file because of it.
- dbVersion 2.0 -> 3.0: schema changed, and inserts carry the new
column, so the version-mismatch folder wipe is the migration.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The reference-pair set was re-discovered from Market Watch on every
build, so adding or removing a terminal symbol silently changed what a
trained model's six cross-asset features meant - the last open
train/serve parity gap from the 2026-08-11 audit. The set a model's
FIRST successful build actually used is now stamped into its .cfg
(append-and-length-guard, adopt-don't-compare - the derived-barrier
pattern) and every later build constructs the panel from exactly that
list; a pinned pair that is temporarily unavailable is skipped, never
substituted.
Also warms SymbolSelect/SeriesInfo for every reference symbol at
InitNeuralNetwork, so the terminal's ~minute of async cross-symbol
download starts at init instead of when the first Build() trips over
an unselected symbol - the source of the startup 'only 0 usable
reference pairs' console failures.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
On a CFD whose base and quote currency match (SP500 -> USD/USD) the FX
encoding degenerated: base and quote strength were the SAME series twice
and the divergence feature collapsed to the symbol's own 20-bar return.
Index mode re-encodes the six slots: denomination-currency strength
(fast/slow), a risk-proxy currency's strength (JPY by fixed preference
order - deterministic across rebuilds), and divergence as own move minus
what the denomination alone implies. FX-pair symbols are untouched.
Fingerprint gains :IDX2 for base==quote symbols only, so index models
trained under the degenerate encoding re-key while FX models keep their
filenames. FORCES RETRAIN on index/CFD charts.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
passTrail demanded m_excTrailScored >= EXCURSION_MIN_SCORED (500), but since
e2c9593 the trail race only scores DISJOINT bars: m_excTrailScored is bounded
by m_excScoredD (~OOS/horizon ~= 256 on SP500 H1) minus the post-ring-clear
warm-up (~8), so every chart failed "[trailing incumbent not warm enough to
race]" at 247-248 of a possible ~256 forever - observed live 2026-08-11 on
all four charts. The counter's statistical population is the same disjoint
sample passDj gates on, so it now takes the same minimum
(EXCURSION_MIN_DISJOINT, 200), reachable with margin after warm-up.
Compile: 0 errors, 0 warnings.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Three findings from the 2026-08-11 audit:
1. The excursion head's trailing-quantile ring was deliberately never cleared
between eras ("a rolling estimate of the market, not of the era") - but
pass 3 re-walks the SAME OOS window every era, so at each walk's restart
the ring still held the outcome masks of the newest OOS bars from the
previous walk: the chronological FUTURE of the bars about to be scored.
For the first ~window+horizon pushes of every era the "trailing" incumbent
was partly a leading one - conservative for the gate (an informed incumbent
is a harder hurdle) but exactly the self-made-artifact class 06d4785 hunts.
The ring now clears at era-score reset; the warm-up bars simply don't score
the trail race, which the m_excTrailN gating already accounts for.
2. skillTrail compared the head's FULL-block Brier (pro-rated by coverage)
against the incumbent's subset sum - valid only if head skill is uniform
across the OOS walk, while the trail-scored subset systematically excludes
each era's warm-up bars. The audit also found m_excBrierHeadD/BaseD/
m_excOosHitsD declared, zeroed and never accumulated (dead since e2c9593
made every scored bar disjoint). The dead trio is replaced by
m_excBrierHeadT: the head's Brier accumulated only on the bars the warm
incumbent also scored, so the race now compares both predictors on an
identical bar set.
3. The AD/Wyckoff feature blocks read GetData with no EMPTY_VALUE guard; a
cold (still-calculating) indicator returns EMPTY_VALUE everywhere, the
sanitize loop rewrote that to 0.0, and the bar SUCCEEDED - so
BufferTempData cached an all-zero Wyckoff block as a success for the whole
bar frame: the one path the f6150ee only-cache-successes rule cannot see,
because it never fails (the ba13eef class, arriving through values that
never fail; a resumed model's era-0 prebuild starts milliseconds after
OnInit). ADIndicatorCold() probes the NEWEST bar - EMPTY_VALUE there means
async warm-up (transient reject, retried), while deep bars beyond the
buffered depth keep the sanitize loop's neutral-fill so degraded history
still trains. Also fixed m_featureCacheValid's declaration comment, which
still described the pre-f6150ee cached-miss semantics.
Compile: 0 errors, 0 warnings.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
1. The expectancy stop was stone dead at shipped defaults. Its only feed -
RecordTradeResult inside CTradeJournalManager::Update() - ran solely under
UseDatabaseRanking, which ships false, so the da54639 halt was armed
(ExpectancyMinTrades=40) and never received a single closed trade. A risk
rule must not be a side effect of an analytics toggle: the journal gains
InitTrackingOnly(), Update() runs unconditionally from OnTick and skips
only the DB insert when no DB was initialized.
2. Below-minimum lots were silently bumped UP to SYMBOL_VOLUME_MIN by
TCNormalizeVolume - correct for a user-entered fixed lot, but in the
risk-sizing path it turned a budget-capped 0.05 into 0.10 on min-0.10/
step-0.01 symbols: double the intended risk, after CapRiskAmount already
clamped, exactly the routine-stop-out-breaches-the-daily-limit scenario
the budget exists to close. CMoneyRiskBase now refuses the trade when the
risk-derived lot is below the broker minimum.
3. All trading was async fire-and-forget (SetAsyncMode(true)) with no
OnTradeTransaction handler and no retry: server retcodes were never
observed. Fail-safe for entries, not for closes - a silently rejected
close rode the position until the next bar (or next day for the timed
close window). Now synchronous, matching the risk-budget flatten's own
already-synchronous CTrade; on an H1 EA the latency is irrelevant.
4. FIXED_LOT bypassed the budget entirely (no CapRiskAmount, no
OpenRiskAtStops) - pre-halt it could commit more than the remaining daily
allowance. A fixed lot cannot be scaled, so the rule is binary: its
loss-to-stop fits the remaining allowance whole or the trade is refused;
unpriceable risk (no SL) is refused while the budget is enabled.
Compile: 0 errors, 0 warnings.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
RefreshLatestSignal ran at the first tick after a bar opens and built its
window at r=0: series index 0 at that instant is a candle with one tick of
data - (close-open)/atr ~ 0, high ~ low, degenerate volume, indicators on a
1-tick bar. Training never produces such a window (every labeled bar is fully
closed, entry at that bar's CLOSE), so the deployed model's final timestep -
the one the LSTM/HYBRID output is keyed to - was out-of-distribution on every
live decision, and pass 3's deploy-gate OOS scores measured a different query
than live executed. The parity index is r=1: the newest CLOSED bar, whose
close IS the current price - the exact instant the label's hypothetical entry
happens. Single backtests shared the old skew (same r=0), which is why the
tester agreed with live while both disagreed with training.
Bookkeeping split that the index change forces: m_lastBarTime/dtStudied stay
anchored to the FORMING bar's open (they gate against SERIES_LASTBAR_DATE;
anchoring at bar 1 would re-fire the refresh every tick), while bt - the
arrow, its High/Low placement, and NMS declustering - anchors to the decision
bar, now matching the rescan path's convention.
Also: a failed refresh no longer trades the previous bar's signal for the
whole bar. RefreshLatestSignal returns success, zeroes dPrevSignal on failure
(no opinion beats a stale one), and RefreshConvergedSignal advances dtStudied
only on success so the next tick retries - the tester path (m_lastBarTime)
already worked this way; this is the live path catching up.
FORCES RE-VALIDATION of deployed models: the effective live query distribution
changes. Bundled with the backprop transpose fix's retrain.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
CaclHiddenGradient computed this layer's gradient as matrix_w[(outputs+1)*i + k]
against a buffer whose actual layout (one row per NEXT-layer neuron, stride
inputs+1) makes the correct read matrix_w[k*(inputs+1) + i]: the transpose for
square layers, and for the non-square boundaries this EA actually builds
(tapered stacks, the 3-neuron head) a mis-strided walk that ran past the buffer
end - garbage on OpenCL, zeroed reads on the CPU DLL, so the tiers did not even
agree with each other. Every gradient crossing a dense boundary on its way down
- the entire learning signal reaching the BN/conv/LSTM front ends - passed
through a fixed wrong matrix: feedback-alignment dynamics, not backprop, which
is why nets still "learned something" and this survived. The book reference
(NeuroNet_DNG) fixed this in a later article version; our kernel descended from
the earlier one. Confounds every model-based negative verdict to date.
Also in this commit, same root cause family:
- per-sample UpdateWeightsAdam (OpenCL): input for slot group j was read at
matrix_i[j] instead of matrix_i[j*4] (corrupted outer product past group 0),
and dispatch dim 1 sized on ceil(inputs/4) left the bias column unreachable
whenever inputs%4==0 - dense biases never trained on OpenCL. Rewritten as a
lane-guarded scalar loop keeping our Adam conventions (sqrt-stored v,
decoupled decay, both clamps, no sign gate). The batched accum path never had
either bug; this kernel is what SetBatchSize(1) runs - including online
continual learning on client machines, where OpenCL is the only tier.
- conv backward passed raw (int)Activation() where the kernels expect
NativeActivationCode(): NONE took the tanh branch (clamping a BN layer's
unbounded z-scores), TANH took sigmoid, PRELU took none. Dormant only because
the conv sits at layer 1 today.
- hidden-gradient dispatch over Neurons()+1 dropped to Neurons(): biases get no
backprop gradient and the extra work-item only ever read past matrix_o.
All three backends (Network.cl, WarriorCPU.cpp, WarriorDML.cpp HLSL) changed in
lockstep; DML gained an `inputs` constant to derive the row stride. New
dense_backprop_check.cpp proves the CPU kernel is central-finite-difference
consistent with the real forward kernel on 8x8, 64x3, 33x64, 5x3 (max diff
3e-9) and that all three activation branches match transcription. All 16 checks
pass. Offline math check only - the in-situ proof remains the per-layer dW/W
report on a real era.
FORCES FULL RETRAIN. Both DLLs rebuilt and redeployed to MQL5\Libraries.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Measured on this machine's actual CPU at the real 760-wide geometry, through a
real DLL boundary (an earlier harness #included the .cpp and the fast-math
build hoisted the timing loop, reporting a flat ~4us for shapes 8x apart).
Every hot kernel is a floating-point reduction. Under the default /fp:precise
MSVC may not reassociate one, so it cannot vectorize one - the dot product was
scalar mulsd/addsd through a single accumulator. Forward pass measured
1.1-1.9 GFLOP/s precise vs 1.7-2.8 GFLOP/s fast, and the same three
neurons*inputs loops (forward, hidden gradient, weight-gradient accumulate)
dominate an era.
batch_accum_check passes on both builds with identical output to every digit
it prints, including the 5-decade optimizer scale-invariance sweep. The
deploy-time CPU-vs-MQL5 self-check tolerance is 1.0e-3, ~11 orders looser than
fast-math drift.
/arch:AVX2 is now explicitly forbidden in the script with the reason. This CPU
is an Ivy Bridge-EP Xeon: AVX yes, AVX2/FMA no. An AVX2 build faults on every
kernel, SehCallFn swallows it per dispatch, buffers are never written, and
every shape takes a flat ~5us - which benchmarks as a 250x speedup until you
check that the outputs are all zero. /arch:AVX alone was measured and bought
nothing; these loops are memory-bound and Ivy Bridge splits 256-bit loads
into 2x128 anyway.
Requires rebuilding WarriorCPU.dll (build_cpu.bat) - the .ex5 is unchanged.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Pass 1 already skipped its feedForward on QUEUED bars, because pass 2 redoes
them. The same argument covers two more bands it was still forwarding:
OOS window (30% of bars) - pass 3 re-forwards every one of them
calibration band (~10% of bars) - pass 2.5 re-forwards every one of them
All three passes derive their bounds from the same helpers and apply the
identical eligibility test, so the bar sets are equal by construction, not by
coincidence. Only the two purge bands and the ineligible edge bars are visited
in pass 1 and nowhere else - those keep their forward pass.
The scan's copy was never the one that survived. Its arrow-cache write was
overwritten by pass 3's (with the thresholded, post-training decision), its
status-label paint was transient, and its predicted-class tally measured
last era's weights. Those tallies move to pass 2.5 and pass 3, on the raw
argmax exactly as pass 1 and pass 2 count it, so the population behind the
panel's "Predicted -> Buy/Sell/Neutral" line is unchanged and stays comparable
with the "Actual" line beside it, which pass 1 still accumulates over every
labelled bar.
Verified unaffected by the cut: dPrevSignal and m_lastBarTime are both written
last by bars 0/1, which are label-ineligible and therefore still forwarded, so
FinalizeTrainRun's `dtStudied = m_lastBarTime` and Lifecycle's newBarPending
sentinel read the same values as before.
Correctness, not just speed: batch norm is UNFROZEN during pass 1 (passes 2.5
and 3 freeze it deliberately), so every scan-time forward on a held-out bar was
advancing the BN running mean/variance from data the model is graded on. Those
running statistics are inference-time model state. It is the mild,
unsupervised kind of leakage - feature statistics, not labels - but it fed the
weights pass 3 then scored, and it is now gone.
Cost: ~40% of all bars lose one forward pass per era, ~16% of net time once
pass 2's backward pass is weighted in. Per-dispatch, so it lands on every
backend.
Both variants compile 0 errors / 0 warnings. Build tag scan-nofwd-v5.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Measured on exc-race-v3: LSTM era 300s -> 1087s (net 272->748s, "other"
30->337s). My estimate had been "single-digit percent". The cost is
per-DISPATCH, not per-FLOP, and therefore hits EVERY backend: the head is
19k weights and ~2.4 GFLOP an era - seconds of arithmetic - but ~48k
forward/backward calls x several layer submits each, and its 760-wide
layer exceeds the CPU DLL's inline threshold so each one pays a real
handoff. The classifier's own net time tripled too, from contention with
a second pool on an already-full box.
Three changes, all backend-neutral because they remove submits rather
than tune threads:
SCORE ONLY DISJOINT WINDOWS (~64x). Adjacent bars share all but one bar
of their horizon, so 16k consecutive bars were always ~250 independent
observations - the full-sample tally was never worth more than the
disjoint one, it just quoted an n that was ~64x too large. Dropping it
costs nothing statistically and removes 63 of every 64 forward passes.
The two parallel tallies collapse into one, which is also less code.
The trailing ring still advances on every bar: it needs the outcome
SEQUENCE, and that is array lookups, not a forward pass.
TRAIN ON EVERY 4th PRIMARY BAR (4x). The target is low-dimensional and
strongly autocorrelated - neighbouring bars carry near-identical
excursion information - so per-bar training buys resolution the target
does not have. Strided on ATTEMPTS, not acceptances, so a stretch of
unlabelled bars cannot silently change the spacing.
OWN TIMING COLUMN. The head's passes were landing in the era line's
"other" bucket, which is how a 3.6x regression read as an unexplained
jump in the one column nobody attributes. A cost that cannot be seen in
the timing line cannot be traded off against anything.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Leftover objects survived deinit because the cleanup list had drifted.
PurgeChart()'s own comment said it removed "our namespaced signal arrows
plus the status-label objects" while the code removed arrows ONLY, and
the panel prefix was swept at OnInit and nowhere else - so an ordinary
deinit left the status line, and any panel straggler, on the chart.
Three scattered call sites and a comment cannot be kept in step. There is
now ONE list - WarriorChartPrefixes() - covering arrows, status label and
panel, and one sweep, WarriorPurgeChartObjects(), used by every path.
Add a prefix there when a new object family appears and every cleanup
picks it up.
Two call sites added:
OnInit, before ANYTHING is drawn (including the status label it would
otherwise delete). Chart objects live in the chart PROFILE, not in the
EA, so they outlive the process: a deinit force-terminated at
MetaTrader's ~4,500 ms budget, a crash, a terminal kill, or an .ex5
replaced while attached all strand objects no later deinit will ever
own - and deleting the EA's files does not remove them, which is why
they read as corruption. Arrows are included: LoadChartSignals restores
them from their sidecar moments later and already opens with its own
arrow sweep, so this only removes orphans the sidecar does not account
for - the ones SaveChartSignals would otherwise ADOPT, since it rebuilds
that sidecar by scanning the chart.
OnDeinit, after ExtPanel.Destroy. Destroy walks an unbounded control
tree and ClearStatusLabel clears text rather than guaranteeing object
removal; either can leave a straggler and nothing looked afterwards.
Bounded work - three prefix deletes and one object-list scan - so it
respects the ordering rule that keeps the cheap visible cleanup ahead
of the heavy save. Arrows excluded: ShutdownChartCleanup already
persisted and removed them and re-deleting would race that write.
The two are complementary: the deinit sweep closes the ordinary case, the
OnInit purge closes the case where MetaTrader never let us finish. Only
the second can help after a starved shutdown.
Both sweeps rescan by name across EVERY object type and delete what the
bulk call missed. ObjectsDeleteAll's return has already been observed
disagreeing with a by-name scan of the same chart microseconds apart, and
object commands are queued on the chart rather than applied inline, so a
returned count is not evidence the objects are gone.
Panel create site now uses WARRIOR_PANEL_PREFIX instead of a literal, so
the name cannot drift away from the list that cleans it up.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Beating a frozen global constant is the weakest admissible bar for
replacing a global constant. The honest incumbent is a rolling rung
frequency: it adapts to the volatility regime - exactly what the head
claims to predict - and needs no model, no 760 inputs and no training.
Implemented as a ring of per-bar outcome bitmasks (32 rungs fit one
ulong), sized horizon + EXCURSION_TRAIL_WINDOW. The newest `horizon`
entries are held back UNRESOLVED: a bar's rung outcomes are only known
one horizon later, so using them would be lookahead and would flatter the
incumbent into an opponent the head could never fairly beat. Pass 3 walks
oldest-to-newest, so "pushed more than horizon bars ago" is exactly
"resolved by now". Each push is O(rungs), not O(window).
The head's decision-rung Brier is pro-rated to the trailing estimate's
coverage before the ratio, since the incumbent only scores bars where its
window is warm.
This line is worth reading on its own, independently of the head: if the
trailing quantile beats the global constant, that is a cheap risk-control
win available with no machine learning at all - and it is the same number
either way, so the run answers both questions in one pass.
The ring is deliberately NOT reset per era - it estimates the market, not
the era, and re-warming 500 bars every era would leave the incumbent
unusable over the first chunk of every scoring pass, handing the head a
free win on exactly those bars.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Second-opinion review killed the +4.2% far-rung result, correctly, and
the mechanism is my own bug. A head trained toward {0.05,0.9} converges
to 0.05+0.85p, so its bias is 0.05-0.15p: negative where p is near 1,
POSITIVE where p < 1/3, growing monotonically as the rung gets farther.
Against a baseline frozen at the IS rate, an upward-biased head scores
positive Brier skill whenever the OOS rate merely sits above the IS rate.
Predicted signature: huge negatives near, ~zero at p=1/3, growing
positives far. Observed: -82% ... -0.6% ... +1.2/+2.7/+4.2. The far rungs
were not the clean end of a distorted measurement, they were the other
face of the same artifact. Everything before 25aca83 is void.
The gate was a bare `skill >= 2%` point estimate over 8 rungs x 4
topologies x N eras, reported per era - a best-of-~300 with no interval
and no multiplicity control, which is the shape of the four traps already
documented here. It now needs FOUR things at once:
DECISION RUNGS only the rungs ExcursionQuantile actually reads at the
live geometry (target 1.62, stop 3.31 ATR), fixed
before looking. Skill at 5 ATR is skill about a
distance no order is placed at - and the TARGET side
currently interpolates 1.5/2.0, which measured -2.2%
and -1.3%.
DISJOINT SAMPLE one bar per horizon. Adjacent bars share 63 of 64
horizon bars, so ~16k scored bars is ~250 independent
ones and every SE over the full set is ~8x understated.
VS ORACLE the best constant achievable ON THE SCORED BLOCK,
closed form from H and n (Brier = H*(1-H/n)). A head
that learned only a LEVEL nearer the OOS rate than the
frozen IS constant scores positive against the old
baseline and <= 0 here. This is the control that
separates per-bar skill from base-rate drift.
MONOTONE CURVE P(reach k) must be non-increasing in k. Nothing
constrained 8 independent sigmoids to obey that, and
ExcursionQuantile returns the FIRST crossing - so a
tangled curve is misread exactly where the head is
least sure. Counted and reported, not silently used.
The pass message now also states what a pass would and would not buy:
expectancy is -costs at zero directional edge whatever the stop distance,
and under prop DD limits LOWER variance also lowers P(reach target before
limit), so "better drawdown" is a choice of failure mode, not a win.
Still owed before any Stage 2: a race against a trailing-quantile
incumbent and a vol-feature logistic. Beating a frozen global constant is
the weakest admissible bar for replacing a global constant.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
ExcursionTargets built its 32 binary targets from the classifier's
LABEL_SMOOTH_HIGH/LOW (0.9/0.05). That caps what the head can ever output
at 0.9, and the near ladder rungs have base rates close to 1.0 - almost
every bar travels 0.5 ATR inside a 64-bar horizon. The Brier comparison
is then decided before the net learns anything:
constant at 0.99 -> 0.99*(0.01)^2 + 0.01*(0.99)^2 = 0.0099
head at 0.90 -> 0.99*(0.10)^2 + 0.01*(0.90)^2 = 0.0180 skill -82%
Which is what the first run reported at rung 0.50: PAI -61.8%,
CONV -146%. A property of the target encoding, not of predictability.
Smoothing earns its place on the 3-class head, where it stops one logit
running away inside a softmax competition. There is no competition here
and this head is scored on calibration, so it has to be free to say 0.99
when the answer is 0.99. Hard 1/0 is safe against the runaway smoothing
guards: this is an MSE-on-sigmoid gradient (calcOutputGradients) whose
(target - output) term vanishes as the output approaches the target, not
the unbounded-logit cross-entropy the classifier uses.
The far rungs, where the artifact is smallest, already showed positive
skill on the two topologies with a sequence stage (LSTM 3.00:+1.2%
4.00:+2.7% 5.00:+4.2%, HYBRID similar), so the verdict was being decided
by the most distorted end of the ladder.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
ba13eef cached a miss unless it was flagged transient, and flagged
exactly two guards: the EMPTY_VALUE open and the cold ATR. Every other
rejection in BufferTempDataCompute - an indicator buffer not yet
calculated, a panel not yet built, a series not yet loaded, a failed Add -
still cached as PERMANENT.
Observed 2026-08-11: the MI pre-scan runs ~3 s after OnInit and touches
all 54k bars while the indicators are still warming. The log announced it
immediately and unmistakably:
feature/label information - ... (0 samples 19 bars apart
= 0 independent blocks over a 64-bar horizon, 0.0s)
Zero usable rows, four seconds in. Training then stalled at era 0 for an
hour with "NOT ONE of 54681 scanned bars produced a usable feature
window" on all four charts. Both charts reporting cross-asset PRESENT and
both reporting ABSENT got 0 samples, so the optional block was not the
discriminator - the cache was.
Enumerating which rejections are "really" permanent is the wrong shape of
fix: it is a list that must be re-audited every time a feature block is
added, and being wrong once costs the whole run silently - which is
exactly how the two-guard version failed. Caching only successes needs no
list and cannot be wrong.
Cost is bounded and small: in steady state the only bars that still fail
are the handful at the deep end of history inside the indicators' own
warm-up, so an era recomputes ~ind_Periods bars rather than 54k.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Era 0 stalls with "NOT ONE of 54681 scanned bars produced a usable
feature window, windows ok=0 failed=54681" and nothing else. That line
reads identically for a cold ATR, a conditionally-missing optional
feature block and an out-of-range index, so it cannot be diagnosed
without one restart per hypothesis.
Two changes:
1. WIDTH CONTRACT in BufferTempData. Every enabled block must emit
exactly m_neuronsCount values on EVERY bar. A block that emits its
values on some bars and skips them on others (indicator, panel or
series unavailable for that bar) does not merely shorten the window -
it SHIFTS every feature after it into the wrong slot, and the net
then trains on silently misaligned inputs that still look like a
valid window to everything downstream. Now rejected, rolled back and
reported once, naming the optional blocks (XA / SPR / swing context)
as the ones carrying an availability test. Worth having independently
of the current stall.
2. BuildFeatureWindow records WHICH lookback slot rejected and how much
of the window was assembled, and the pass-1 stall report renders it:
"slot 0 of 20 REJECTED (window had 0 of 760)" is an indicator warm-up
or history-edge read; "every lookback bar ACCEPTED and the window was
still short: 640 of 760" is a missing 6-value block.
No behaviour change on a healthy run: the width check is an equality
that already holds, and the diagnostics render only inside the
total-failure branch.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Direction is closed - normalised asymmetry fails on three instruments
with a working positive control, and the classifier's own best-of-999
era-cap test agrees (+0.9pp = 1.48 sigma, family-wise p=1.0000). SIZE is
a different question and RANGE clears at ~4x its null.
Checked the denomination before building on that, since the source memo
warns to: m_excUpCache holds (maxHigh - fill)/ATR, so "RANGE is
predictable" is a claim about travel RELATIVE to current ATR, not a
restatement of "ATR is autocorrelated". It is exactly the part a fixed
multiple (stop 3.31*ATR, target 1.64*ATR) discards.
A second small CNet, 760 -> 24 -> 32 sigmoid outputs = P(price reaches
ladder rung k) upward and downward. Survival parameterisation rather than
regressing the multiple, because it needs nothing new from CNet: sigmoid
outputs and the per-neuron delta the `total != 3` branch already applies
(a quantile head would need a linear activation and a pinball gradient in
Network.mqh, Network.cl and the DirectML path, on a class four topologies
share). Targets are free - m_ladderUpAt already records first-touch age
per rung with 0 meaning never reached.
Separate net, not extra outputs on the classifier: more outputs would
change m_outputNeuronsCount, the .nnw shape and the fingerprint, and push
the count off 3 - the exact condition backProp uses to select the joint
softmax gradient the 3-class head depends on. The classifier is
bit-for-bit unaffected and this is removable without trace.
STAGE 1 PLACES NO ORDERS. It reports a Brier skill score against the
constant per-rung base rate - the baseline a fixed ATR multiple already
assumes - with both predictors fitted IS and evaluated OOS, so neither
gets a look at the test set. Positive skill justifies Stage 2 (drive
SL/TP and sizing off ExcursionQuantile, which is defined and deliberately
uncalled). Zero or negative means ATR already carries everything and
Stage 2 must not be built.
Trains only on primary occurrences: the replay queue oversamples for
CLASS balance, and a direction-balanced sample is a biased SIZE sample.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
coveragePct, dirPrecPct and the declustered TRADED tally were all computed
from oPrevSignal - the RAW argmax - while the live order, the arrow and the
panel all run on oDeploySignal, which is argmax AFTER the confidence
threshold. The gate was certifying a strategy the EA does not trade.
Invisible until now: the threshold sat at ~0.02, so the two populations
were the same set. The held-out calibration slice (2189316) moved it to
0.14-0.40 and the gap opened immediately - PAI era 256 graded 100% coverage
while its traded population was 21% (3,399 of ~16,200 OOS bars).
Consequences that were being hidden:
- coveragePct >= minCoveragePct was tested against the wrong population,
so a model whose TRADED coverage falls under the 24.8% floor still read
as clearing it
- precSE = sqrt(p(1-p)/n) used n ~16,000 instead of n ~3,400, so the
EDGE_MIN_SIGMAS bar was ~2.2x too lenient on the real evidence
- the NMS replay declustered a different, larger stream than live, so
threshold-rejected bars consumed cluster slots and set alternation state
Gate quantities now read m_oosBuyFired/m_oosSellFired (the thresholded
population, already tracked for the live-precision line) and the NMS replay
runs on oDeploySignal. The threshold can only turn a direction into Neutral,
never flip a side, so the fired set is a strict subset and every per-bar
outcome is the one already computed.
Recall and logBuyPrecPct deliberately stay on the raw argmax: they measure
intrinsic class separation, and thresholding them would conflate "cannot
separate the classes" with "declines to act on the separation it found".
This is the 9a7c37f defect class, and the NMS block carried a comment
warning about it while committing it three lines above.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
FitDirConfThreshold harvested its margin histogram from pass 2's own
backprop samples. Pairing every fit against the same era's OOS result
shows what that measured:
PAI era 1 IS 25% cov @ 66.1% (-0.8pp) -> OOS 64% (-3pp) gap +2.1pp
PAI era 76 IS 90% cov @ 79.6% (+12.7pp) -> OOS 65% (-2pp) gap +14.6pp
LSTM era 9 IS 77% cov @ 81.6% (+14.6pp) -> OOS 63% (-4pp) gap +18.6pp
The gap grows monotonically while OOS stays flat, so within a handful of
eras the curve stops describing behaviour on unseen bars. That is fatal
here specifically, because the objective branches on the SIGN of
(p - break-even): the memorized curve reads +12pp at 95% coverage, so
coverage x (p - p0) correctly maximises coverage and returns ~0.02 - fire
on every bar. The "p < p0 -> get more selective" branch, which is the
actual regime and the entire point of 983a6a3, could never fire because IS
never showed p < p0.
Carve a calibration slice out of the IS span - DIR_CONF_CALIB_PCT_OF_IS,
purged from backprop by one label horizon on BOTH sides (the far-side
purge is not optional: without it the newest training bars carry labels
partly decided by price action inside the slice, putting the memorization
straight back into the curve). Score it in a new chunked pass 2.5, after
pass 2 has trained and before pass 3 grades - the only position where the
histogram is simultaneously not-trained-on, not-graded, and current with
the weights it will be applied to.
Costs 15% of the training data. Worth it beyond honesty: the deploy gate
needs dirPrecPct > chance + EDGE_MIN_SIGMAS*SE, and a threshold pinned
near zero dilutes any edge concentrated in the confident bars across every
bar the model calls, driving dirPrecPct toward chance by construction. A
threshold that can be selective is the only mechanism by which a small,
concentrated edge could ever clear that gate.
Also: a sparse histogram now KEEPS the previous threshold instead of
resetting to 0.0. A failed measurement must not decay to the most exposed
setting in the range.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
SignalClusterWindow 3 -> 10 for all topologies. On H1 a 3-bar window
collapsed only the tightest runs and left visible clusters at every
turn; 10 bars is closer to the spacing of genuinely distinct setups.
ALTERNATION. Rule 1 only collapses a same-direction run INSIDE the
window; past it a second Buy is emitted with no Sell between, giving
Buy/Buy/Buy/Sell. With both directions tradeable that sequence is the
model re-entering a move it is already in rather than finding a new
one. The kept sequence must now alternate: the first signal passes,
and after that a direction passes only if the last KEPT signal was the
opposite one.
Added to ALL THREE consumers, with identical logic, because they must
agree:
- NmsLiveAccept -> the live trade
- pass 3's OOS replay -> the tally the deploy gate grades
- PruneDirectionalClusters -> the drawn history
A rule applied to only some of these certifies one strategy and trades
another - the same defect class as the geometry the gate certified
while OpenParams placed something else (9a7c37f) - and would draw the
user arrows the EA would never have taken.
Deliberately NOT applied to the LABEL. The barrier target has no "must
flip" invariant: consecutive Buy labels are routinely correct, and an
earlier alternation gate was removed with the triple-barrier relabel
for exactly that reason. This filters what is ACTED ON, which is what
"applies to training" can honestly mean here - pass 3's declustered
tally is the training-side number that decides deployment.
BothDirectionsTradeable() is the stated precondition (with one side
disabled there is no opposite to wait for, so alternation would
suppress everything after the first call). This build has no
long-only/short-only input, so it is constant true - kept as a named
predicate so a future direction restriction has one place to change
rather than three call sites silently assuming both sides.
Build tag -> nms-alternate-v4.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
FitDirConfThreshold walked from the most selective bin down to bin 0
keeping `precPct >= bestPrec`, with the stated intent that a plateau
should walk toward more coverage. The failure mode is the models that
need a threshold most: a net with no edge scores its base rate at
EVERY threshold - a perfect plateau - so the walk ran all the way to
bin 0 and returned 0.0, i.e. fire on every bar.
Reported as PAI "overshooting signals" while the other three stayed
selective. PAI has the flattest plateau because its margin
distribution is the most degenerate: its OOS outputs span the full
0.000..1.000 where CONV sits at 0.214..0.814, so nearly every call
lands in the top bins and precision barely moves as the walk descends.
The deeper problem is that precision is not the money quantity. For a
k:m barrier with p0 = m/(m+k),
EV = (p - p0) * (k + m) => EV per bar = coverage * (p - p0) * (k+m)
and (k+m) is constant across thresholds, leaving coverage * (p - p0).
That objective needs no tie-break and behaves correctly everywhere:
p > p0 everywhere -> takes the coverage (the old outcome, now for a
reason rather than as a plateau artifact)
p flat at p0 -> every point scores 0, the coverage floor decides
p < p0 everywhere -> the LEAST coverage loses the least, so it gets
MORE selective instead of trading everything
The last case is the current reality for all four models (-1 to -4pp
against break-even) and is the exact opposite of what the old rule
did. The comparison is sound: the histogram is already fitted on wins
(qTradeWon), not label agreement, so precision and break-even measure
the same quantity.
Ties now keep the more selective point - the loop reaches it first and
the test is strict >.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The risk/reward explanation was a second, wrapped line and cost more
vertical space than it earned - that detail belongs in the journal,
where the geometry is already logged in full.
The comparison itself stays, in two words: "64% (unseen data, need
67%)". Without it the win rate reads as skill when it is the barrier
geometry's own base rate, which is exactly how four models sitting at
chance came to look like four models at 65%.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Both made the run structurally unable to succeed, independently of any
signal in the data. Found by reading the 13:01 log.
RECALL GATE. m_objectiveMet required Buy, Sell AND Neutral OOS recall
each >= 40%. First-touch resolution (ce52654) collapsed Neutral from
the ~94% majority it was under exact-pivot labels to a same-bar-tie
residue - 250 of 38,261 bars, 0.65% - so the floor was asking the model
to identify 40% of coin-flip ties before it could converge. Measured:
CONV, LSTM and HYBRID all logged "Neutral:0% (need >=40% each)" on
every era. No model could ever satisfy it; every run was destined for
the plateau ladder or the era cap.
Only the DIRECTIONAL floors are load-bearing for the anti-collapse job
the gate exists to do: an all-Neutral model shows Buy and Sell recall
at 0% and is blocked by them. Neutral's own floor guarded the mirror
bias (over-calling Buy/Sell at Neutral's expense), which was real at
94% prevalence and is not at 0.65% - there, almost never calling
Neutral is correct rather than biased.
Prevalence-guarded rather than hardcoded off, so it returns by itself
if a future label rule makes Neutral substantial again. Deliberately
NOT extended to Buy/Sell: exempting a thin directional class reopens
the era-44-46 hole, which directionalRecallMeasured only half-covers -
it checks those classes were MEASURED, not that they passed.
ETA DECAY. A regressing era restored the checkpoint, reset the
optimizer and cut eta - all on the FIRST regression. The next era then
started from an identical state with a smaller step, regressed again,
and got the same treatment. The loop is self-sustaining and cannot
discover anything, because rolling the weights back is exactly what
removes the exploration that would end it.
Measured on PAI: eras 2-11 every one a regression against era 1, eta
0.000594 -> 0.000024, dW/W 0.000%/0.000% from era 2 onward. Ten eras,
~45s each, reproducing era 1 exactly and unable to do anything else.
Now requires ETA_DECAY_PATIENCE_ERAS consecutive regressions - the
standard ReduceLROnPlateau formulation. A single bad era is noise, and
an improving era clears the counter so alternating runs never
accumulate into a decay.
Build tag -> gate-patience-v3. It had not moved in six commits, which
is why the running binary could not be identified from its own log.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Chart objects live in the MT5 chart PROFILE, not in this EA's files.
They survive a terminal restart, a recompile, and deleting every
.nnw/.cfg/.stats/.arrows on disk. Only a deinit that RUNS TO COMPLETION
removes them - and MetaTrader force-terminates OnDeinit at roughly
4,500 ms, so a run killed mid-cleanup orphans them permanently with no
owner left to clean up after. That is the "deleted every file,
recompiled, restarted, old arrows and a stale panel still there"
report: nothing was wrong with the files and deleting them could not
have helped.
Both halves are fixed.
STOP OVERRUNNING THE BUDGET. OnDeinit used to finalise the in-flight
run (StopTraining -> FinalizeTrainRun: checkpoint restore, live-state
re-seed) and then write two full nets per chart. On four charts that is
the bulk of the budget, spent to preserve a PARTIAL era that was never
scored, never checkpointed and never deployable. FlushTrainRun()
discards it instead - drop the resumable bookkeeping, leave the net
neutral (unfreeze BN, flush the batch, batch size 1), skip the save -
and training resumes from the last completed era, which the era-end
save and the periodic autosave have already put on disk. What is
discarded is bounded by one era.
A CONVERGED model keeps the old finalise-and-save path: its weights can
carry online-learning updates made since the last era boundary, and for
a deployed model no further era boundary is coming to persist them.
MAKE CLEANUP SELF-HEALING. Every purge sat behind a branch - no model
loaded, sidecar missing - so the common paths returned leaving whatever
the previous instance stranded. LoadChartSignals now sweeps the arrow
namespace unconditionally before restoring, so the post-init chart
holds exactly what the sidecar holds whichever branch runs, and the
panel gets the same treatment before Create() (CAppDialog namespaces
its controls, so a killed Destroy strands the lot and the next attach
draws a second panel on the corpse).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Corrects the premise of the previous plan. Break-even is NOT a ceiling.
If the model shifts the win probability on the bars it selects from
p0 = m/(m+k) to p0 + d, then
EV = (p0+d)*k - (1-p0-d)*m = d*(k+m)
because p0*k - (1-p0)*m is zero by construction. The stop:target RATIO
is expectancy-neutral - a punishing break-even is exactly repaid by the
payoff - and only the real edge d and the TOTAL WIDTH (k+m) move EV.
Width matters because the spread is charged once per trade however wide
the barriers are, so a narrow barrier spends much of its own range on
costs. DeriveBarrierGeometry's own comment already said the ratio buys
nothing; the objective just never followed from it.
Blocker this had to solve first: m_excUpCache/m_excDownCache hold only
MAXIMUM travel each way, and a maximum cannot say which side was
reached FIRST - so any geometry other than the walked one was
undecidable on precisely the bars where both barriers were touched,
~28% of the sample.
- BARRIER_LADDER: per bar, the first-touch AGE for 8 travel distances
in each direction, filled during the walk the labels already run.
Cursors keep it O(1) amortised per walked bar rather than 16
comparisons. Levels are travel FROM ENTRY, not barrier prices, so one
ladder serves both directions and the spread is applied analytically
when a level converts back to an SL/TP multiple - storing prices
would need four ladders and bake today's spread into the cache.
Sized, invalidated and validity-gated with the label caches.
- ReportGeometryExpectancyScan: every ladder pair priced exactly off
that cache - width in ATR and in SPREADS (cost efficiency, knowable
without knowing d), break-even, both base rates, the share of bars
resolved inside the horizon, and EV per unit of edge. Compares the
widest resolvable pair against the quantile rule's pick.
MEASUREMENT ONLY - the quantile rule still chooses. Nothing here can
measure d, and width buys nothing if the wider target is less
predictable. Base rates are printed beside each break-even because a
persistent gap is DRIFT and must not be credited to the model.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
"Buy/Sell calls correct: 65% (unseen data, lifetime)" is a WIN RATE -
m_cumOosCorrect advances on oTradeWon, did the implied trade reach its
target before its stop - not label agreement. A win rate means nothing
without the barrier that produced it.
With the measured geometry a trade risks slMult*ATR to make tpMult*ATR,
so under a driftless walk ANY directional call wins
slMult/(slMult+tpMult) of the time for free. On the shipped 3.33/1.62
pair that is 67.3%, and the empirical long-win base rate on this window
is ~65.7%. All four topologies read 65%: at chance, and below
break-even, while the panel announced "65% correct".
Four models with completely different trade counts agreeing on one
number was the tell - a win rate fixed by the geometry rather than
produced by the network. The deploy gate already benchmarks against
this null (chancePrecPct); the panel did not, and the panel is what a
buyer reads.
The line now carries its own break-even and the risk/reward that sets
it, so the number can never again be read as edge on its own.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
BufferTempData cached EVERY failure - m_featureCacheHasValue[idx]=true
with m_featureCacheValid[idx]=false - and the cache never re-tries a
miss. So a single feature read taken before the terminal had finished
calculating the indicator buffers marked those bars unusable for the
rest of the process, even though the data arrived milliseconds later.
MT5 fills an indicator's buffers asynchronously after the handle is
created, and a cold ATR returns 0 for EVERY index, not just its warm-up
tail. BufferTempDataCompute rejects a bar with no ATR (correctly - the
price features would be meaningless), so the whole window failed, and
the whole cache was poisoned.
Only resumed models were hit, because only they read features that
early. Topology.mqh sets m_warmupPassesRemaining = netLoaded ? 0 : 3:
a fresh start sits through three separately-scheduled Train() calls
before anything touches a feature, which is exactly what those passes
are for. A resumed one skips them and TuneIndicatorsAndTrain drives
StartLabelCachePrebuild and the MI report from the first chart event.
Its rationale - "a restart already has a proven-synced history" - holds
for HISTORY and not for INDICATORS, which are recreated every process
start.
Downstream: BuildFeatureWindow failed on every bar of every era, so
add_loop never went true, so pass 2, pass 3, the era counter and the
checkpoint were all skipped and pass 1 swept 0->100% forever. The
"0 samples" MI report line at startup was the same failure, four
seconds earlier, already visible in the log.
- a miss is now cached only when it is PERMANENT; the two "not ready
yet" guards mark m_featureFailTransient and are recomputed on the
next visit. Steady-state cost is ~ind_Periods bars per era, not 54k.
- an era that discards itself now drops the feature cache before
restarting, so any remaining cause of this state self-heals instead
of looping.
Deleting the .nnw "fixed" this only by turning the model back into a
fresh one.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
add_loop is exactly "at least one bar produced a usable feature
window". When it stays false, pass 2, pass 3, the era counter, the
checkpoint and every log line in the era-end block are ALL skipped:
Train() returns having done nothing, m_eraResumePending is still false,
and the next call restarts the SAME era from bar 0. That is an
infinite 0->100% "scan" loop that prints absolutely nothing - the only
remaining silent restart path in Train(), and it matches the reported
symptom exactly.
Pass 1 now counts usable vs unusable windows and reports at the pass
boundary, which demonstrably executes:
- total failure routes through ReportTrainStall (already capped at
one line a minute, and carries the run-state flags) naming the
counts, the required window width and the bar count
- success prints how long the scan took and how many samples it
handed to pass 2, but only once the era has passed 10s - a fast
era stays as quiet as before, a slow one distinguishes "advancing"
from "sweeping the same bars forever"
A PARTIAL failure is normal and deliberately does not shout: pass 1
walks oldest-to-newest and the deepest bars predate the indicators'
warm-up, so those windows fail and are cached as misses.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>