- Added bulk read/write methods for feature caches in IFeaturesView and its implementations to optimize performance.
- Introduced LabelCacheInvalidateAll method to manage label cache invalidation alongside feature cache.
- Implemented PooledIndependentBars method in topology interfaces to account for additional independent observations.
- Enhanced risk budget management with throttling for peak-equity updates to reduce unnecessary file operations.
- Improved error handling and logging for ATR trailing stops to ensure better visibility of issues.
- Updated alt-data handling to prevent unnecessary operations during testing and optimization phases.
THE OPTIMIZER ("0.1% an hour per agent", 0 of 39 passes in 78 min,
12 agents): the tester fires OnTimer on SIMULATED time, so the live
chart's 500ms EventSetMillisecondTimer over a 2016-2026 pass is ~600
MILLION OnTimer calls - each walking 4x PollTraining, the vote
readout's string build, the overlay advance and the deployed census.
None of it serves an inference-only pass: training never runs, per-bar
inference is driven by OnTickHandler off the tick stream, the risk
budget re-checks in OnTick, and there is no chart to keep fresh.
StepSetTimer now arms EventSetTimer(3600) in tester/optimizer/forward
(~2,600 calls per pass) and keeps the 500ms timer for live charts.
Plus a TESTER PASS SELF-PROFILE: per-tick buckets (pre / Expert.OnTick
/ journal) and the timer total, printed once at the pass's OnDeinit -
so if a pass is still slow it names its own consumer instead of being
diagnosed from outside.
OFFLOAD (operator: "as much calculation as possible to DLL/OpenCL"):
batch norm was the ONE stage still host-side on the DLL tier - the
device path was OpenCL-only, so every sample crossed the bus twice per
BN layer and normalized in interpreted MQL5 (and every model runs
batchnorm ON). Four new exports mirror AI\Network.cl's BatchNorm*
kernels 1:1 in DOUBLE precision (closer to the host reference than
the float OpenCL kernels): forward with running stats + frozen flag,
hidden gradient with the clamp derivative, gamma/beta accumulate, and
the batch-mean apply (no weight decay, moments-before-skip ordering,
sqrt-stored v). BnDeviceEligible/EnsureBnDeviceBuffers/all four
Dispatch* now route by backend; the EXISTING in-situ self-checks
(host-vs-device on the first real sample, latch-off + host fallback on
mismatch) verify the DLL kernels exactly as they verified OpenCL ones.
batch_accum_check regression: ALL CHECKS PASSED on the rebuilt DLL.
Same deployment coupling as bd46374: the .ex5 imports the new exports
- copy DirectML\WarriorCPU.dll into MQL5\Libraries (terminal closed)
together with the new .ex5, and re-copy it to the tester agents (or
just run DirectML\build_cpu.bat once with everything closed - it
deploys to every discovered Libraries folder).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
"Hundreds of times slower than a regular EA" decomposed into two
multiplied factors, both measured:
1. THE OPTIMIZER STEP RAN IN INTERPRETED MQL5. The CPU tier shipped
the F4 accumulate exports with deliberately no matching apply
(WarriorCPU.h said so), so on the DLL backend - this box - every
TRAIN_BATCH_SIZE=8 batch fell to the host loop in ApplyAccumToBlock:
a per-weight MQL5 pass through CBufferDouble.At()/Update() plus four
full weight-matrix BufferRead/Write round trips. The 2026-07-26
profile had already shown the per-sample Adam step at 81% of ALL
runtime (feedForward: 8%; feature building: 0.35%) - sqrt+divide
per weight vs one multiply-add; moving it into MQL5 made it worse.
New CPU_ApplyAccumAdam / CPU_ApplyAccumMomentum: one element-wise
ParallelFor takes the batch-mean step and zeroes the accumulator
DLL-side, generic over any flat block (dense/conv/LSTM/batch-norm -
all apply paths funnel through ApplyAccumToBlock, which now tries
the DLL first, with the same one-warning failure latch as the
OpenCL fast path). Math is the shipped step to the last clamp:
sqrt-stored v, ClampDelta, AdamW decay, ClampWeight.
batch_accum_check extended (check 6) and ALL PASS: apply == host
reference at B=8/B=4, accumulator zeroed, and B=1 accumulate+apply
== the unbatched Adam kernel BIT-EXACTLY (kernel-vs-kernel, no
transcription). DLL rebuilt with the shipped /fp:fast recipe.
2. A 24% DUTY CYCLE. Train sliced 120ms per 500ms timer period
(30ms/member x4), leaving the chart thread idle 76% of the time.
Now 300ms total (75ms/member): ~60% duty, ~2.5x, click latency
bounded at ~300ms while training runs - between the fully-reactive
120 and the documented "sticky drag" 480.
DEPLOYMENT COUPLING: the new .ex5 #imports the new exports, so it will
NOT LOAD against the old WarriorCPU.dll ("cannot find function"). Copy
DirectML\WarriorCPU.dll into MQL5\Libraries (terminal closed) in the
same step as deploying the new .ex5.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The 18:23 terminal close (20260825.log) killed two of six charts inside
OnDeinit: they printed "shutting down" then nothing for 5.9 s until
"Abnormal termination", stranding ~700 objects each - including the one
family no prefix sweep can reach, the control panel (CAppDialog names
its 15 objects <numeric instance id><control>, and a re-attach mints a
new id, so a killed panel is a permanent ghost; XTIUSD carried one
across sessions). The stall sat in the two file writes that preceded
all visible cleanup while the four sibling charts flooded the same
2013-era disk - the ~4x18MB-per-chart shutdown weight saves.
Three changes:
1. OnDeinit touches no file until the chart is clean. CVoteArrowStore
splits Save() into Snapshot() (the chart scan, in memory) and
WriteSnapshot() (the disk half, consuming). New order: status label,
vote-arrow snapshot, prefix sweep, panel destroy - all object ops -
then member sidecars, final sweep, timings, and only then the
visibility file, the vote-arrow write and the weight saves.
2. PurgeOrphanedPanelObjects() at OnInit: deletes numeric-prefix
CAppDialog ghosts by name (6 chrome + 9 buttons), qualifying a
prefix only when >=4 of OUR button names carry it, so a foreign
dialog sharing stock chrome names is never touched.
3. m_netDirty: set by every net mutation (both backProp sites, both
RestoreWeights sites, online learning conservatively, panel reset),
cleared only on a successful Net.Save. Shutdown AND the per-bar
autosave now skip the ~18MB write when the net is provably unchanged
- for converged ensembles that is every save - which removes the
very flood that starved the sibling charts. .stats still writes
every time (small; carries the vote record and calibration). A
skipped save leaves the .nnw header dtStudied stale, which is the
already-handled attach-after-offline-gap case.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
"Vote win rate: measuring..." never resolved on a deployed chart whose
.stats predate the WST7 ensemble record: g_ensCumOosTotal is fed only by
the era-end combined-vote scorer (Training.mqh), and a deployed ensemble
runs no further eras. The replay pass rebuilt every MEMBER's ladder
(64-71% each, per the 16:12 log) but nothing ever scored the COMBINED
vote, so the aggregate line sat on "measuring" while 300+ arrows drew.
The overlay sweep already reconstructs the vote per bar with the live
threshold and direction policy - so it now also tallies, BEFORE
declustering (NMS thins arrows, not calls), each threshold-clearing bar
against the inline swing-pivot label (same resolution ScoreReplayFromCache
uses, same window-mismatch reason). On sweep completion Warrior_EA.mq5
harvests the tally through a consuming one-shot read and adopts it ONLY
when the record is empty and the models are deployed - a training-time
sweep can never pre-empt the era scorer, and a restored record always
wins. The result is persisted immediately into every member's .stats.
Also verified against the same log: the sweep does NOT ignore
DrawUnfilteredSignals - 4986 voter bars -> ~300 arrows, all gated on the
25% open threshold. The arrow increase vs the restored set (41-312 saved)
is the replay-minted ladder reading stronger (partly in-sample), plus the
reconstruction deliberately not replaying order validation/session hours
(tooltip says so); the backfilled record carries the same caveat and is
labelled so in the log.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Drops "peak N%", "need N%" and the "armed (bar still open)" middle
verdict state - three pieces that were useful while tuning
Signal_ThresholdOpen but add nothing once a chart is settled and running.
m_votePeak is still tracked (nothing programmatic reads it via this
line), just no longer printed.
The verdict collapses back to two states: "training, not tradable yet"
(undeployed) or TRADE/no trade (deployed) - fires is already forced
false on a prospective vote, so "no trade" falls out for a bar that
hasn't closed without a separate word for it.
Compile-verified in _claude_stage: 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The 15:13 session proved the replay pass ran end-to-end on all 24 models
and scored ZERO labelled bars on every one of them, while each rescan sat
on ~5000 scored predictions (~2755 Buy / ~2232 Sell). The two windows
never overlapped:
StartLabelCachePrebuild deliberately keeps a CONVERGED model's
dtStudied watermark (it gates inference recency and must not move), so
the prebuild's window was the handful of bars since the last studied
bar - all with uncommitted pivots, hence "label cache pre-built -
Buy: 0 | Sell: 0 | Neutral: 0" on every member.
The label never needed a cache. SwingPivotDirectionLabel(idx) is a pure
function of the ZigZag/Close/ATR buffers the rescan itself refreshes over
exactly the scoring window, and m_lastLabelLifespan == 0 is its own
unresolved flag - the same finality gate the cache applies, applied
directly. ScoreReplayFromCache now resolves each bar's label inline and
the label-prebuild stage is deleted from the rebuild state machine
outright; going through a cache built for a different window was
indirection that changed the answer.
Also splits the empty-result diagnostics: "no resolved labels" (a
windowing/data fault) is now distinguished from "labels present, every
call Neutral" (a calibration verdict). The first version reported the
second message for both, which mislabelled this very bug as a calibration
outcome in the same breath as reporting scored=0.
Honest limitation, stated in the code too: the replay window includes
bars the model trained on, so a replay-minted ladder is measured partly
in-sample and will read stronger than a holdout-measured one. It is
replaced by the genuine article at the next completed scoring pass; until
then it is what makes a restarted deployed model able to vote at all.
Compile-verified in _claude_stage: 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The previous commit persisted the tier ladder, which fixes this going
forward but did nothing for models whose .stats predates WST7 - they
still had to retrain to mint one. They never did. Every number a
converged model needs in order to vote is a pure function of weights
already on disk plus labels derivable from the chart, so replay them:
stage 1 build the label cache (existing chunked prebuild)
stage 2 rescan history (existing chunked rescan, deployed net)
stage 3 score + rank + persist (one walk over two arrays)
ScoreReplayFromCache() walks m_arrowSignalCache against
m_labelCacheBuy/Sell, fills the same m_oosTierFired/Hits and per-class
totals pass 3 fills, and hands them to RankTiersFromOos() - deliberately
feeding the existing ranker rather than reimplementing it. The shrinkage,
the chance reference and the module trust weight are subtle enough that a
second copy would drift, and a ladder measured by a slightly different
rule would be silently incomparable with every ladder training produced.
AdvanceDeployedRebuild() sequences the three stages off the timer. It has
to be a sequence: stages 1 and 2 are each minutes of work draining in
time-boxed slices, and stage 2's output is meaningless until stage 1 has
labels to score against. The previous version ran the rescan with no
labels at all, which is why it could only ever rebuild arrows and never
the ladder - the thing actually blocking the vote.
The result is written to .stats immediately. The failure being repaired
is state that lived in memory and was never written down; recomputing it
and not saving it would repeat that exactly.
Also routes every rescan completion through one hook, so there is a
single place that knows what a finished rescan means - republish for a
manual one, score and rank for a rebuild.
Compile-verified in _claude_stage: 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
THIS IS NOT A DISPLAY BUG. A deployed model could not vote, or trade, at
any point after a terminal restart, and never would have.
LiveVoteContribution() returns 0 for every call until m_tiersSelfRanked
is set - deliberately, and correctly: before RankTiersFromOos() runs,
m_pattern_0..3 hold the constructor's stock 25/50/75/100, which since the
2026-08-18 currency change is the WRONG UNIT rather than a weak opinion,
and one unranked member would drag the whole ensemble over any threshold.
But that ladder is produced ONLY by a completed pass 3, and it was never
persisted - the code comment at LiveVoteContribution says so outright.
A converged model runs no further passes. So on every restart it lost its
entire vote permanently:
LiveVoteContribution -> 0 => no live vote ("0 vote/4 flat")
ReconstructionWeight -> 0 => overlay divisor 0 ("0 had a snapshot")
=> no arrows
=> no fired bars, so g_ensCumOosTotal stays 0
=> "measuring..." forever
Every symptom reported over the last three exchanges is that one cause.
The log is unambiguous: six H4 charts resumed at era 70/71, all 24
rescans completed with ~2700 Buy / ~2200 Sell per model, and the overlay
then swept 4999 bars finding "0 had a snapshot". The calls were there;
nothing was permitted to count them.
WST7 now stores the four tier weights, the module trust weight and the
self-ranked flag beside the model. Restored only when the stored flag
says the ladder was MEASURED - a .stats written before a model's first
pass 3 holds the stock ladder, and adopting that as if measured is the
exact error the flag exists to prevent.
A .stats predating WST7 has no ladder, so existing converged models stay
silent until their next scoring pass mints one. That case now prints a
warning naming all three of its symptoms, because each one independently
looks like a different bug.
Compile-verified in _claude_stage: 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The sidecar added in 484a9d8 restores the vote arrows from the previous
session - but there was no previous session to restore from, and a
deployed ensemble could never produce one.
The overlay that draws the vote layer replays each member's
m_overlaySigSnap, published in exactly one place: RankTiersFromOos, at
pass-3 completion. A converged model runs no further eras. So after a
restart every member's snapshot was empty, would never fill, the sweep
had nothing to replay and the chart stayed blank permanently - no route
back by any path.
The chart rescan is the route: it runs the DEPLOYED net forward over
history and rebuilds the per-bar cache, which is the same quantity pass 3
produces, obtained without training. It already existed for the panel's
Show-Signals button; it just never handed its result to the overlay, so
on the default filtered view a rescan rebuilt only the RAW per-member
layer - the one that is hidden - and appeared to do nothing.
- PublishOverlaySnapshotFromCache() extracted from RankTiersFromOos, so
the era end and a completed rescan publish through one implementation.
- A completed rescan now calls it, which also arms the sweep.
- PollTraining auto-arms one rescan for a model that is converged, has no
snapshot, and is on the filtered view. One-shot: a model that
legitimately calls Neutral everywhere must not rescan forever chasing a
snapshot that is correctly empty. On the timer, not in OnInit - it is a
full feedForward per bar over up to 5000 bars and drains in the same
time-boxed slices as a manual rescan.
Together with the sidecar this closes both halves: the rescan covers the
first session and any chart whose file was lost or invalidated by a
threshold change; the sidecar covers every session after one is saved.
Compile-verified in _claude_stage: 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Four reported symptoms, three of them one root cause: the ensemble's
certified record was session-scoped and written ONLY at pass-3
completion. A deployed ensemble runs no further eras, so every restart
lost the aggregate win rate, the aggregate panel line and the overlay
snapshots - and could never regenerate them, because regeneration only
happens at an era end that will never come.
THE SELF-CONTRADICTION. Member rows read "Live - learning from new bars"
(from m_trainingComplete) while the line under them read "training, not
tradable yet" (from `prospective`, which means "this number came from
ProspectiveVote() rather than a real Direction() call" - what happens on
any bar where every member abstains, and which says nothing whatever
about training state). Both now resolve through one predicate:
WarriorChartModelsDeployed(), fed by members publishing their own state
on the same slot and cadence as their vote. Adds a third verdict word,
"armed (bar still open)", for a deployed model on a prospective
recompute - the case that used to claim it was training.
DEPLOYED PANEL. Once every published model is converged the per-member
rows are dropped: what ships is the aggregate vote win rate, the live
vote, and the verdict. While training the rows stay - they are the only
way a collapsed or lagging member is visible, since a collapsed member
abstains and so is invisible in the aggregate by construction.
ACCURACY NOW RESPECTS THE ENTRY THRESHOLD. The panel's "precision 65%"
came from m_cumOosCorrect/m_cumOosTotal, which counts every bar a model
called Buy or Sell - threshold-blind, and per-model rather than
per-vote. The correct number already existed (votePrecPct: bars where
|vote| >= threshold and the direction policy allows) and is now what the
panel shows, with the threshold named in the text because the number is
meaningless without it.
VOTE ARROWS PERSIST. With DrawUnfilteredSignals off - the default - the
chart shows SIG_VOTE_PREFIX arrows, and nothing saved them:
CChartUI's .arrows sidecar is member-scoped and never saw that layer.
New CVoteArrowStore mirrors them to a chart-keyed sidecar and restores
them progressively at init, on the same budgeted non-blocking path.
The header stores the open/close thresholds; a mismatch on load DISCARDS
the arrows rather than redrawing a picture of a strategy no longer
configured - stale arrows are worse than none, because none is visibly
empty and stale is confidently wrong.
Also: .stats bumped to WST7 carrying the ensemble record (guarded on
threshold match, most-complete-copy-wins), and the loader's version
tests collapsed from an or-chain to ">=" - the magics are ASCII 'WST1'..
'WST7' so they are already ordered, and a missed arm in that chain reads
the NEXT field's bytes into this one, which fails as plausible numbers
rather than as an error.
Compile-verified in _claude_stage: 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This repo now holds only EA (MQL5) sources. The Python research scripts and the
third-party MQL5/PDF reference material live in a sibling workspace folder,
..\Warrior_Research\, with their own git repo (initial commit ec2214a there).
Nothing in the EA depends on either folder at build or run time, and the research
scripts address ..\Market Data\ and the MetaTrader Common\Files directory by
absolute path, so the relocation breaks no path. EA comments that cite scripts by
name (research/edge.py, research/altdata/export.py, research/test_spread.py, ...)
stay accurate - only the parent folder moved.
.gitignore drops the two rules that only existed for the moved trees
(references/*.pdf, research/edge_rows.npy); they were carried over to the new repo.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Refresh() returning false skipped the whole of Processing(), and
Processing() is where CheckClose() and CheckTrailingStop() live. So on
any tick with unusable quote history, a failed RefreshRates(), or a
period-flag mismatch, an already-open position got no exit check at all -
it rode. Invisible by construction: nothing logged, no order sent, and
next tick the position looks exactly as it should. The only trace is a
stop that should have moved and didn't.
ProtectOpenPosition() now runs the CLOSING half of Processing() on those
ticks. Only the closing half, on purpose: CheckReverse() and the
pending-order block both OPEN exposure, and opening on data just declared
unfit to trade on is the opposite of the point. Closing on an imperfect
quote reduces risk even when the quote is wrong; opening on it does not.
When it acts, it says so in the journal - a degraded-path exit should
never be silent.
Also pins the invariant at the Expert_EveryTick gate: that input
throttles how often the EA forms an OPINION, never how often it can act
on a position it already holds. Exits stay above the gate, and the
comment now says so to the next person editing it.
Pre-existing hole, not introduced by the EveryTick work in 6d48fdb -
that change is what made it worth reading the tick path closely.
Compile-verified in _claude_stage: 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Records the rule the removal exposed: validation catches an enum value that no
longer exists, but not one that silently now means something else. Also drops
the stale claim that SL_Mode/TP_Mode define the training target - the
swing-pivot label is geometry-free.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Five modes went, all of them staking real risk on the model's confidence:
Intelligent entry (ENTRY_INTELLIGENT), stop (SL_INTELLIGENT), target
(TP_INTELLIGENT), trailing (CTrailingIntelligent) and lot size
(CMoneyIntelligent's quarter-Kelly). With them, the Confidence_Source
input and the CONFIDENCE_SOURCE enum, whose only job was choosing which
number those five read.
The reason is calibration, not correctness: the confidence magnitude is
known to be miscalibrated against the label prior, so every one of these
modes multiplied money by a quantity whose units were never established.
The DB arm had a second, independent defect - since the tester DB guard
(SignalDatabaseActive) it reads 0 in tester and optimizer but non-zero
live, so any backtest of CONF_DB/CONF_BLENDED could not reproduce live
trading. And what the DB produces is a filter-RANKING win rate, not a
per-trade win probability.
Both confidence numbers are still recorded per trade (aiConfidence /
dbConfidence) and still bucketed against outcome in TradeJournalReport.
Recording is what keeps the question answerable; acting on it was the
part with no evidence behind it. ConfidenceBridge.mqh now carries an
explicit telemetry-only rule at the top.
ENUM ORDINALS PINNED. Removing a member vacated a value in four enums at
once and MT5 does not validate an enum input replayed from a saved .set
or a stored optimization pass. TRAILING_STRATEGY and
MONEY_MANAGEMENT_STRATEGY now carry explicit values so the survivors keep
the numbers they were saved as, and ValidateBarrierInputs is widened into
ValidateTradeManagementInputs covering SL_Mode, TP_Mode,
Entry_Multiplier, TrailingStrategy and MM_STRATEGY. Without that gate a
chart saved with the Intelligent stop would feed SL_Mode = -1 into a
multiplier now used verbatim, placing the stop on the wrong side of entry.
RETRAIN-NEUTRAL: neither SL_Mode nor TP_Mode appears in
BuildModelFingerprint() or ComputeDbConfigFingerprint() since the
swing-pivot target replaced the barrier labels. No .nnw, .cfg or .db
re-keys. Also drops the now-dead g_TradeRewardRiskRatio bridge, the
CMoneyRiskBase::AdjustRiskAmount hook and the unsigned AIConfidence().
Compile-verified in _claude_stage: 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Three related changes, all aimed at work being repeated at a frequency
nobody chose.
1. OnDeinit gets a tester/optimizer fast path.
Everything in the live teardown exists to leave a CHART clean and a live
model's state on disk. An optimization agent has neither. It was still
running, on EVERY pass: a per-signal arrow-sidecar WRITE
(ShutdownChartCleanup -> PersistAndClearChartSignals) plus two full
chart-object scans plus a ChartRedraw. At optimization scale that is
hundreds of thousands of pointless file writes per agent, against a
~4,500 ms budget MetaTrader force-terminates on - the shape of thing
that stalls an agent rather than failing it.
The fast path keeps MarkShutdown() and FlushTrainRun() (so a killed pass
never leaves a half-written era) and still calls dbm.Deinit() and
Expert.Deinit() - leaking the signal tree or a handle across passes is
its own way to accumulate into a stall. The two now-unreachable
!isTesterRun guards further down are folded away.
2. All four tester handlers are present and documented by WHERE THEY RUN.
OnTesterInit/OnTesterPass/OnTesterDeinit run in the CONTROLLING TERMINAL
once per session; only OnTester runs on the agent, per pass. OnTesterPass
was missing entirely - added empty and deliberately so: it only fires for
passes that shipped FrameAdd() data, which this EA never sends, and
reading frames there would put per-pass work on the terminal's critical
path. Declared so that adding frame-sending later fails loudly instead of
silently dropping every frame.
3. Expert_EveryTick is now actually enforced.
It was passed to Expert.Init() and only ever reached StartIndex() - which
bar a signal READS. The whole pipeline still ran on every quote. It now
gates m_signal.SetDirection() in CExpertCustom::Processing(): that call
drives Direction(), which is a TRANSACTION (NN forward passes, DB rows,
chart arrows, one-shot vote state), and re-running it on every tick of a
4-hour bar repeats all of it.
Scoped deliberately. Everything after that line still runs per tick -
CheckReverse/CheckClose/CheckTrailingStop and pending-order maintenance
are risk management, and a stop that only trails at bar boundaries is a
different strategy, not a faster one. The scheduled close-all in OnTick()
matches a +-1 MINUTE window, so bar-gating it on H4 would step straight
over the thing 100% of label timeouts already resolve against.
g_riskBudget.Update() also stays at quote frequency, by design.
System/NewBar.mqh becomes CNewBar, a class. The free function it replaced
had zero callers and kept its watermark in a `static`: ONE watermark
shared by every caller, so the first caller each tick consumed the
transition and every other caller was told "no new bar" for a bar that
had just opened. Per-instance state fixes that; first observation counts
as new, so a fresh attach acts immediately instead of idling up to a full
bar.
Compile-verified in _claude_stage: 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
~2,300 lines. META had real, repeatedly measured ranking skill and ZERO
operating points that ever cleared break-even (0/350 H1 eras, 1/999 H4
pre-2-sigma, 0/8 pooled fitted points). The clinching arithmetic was edge x
width = 0.095 ATR/trade against spread 0.099 ATR/trade, and the
dose-response showed the high-conviction tail is temporally unstable -
the precision-vs-threshold slope flips sign between calib and test on 3 of
4 symbols, so no ex-ante threshold rule exists. It shipped default-off and
never gated a live entry. The self-measured tier weights are what actually
rank the vote, and all six H4 instruments converged on them alone.
RETRAIN-NEUTRAL, and that is the property that made this safe:
- The weights fingerprint emitted "|TGT:META2" or "|TGT:SWG1" from an
if/else. Every direction model already took the SWG1 arm, so
collapsing it to an unconditional append is byte-identical. No .nnw or
.cfg is orphaned or re-keyed.
- NetInputWidth() lost its "+ MetaDescWidth()" term. MetaDescWidth()
returned 0 for every direction model, so the input layer is unchanged.
- DbLegacyAiSlot()'s slot 5 was reachable only with all four Use_* NNs
off AND meta on - a config that never shipped. Every existing .db keeps
its filename.
Deleted outright: Signals/SignalMETA.mqh, Expert/Trading/MetaGate.mqh (the
directory is now empty), Expert/Training/{MetaCorpus,MetaCandidateStore,
MetaFamilies}.mqh, Tests/Test_MetaFamilies.mq5, Meta_Labeling_Design.md.
Unwound in place, the delicate part: Training.mqh carried four
IsMetaTarget() branches whose else-arm WRAPPED the direction body (pass 1
queueing, pass 2 backprop, pass 2.5 calibration, pass 3 OOS scoring). Each
wrapper is removed and the direction body promoted back to its original
nesting - the bodies were never re-indented when the wrappers were added,
so the promoted code is byte-identical to what ran before META existed.
Also gone: the ensemble verdict's meta-veto replay and its
approved/vetoed/unscored counters, the per-family/per-side OOS
decomposition arrays, the m_isTrainQueueCand parallel queue and its
lockstep shuffle, and the S2 era report.
Also removed: the CMetaGate abstraction and the live CheckOpenPosition
veto; m_gates plus AddFilter's non-voter routing and IsVotingSignal()
(META was the only non-voting child, so m_gates was always empty);
m_parentSignal/SetParentSignal (existed only to reach the root's gate);
SweepPrepare/SweepPrepareIndicator (only caller was the corpus sweep);
IsMetaTarget() from all four view interfaces and their adapters;
Use_MetaLabeling, EnableMETA, Meta_ExportDataset, m_trainTarget.
EvalShift is KEPT - HistoricalNetVote() uses it for the filtered overlay,
not just the corpus sweep; only its comment changed. The 2-output softmax
arm in NetForward.mqh is kept too: it costs nothing and is the reusable
binary-head path, now commented as unclaimed rather than as META's.
Compile-verified in _claude_stage: 0 errors, 0 warnings, matching the
pre-edit baseline.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Two removals of work that a backtest was paying for and never using.
1. SignalDatabaseActive() gates the signal DB off in tester/optimizer.
A backtest opened the fingerprinted SQLite DB under FILE_COMMON - and so
did every parallel optimization agent, against the same file, with the
per-tick journal Update() behind them. Measured 2026-08-25 on a 12-agent
SP500 H4 run: zero passes completed in 75 minutes.
It bought nothing, for a reason specific to this EA's current shape: the
DB's only effect on a trading decision is ApplyPatternWeight overriding a
filter's module weight, and that is declined for any self-ranking filter
(CExpertSignalCustom's !filter.SelfRanked() guard). The AI members
self-rank once their tiers are measured, and the classic votes that DID
consume the ranking are gone - so a tester run's DB was written and never
read. Skipping it changes no decision.
One predicate, not two inline guards: OnInit asks the question twice
(InitDatabaseAndJournal, then VerifyDatabaseTransactionCycle) and a run
where those disagreed would try to open a database it never initialised.
The tester now takes journal.InitTrackingOnly(), so close detection,
MAE/MFE and the expectancy-stop feed still run - only the SQLite half is
dropped, and Update() already skipped its INSERT when there is no DB.
Caveat recorded at the predicate: if a future filter consumes DB ranking
WITHOUT self-ranking, this needs revisiting - a backtest would then stop
reproducing live.
2. ExportFeaturesOnly and its two exporters are gone.
Research-only CSV dumps (feature matrix + a hardcoded 8-symbol x 5-TF raw
rates grid), superseded by the research/ python path that reads its own
data. Removed the input, m_exportFeaturesOnly, the setter, both method
declarations, ExportFeatureMatrix()/ExportRawRates() (111 lines in
AutoTune.mqh), the OnTick early-return, and the ctor initialiser.
The config-lock bypass it owned collapses to the plain tester test:
`if(!inTesterOrOpt && !AcquireConfigLock())`. Shared helpers it called -
ServableBars, EnsureBarCachesCapacity, ResizeBuffers, RefreshData - all
have other callers and are untouched.
Compile-verified in _claude_stage: 0 errors, 0 warnings, identical to the
baseline taken before either edit.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Use_Training_Pool gates a fully-built, fully-wired mechanism
(Expert\Training\TrainingPool.mqh + the Add/Adopt/Publish call sites
already in Training.mqh) that shipped false. Nothing to build - the
writer/reader/atomic-file/compat-gate/age-gate/lookahead-purge were
all already there, measured +2.02pp of paired skill at H4 (research/
edge.py, 2026-08-24). Flipping the default is the whole change.
Compile: 0 errors, 0 warnings (stage).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Two chart-display fixes reported after watching a converged 4-model
ensemble: the ensemble panel's trailing "(era 69, 4 models,
DEPLOYING)" was frozen at whatever era the ensemble happened to
deploy on, and the separate top-right HUD (one line per model, raw
B/S/N + weight + era + error) was clutter once the vote itself is
what matters.
Root cause of the freeze: g_ensembleVoteLine is written once per era,
at pass-3 completion. A deployed/converged ensemble runs no further
eras (ScheduleTrainingIfNeeded's trainingComplete branch skips
Train() entirely), so that line could never update again - the era
count and "DEPLOYING" marker were permanent set-dressing from the
deploying era, not a live reading.
- EnsembleScoreCombinedVote() drops the era/DEPLOYING tail once
g_ensDeployApproved - nothing left there worth freezing.
- UpdateVoteReadout() (the aggregate "VOTE ..." line, previously its
own top-right chart object) now writes g_liveVoteLine instead of
drawing anything. Both status-label builders - PublishEnsembleStatus
for the ensemble panel, PublishStatus's choke point for the solo
panel - append it as one line, refreshed every tick/timer exactly
as the old HUD was, so the live vote replaces the frozen era tail
in the same visual slot.
- RefreshVoteReadout()'s per-member loop (DisplayHudLine, one
ObjectLabel per model) is deleted outright rather than folded in -
the operator asked for the aggregate only, "without telling me each
individual network".
Follow-on dead-code removal, since DisplayHudLine was the only
caller: the DispProb/DispSignal/MetaGateArmedNow/MetaHasScore/
MetaLastP/MetaLastBe/MetaApproved/MetaVetoed leg of IChartView (and
its AIBaseChartView/AIBaseChartViewImpl/ExpertSignalAIBase forwards)
had no other reader. The underlying data survives untouched -
m_metaTelemetry is still populated live by SignalMETA.mqh,
m_dispSignal still feeds ProspectiveVote - only the chart-view
forwarding that existed solely to reach the deleted HUD is gone.
Compile: 0 errors, 0 warnings (stage).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
DIRECTION_INTELLIGENT and the drift verdict it fed were removed in
the step-3 demolition (8f21646); WarriorDirectionAllows() now
resolves purely from tradingdirection (LONG_ONLY/SHORT_ONLY/BOTH).
Two comments in the OOS-verdict certification path and the filtered-
overlay reconstruction still described the deleted mechanism -
found while auditing both paths for correctness. No behavior change.
Compile: 0 errors, 0 warnings (stage).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The capacity budget is stated in weights per INDEPENDENT observation
and divides by the mean label lifespan to get there. It never once
did: EstimatedInSampleBars() deflates via m_labelOverlap, but it is
only ever called from InitNeuralNetwork, where the label cache does
not exist yet (that same function sets m_labelCachePrebuilt = false
a few lines below), so MeanLifespan() returned its "nothing measured"
default of 1.0 at every call. Every fresh model was sized as though
its labels did not overlap - over-budgeting the first dense layer by
a factor of L, which is several rungs of a power-of-two ladder. The
"expect overfitting, reduce the feature set or pool instruments"
warning is the branch that should fire on H1 and structurally could
not.
Fixed at the source rather than by reordering the boot sequence (the
prebuild is chunked across Train() calls and cannot complete inside
init): MeasureSwingGeometry() walks the ZigZag ONCE at init and
answers both questions from it - the median leg gives the window,
and the leg series gives the mean label lifespan analytically.
SwingPivotDirectionLabel resolves bar i when the SECOND pivot after
it commits, so a bar d bars before pivot P waits d + (the leg
leaving P); summed over every bar of every leg that is exactly the
mean the label walk accumulates.
That also closes the coherence gap the swing target opened: the
window was measured with a private +/-12-bar fractal while the label
aimed at ZigZag(12,5,3) pivots, so it was sized against a leg
distribution the label never used. One pivot source now, the
label's.
Also:
- ResetWeights() re-derives the shape. It rebuilt from the members a
history-starved init had pinned and re-saved them - so the "let
history download, then reset from the panel" advice in both
fallback warnings did nothing at all.
- The CAPACITY line prints the measured lifespan beside the one the
topology was sized for, and warns when they differ by more than a
ladder rung. That is the check that makes the estimator falsifiable.
- Topology reads the view's symbol, not _Symbol (latent for pooling).
- Unmeasured geometry defaults to HISTORY_BARS_FALLBACK, never 1.0:
under-sizing is recoverable, over-sizing silently is not.
Compile: 0 errors, 0 warnings (stage).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The MI suite kept its one irreplaceable job - the label-alignment
lookahead scan, whose margin is priced by the headline permutation
null and whose validity is proven by the positive control. Everything
that judged or vetoed on top of that measurement is gone:
- m_dirEvidence deploy veto deleted from all four deploy sites. The
policy is that screens are priors, not gates; the family-wise
selection test on held-out precision is the deploy protection, and
a marginal per-bar MI test cannot veto a model that reads the
window jointly (the report itself said so on every print).
- Per-column CFeatureSelector deleted; BlockPermuteLabels (the null
engine ScoreMiSample depends on, ragged-tail fix intact) moves to
AutoTune.mqh as a free function.
- Feature-lag profile deleted, with its MI_LAG_* constants and
BuildMiSample's featureBarOffset; MiShiftPad no longer pads by
m_historyBars.
Compile: 0 errors, 0 warnings (stage).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Step 3 of the swing-pivot plan, whole-hog. The swing label is now the ONE
target and the era verdict is precision + recall per class against the
label's own base rate - no win rate, no break-even, no expectancy, no
geometry anywhere in training.
DELETED
- Expert/Excursion/ (4), Expert/BarrierHorizon/ (4), GeometrySweep,
FirstPassageLadder, Labeling/TripleBarrier.mqh (CLabelOverlap survives
in Labeling/LabelOverlap.mqh), 3 test EAs.
- TripleBarrierLabel + walk, fractal label, geometry derivation/scan/
adoption, exit-policy replay, excursion MI targets, the drift verdict
(DIRECTION_INTELLIGENT), the recall floor, balanced-accuracy telemetry,
the barrier defines, the .cfg geometry adopt (slots kept as zeros for
the positional layout), the derived-geometry live-order override.
- TRAINING_TARGET input/enum: direction models are always swing; META2
re-keys the meta head onto label agreement (descriptor loses its two
geometry slots).
REWORKED
- Labels.mqh (1795 -> ~370 lines): AdvanceSwingLabelState with
FINALITY-GATED CACHING - an unresolved bar (pivot pair uncommitted) is
never cached, so it can never freeze as a false Neutral; training,
calibration, OOS scoring and online learning all skip unresolved bars.
- SDeployVerdict: significance-only; SOosTally chance = larger
directional class share; pooled gate poolability = timeframe (record v2).
- Purge/embargo/declustering gaps: the measured mean label resolution
lag (LabelResolutionBars), not a barrier horizon.
- Pool purge key + backfill DB rows: marked at the bar the label
resolved on (m_labelResolveAge), not a fabricated barrier touch.
- Online learning frontier: finality, not a horizon delay.
- m_bestBalancedOos -> m_bestSelectionScore, m_erasSinceBestBalanced ->
m_erasSinceBest, ensemble vote outcome arrays -> label arrays.
STEP 4 folded in: Entry_Multiplier / SL_Mode / TP_Mode / tradingdirection
are inputs again - trade management is the tester GA's search space.
Fingerprints: every direction model re-keys (TGT:SWG1 now unconditional,
CUT token gone); META1 -> META2. Full retrain, as planned.
Compile-verified in _claude_stage: Warrior_EA + both surviving test EAs,
0 errors, 0 warnings each.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- TrainingTarget defaults to TARGET_SWING.
- LogitAdjustTau input, preset enum and all plumbing deleted: tau is fixed
at 1.0 (the full log-prior, Menon et al.'s consistent value); the
delivered strength is capped to the head's usable logit range from the
priors the prebuild measures. The CAPPED journal line is the step-1
measurement. |LA💯BS becomes a frozen legacy fingerprint slot, so no
existing model re-keys.
- The swing label measures its own resolution lag (idx - P2, the earliest
bar P1 can be final on) into the overlap/SE machinery, capped at
SWING_SCAN_CAP_BARS instead of a barrier horizon it does not have.
- The prebuild line is target-aware: both-won, timeout and horizon-lifespan
fragments are barrier-walk facts and no longer decorate swing counts.
Compile-verified in _claude_stage: 0 errors, 0 warnings.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Operator's observation, verified against Examples/ZigZag.mq5's selection loop:
the only erasures it performs are ZigZagBuffer[last_high_pos] while hunting a
bottom and ZigZagBuffer[last_low_pos] while hunting a peak. A pivot therefore
leaves the erasable slot permanently the moment the OPPOSITE pivot is committed,
and can never move again - the opposite pivot does not itself need to be final.
SwingPivotDirectionLabel now waits for that event instead of for
m_swingConfirmationBars. The bar aims at P1, so it becomes trainable once P2
exists; pivots alternate by construction, so P2 is the next non-zero bar and
needs no type test. Until then the label is not knowable and the bar is Neutral.
Exact rather than a guess, and it removes the need to measure a repaint-lag
distribution at all. SwingConfirmationBars keeps its other uses; it is no longer
this target's lookahead control.
Compile-verified in the staging copy: 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Series indices are relative to now, so one new bar moves every cached bar's
index by one. EnsureBarCachesCapacity answered that by wiping the label cache,
the excursion caches, the ladder and the feature cache and rebuilding the whole
prebuild from scratch - on any timeframe where a bar closes before a run
finishes, the labels were being recomputed continuously and the training set
never held still.
The labels do not change when a candle closes. ShiftBarCaches moves every
per-bar cache up by the number of new bars, marks only those newest bars as
unfilled, and leaves the rest exactly as computed. CFirstPassageLadder gets a
matching Shift (resizing directly rather than through Allocate, which zeroes the
ages this is preserving).
Refuses, falling back to the full rebuild, when a prebuild is mid-flight: its
cursor is an index into the array being moved.
Compile-verified in the staging copy: 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
TARGET_SWING: the direction models learn which way the next CONFIRMED SWING
PIVOT lies from the current close. Geometry-free - the label owes nothing to a
stop, target or horizon - which is what lets trade management be tuned
separately instead of being baked into what the net learns.
SwingPivotDirectionLabel reuses the ZigZag pivot the horizon and leg-size
measurement already walk, so there is ONE notion of "pivot" in the codebase. It
walks forward in time and stops at m_swingConfirmationBars: a pivot nearer than
that is still repainting, so its label is not knowable yet and the bar stays
Neutral. That boundary is the whole lookahead control for this target.
TrainingTarget input is back (TARGET_BARRIER default, unchanged behaviour) with
TARGET_FRACTAL and TARGET_SWING beside it; |TGT:SWG1 joins the fingerprint so
switching trains a separate model rather than relabelling an existing one.
ADZigZag was renamed to ZigZag throughout (30 identifiers). It has loaded
MetaTrader's stock Examples\ZigZag at its stock defaults for some time - the
migration was done, only the name was left behind, and a name that says "AD"
about a stock indicator is exactly the legacy pointer this codebase should not
carry. No behaviour change: same #resource, same params.
Compile-verified in the staging copy: 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
First 35 eras across both charts, this run:
shipped 1.21/2.43 (SP500) and 1.26/2.52 (USDJPY): mean -0.0525R,
positive in 6 of 35 eras
best plateau after the neighbourhood guard: mean +0.0292R,
positive in only 17 of 35
most-recommended pair: 20.00/0.50, seven times - a ~40:1 lottery that is
simply the least negative cell in an all-negative grid
The recommendation jumps between opposite corners of the ladder between
consecutive eras, which is a grid fitting noise rather than a geometry worth
adopting. Two changes so the line cannot be misread:
- GEOSWEEP_MAX_TIMEOUT_SHARE (0.70): a cell where most trades never touch
EITHER barrier is not a geometry being tested, it is the horizon close being
measured. 20.00/20.00 timed out on 100% of trades and was still selected.
Excluded from SELECTION only; the cell stays filled and readable.
- When the winning plateau is <= 0 the line now says so in those words:
"NOTHING ON THE LADDER PAYS ... the pair below is the LEAST NEGATIVE cell,
not an edge."
Still measurement only - nothing reads the recommendation and no geometry moves.
Compile-verified in the staging copy: 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Step 1 of decoupling SL/TP from training. The geometry is currently chosen
BEFORE the model exists - excursions -> stop at a quantile -> target at the
policy minimum ratio -> labels -> the net learns those labels - so it has never
been asked which pair maximises expectancy GIVEN WHAT THE MODEL CAN PREDICT.
The scan meant to answer that reports "0 ELIGIBLE candidates" on this config
(every rung disqualified by the close-all clamp), so nothing has ever compared
the shipped pair to an alternative.
This needs no retrain and no backtest. CFirstPassageLadder already stores the
first-touch AGE of every rung on both sides and OutcomeR() resolves ANY pair
exactly with the spread charged the way the fill charges it - so 14x14 pairs
over one era's OOS calls is a few thousand array reads.
- Expert/Training/GeometrySweep.mqh: CGeometrySweep accumulates (n, sumR,
sumR^2, timeouts) per rung pair from the model's own directional OOS calls.
Reads no chart, holds no net, opens no file - exercisable against a
hand-built ladder, same doctrine as SDeployVerdict.
- Best() ranks on the 3x3 NEIGHBOURHOOD mean, not the cell itself. A 14x14 grid
read at its single highest cell is a best-of-196 maximum, biased upward by
construction - the same selection problem the deploy gate corrects across
eras. A pair whose neighbours also pay is a plateau; a lone spike is a lucky
run of trades and does not survive the next window. GEOSWEEP_MIN_TRADES (30)
keeps thin cells out of the selection entirely.
- Wired into pass 3 where the call and the bar index are both in hand, reset per
era, reported at pass-3 completion beside ReportCandidateGeometry. ONE line,
and only when the recommendation CHANGES - it prints the shipped pair's
expectancy and the best pair's on the SAME trades, so "better" is a difference
rather than two numbers from two populations.
Measurement only: nothing reads the recommendation yet and no geometry moves.
Compile-verified in the staging copy: 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
selectionScore used to be a win rate in percentage points and printed at one
decimal everywhere. Under DeployOnExpectancy it is expected value in R, so
"%.1f" rendered every real score as "0.0" - era 2's +0.05R and a genuine zero
looked identical, which makes the journal useless for watching the ranking the
plateau ladder is doing.
One formatter, DeployScoreText(), next to the score it formats: "%.3fR" under
expectancy, "%.1f%%" under significance. Routed all nine print sites through it
(ensemble era line, best-so-far, panel, regression, new-best, era-cap prompts,
the convergence line, the deploy dialog) and dropped the "%" suffixes they had
hardcoded. No new prints, no new log lines.
Compile-verified in the staging copy: 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
TWO CHANGES, both of which turn a permanent "nothing happens" into a decision.
1. THE DEPLOY GATE ASKS THE WRONG QUESTION. tradeable required the win rate to
clear chance by EDGE_MIN_SIGMAS - "can I PROVE an edge exists" from one OOS
window. On H4 that asks ~66% against a market supplying ~53%, so it is
unreachable by construction and no run has ever deployed through it.
SDeployVerdict now also carries the economics of the geometry actually being
traded - cost-adjusted break-even and reward:risk, both from the new
CostAdjustedGeometry() so a spread convention cannot be applied to one and
missed on the other - and derives
E[R] = (p - p*) * (1 + RR)
which is exactly zero at break-even by construction, so "profitable" and
"beats break-even" can never disagree. Under DeployOnExpectancy (new input,
default ON) tradeable becomes E[R] > 0 and selectionScore ranks eras by
expectancy instead of precision. Coverage and both-sides-live still gate
both: an expectancy over a handful of one-sided calls is not tradeable.
The struct also publishes scoreSE - the SE of selectionScore IN THE SCORE'S
OWN UNITS - because the score changes units with the objective (win-rate
points vs R). Both plateau bands now read it instead of precSE, which was
right for one objective and dimensionally wrong for the other.
Setting DeployOnExpectancy=false restores the previous behaviour exactly.
2. THE FILTERED VIEW COULD NOT DRAW WHILE ANY MODEL WAS TRAINING.
HistoricalNetVote built its divisor from VoteCapableWeight(), which answers
"may this member move real money" and returns 0.0 for an AI member until the
whole run converges. So the reconstruction's divisor was zero on EVERY bar,
every bar was skipped as "nobody looked", and the chart drew nothing at all -
for the entire training run, which before the plateau noise band was forever.
Reported as "no signals drawn since the refactor".
New ReconstructionWeight(): the same weight WITHOUT the converged-run
requirement, overridden on the AI member to ModuleWeight() gated on
SelfRanked() only. The overlay is a picture of what the vote WOULD have
shown, which a mid-training model can answer - the chart HUD already says so
with its "(trn)" marker. Live Direction() still uses VoteCapableWeight(), so
no untrained model gains a say in an order.
Compile-verified in the staging copy: 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
TWO INDEPENDENT BLOCKERS, both of which make the EA look like it is working.
1. THE LADDER NEVER ADVANCES. isBetter/isBetterEra compared selectionScore with
a bare `>`. selectionScore is a win rate over a few hundred independent
calls, so it moves several points era to era on noise alone - measured on
SP500 H4 today: 32.8 / 32.2 / 31.6 / 29.6 / 31.4 across consecutive eras, a
~3-point spread with no trend. Any upward blip was recorded as a new best,
which reset BOTH the plateau counter and the stage, which re-armed a x5
learning-rate warm restart, which injected fresh noise and produced the next
blip. The search sustained itself on its own variance and never reached
PLATEAU_STAGE_DEPLOY - the reported "thousands of eras without converging".
A new best now has to clear the incumbent by PLATEAU_NEW_BEST_SIGMAS (2.0)
times precSE, which the deploy gate already computes. 2.0 rather than 1.0
because incumbent and challenger are both noisy, so the SE of the difference
is ~sqrt(2) x SE, and a 1-SE band was already measured too narrow in a
noise-dominated search. Applied at BOTH ranking sites - the ensemble's and
the solo member's - which are documented as the same ordering. The first
scoring era still checkpoints unconditionally.
2. THE BLANK-CHART CENSUS WAS LYING. It printed "No member has a completed era
yet (snapshots fill at each member's first pass-3 completion)" while the
members were on era 23, because it inferred the cause from m_overlayVotedBars
alone - and that counter requires BOTH a non-zero divisor AND a non-zero net.
Three different states collapsed into one sentence. Split out
m_overlayHadDataBars (divisor non-zero) so the line names which it is:
hadData == 0 -> nobody published a snapshot: publication/index
hadData > 0, voted == 0 -> members looked and abstained: calibration
voted > 0, drawn == 0 -> the vote never cleared the threshold
Diagnostic only. It does not fix the missing arrows - it identifies which of
the three is happening, which the current line actively obscures.
Compile-verified in the staging copy: 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Every era was a ~1,200-bar chunk of a 16,264-bar window, and the oldest 90% of
the history was never reached.
All four passes yield mid-chunk on the 120ms budget: each one calls
StashEraResume (the single writer of m_eraResumePending) and returns. Those
used to be returns from Train() itself. When the passes were extracted into
their own methods (08c2cec) they became returns from a void helper, and Train()
carried straight on - reporting pass 1 "done" after one budget, running pass 2
over the sliver pass 1 had queued so far, scoring an OOS slice of it, and
letting AdvanceEra count an era. The extraction moved one side of the binding
and left the reader behind.
Measured on SP500 H4 (VerboseMode, 2026-08-24 15:05-15:14):
era 0 TRAINING WINDOW = 16264 bars ... Bars(series) = 16264 <- window fine
era 1277 pass 1 done in 0s - 1144 of 1193 bars usable <- sweep is not
era 1296 pass 1 done in 0s - 3117 of 3166 bars usable
era 1318 pass 1 done in 0s - 1391 of 1440 bars usable
~1,400 eras in ten minutes, the count varying with how many bars a 120ms budget
happened to buy. Downstream: each member held a different tiny OOS slice, so
the combined vote's shared-bar intersection collapsed ("0 shared OOS bars" on
nearly every era, score 0.0), and the plateau ladder counted 46 ungraded eras
as a plateau and fired a boosted warm restart on all four models.
Train() now returns whenever m_eraResumePending is set - after pass 1 (before
ReportPass1Outcome, which has no verdict to give on a yielded sweep), pass 2,
the calibration walk and pass 3. m_modelEta is already saved inside
StashEraResume, so the early returns keep the learning-rate trajectory.
The resume machinery itself was correct and is unchanged: BeginEra's resume arm
restores the cursor, m_passWindowOk/m_passWindowFail accumulate across chunks,
and the m_isPass2Active/m_isPass2Done guard already routes a resumed call to
the right pass.
Expect era numbers to advance slowly now. That is the fix, not a new stall.
Compile-verified in the staging copy: 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
With VerboseMode on, pass 1 reported eras of 422 / 949 / 1358 / 2562 bars on
SP500 H4 - four models, same chart, same second - against a series holding
~16,264 bars, and the number moved every era (CONV: 2562, 3671, 3405, 3532,
2830). Nothing in the journal said so. ReportDetectability and the CAPACITY
line both quote EstimatedInSampleBars, which is derived from the configuration
and not from the era, so they kept reporting "11385 in-sample rows / OOS window
4874 bars" for a window that was a tenth of that.
era.bars is MathMin(Bars(symbol, PERIOD_CURRENT, dtStudied, now) + historyBars,
Bars(symbol, PERIOD_CURRENT)). A short era is therefore either a dtStudied that
is too recent or a short price series, and those need opposite fixes - so the
new line carries all three quantities plus the resolved dtStudied and
SERIES_FIRSTDATE, not just the result.
Reported on change only: an era over a warm feature cache runs in a fraction of
a second here, and a per-era line would bury the journal.
Diagnostic only - no training behaviour is changed by this commit.
Compile-verified in the staging copy: 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Two charts (SP500 H4 + USDJPY H4) ran with the pool enabled and produced no
TrainPool directory, no adopted rows and not one journal line. The pool was
inert and there was no way to tell that from "the feature is off".
It could never have fired: the fingerprint is not symbol-invariant. It hashes
NeuronsCount, which counts the alt-data columns - and those are per-symbol
(SP500 carries cot_spec_net, the FX majors cot_idx_1y/3y/chg_4w) - and the
cross-asset block appends ":IDX2" when base currency == profit currency, true
of an index and false of a pair. SP500 came out 50 features wide under
XA:6:IDX2, USDJPY 52 wide under XA:6. Compatible() gates on both, so adoption
was zero by construction.
- STrainPoolHeader::MismatchReason() replaces the bare Compatible() predicate
and names the mismatch; Compatible() now delegates to it, so "may I adopt"
and "why not" can never drift apart.
- CTrainPoolReader::Adopt() reports its own verdict - adopted, alone, or every
peer rejected with the reason per file - and reports it on CHANGE only. An
era over a warm feature cache runs in a fraction of a second here, so a
per-era line would bury the journal. The duplicate Print in RunPass2 is gone;
pool state is now reported from exactly one place.
- CTrainPoolWriter::Publish() rate-limits to TRAINPOOL_MIN_PUBLISH_SEC (300s).
Every era re-derives the same rows from the same in-sample span, so per-era
publishing rewrote a multi-megabyte file continuously for no new information.
The first publish is never delayed.
Compile-verified in the staging copy: 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Peer rows join m_isTrainQueue as NEGATIVE sentinels before the shuffle, so they interleave with
this chart's samples instead of training in a block at one end. A block would be a curriculum:
whatever the optimizer saw last would decide where it landed.
TrainPoolStep is a separate path on purpose. Everything in pass 2's local branch after the
forward pass reaches for something indexed by a LOCAL bar - m_labelCache, m_winLongCache, the
excursion target, the arrow cache, m_Time - and a peer row has none of those. Sharing the path
would mean inventing values for all of them, which is how another instrument's outcomes end up
inside m_cumIsCorrect and the operating point gets fitted to them. The IS-vs-OOS gap is read as
THE overfitting signal, so polluting the IS side would not crash anything; it would just quietly
stop meaning what it says.
The purge key reuses the label walk's own two bounds - the horizon and NextScheduledCloseAll -
rather than approximating with a bar offset. A second horizon model here would drift from the
real one, and this project already measured that the close-all, not the nominal horizon, is what
actually terminates labels. Cutoff is the OLDEST OOS BAR'S TIME, in wall clock, because bar
indices cannot be compared across instruments that each have their own calendar.
Contribution happens while the window is still in TempData and before the forward pass
overwrites it, and is gated to direction models: the meta head trains a different target on a
wider input, which the fingerprint gate alone would NOT catch, since a meta model's fingerprint
matches its own peers perfectly well.
Use_Training_Pool ships false and does nothing until a second chart runs a matching fingerprint.
Compile-verified against a BASELINE of the same tree without the wiring: both produce 12
errors, all error 313 invalid-resource-path from #resource directives that cannot resolve in a
headless staged build (stock Controls res\*.bmp, plus the pre-existing Network.cl). Code errors
0, warnings 0, identical to baseline. Staging copy and junctions removed; the live .ex5 was
never touched.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
PooledGate pools the DECISION; this pools the DATA. Measured in research/edge.py with both arms
sharing calendar folds, exit-time purge, benchmark and scoring so training breadth is the only
variable: H4 k=2 gap +2.02pp at t_mkt 3.97, which CLEARS the Sidak bar of 3.69 over five feature
sets at df=6, replicated independently at D1 k=1 (+2.03pp, t_mkt 2.79). The per-instrument arm
was NEGATIVE on every feature set at both timeframes - it loses to "always take the drift side".
This EA trains one net per chart, which is that arm.
Rows, not symbols. Pointing the feature stack at another symbol needs per-symbol indicator
handles and this project has been bitten there twice - the handle leak that never released the
old handle, and the twelve "dead" handles that were one shared refcounted iMA. Each chart
instead computes its own features with its own handles and shares the NUMBERS. Sound only
because FeatureBuilder already ATR-normalises every price-unit feature, for exactly this reason
("instead of feeding e.g. 0.0005 on EURUSD").
Not a fingerprint participant: pooling changes what the model is trained ON, not what it IS, so
adding it would re-key every .nnw to record something outside the model's identity. The
fingerprint instead GATES adoption - it is the assertion that column k means the same thing in
both files - alongside a width check (a fingerprint match with a width mismatch means one side
pinned an older layout) and an exit-TIME purge, since a bar index cannot be compared across
instruments that each have their own calendar.
Writer and reader are separate classes: different reasons to change, different lifecycles, and
one class would carry the export buffers through every read. The file layout lives in one
STrainPoolHeader used by both sides so a layout change cannot be applied to the writer and
missed in the reader. Staging goes through System\AtomicFile rather than a second hand-rolled
temp-and-rename.
Compile-verified in isolation: 0 errors, 0 warnings. Staging junctions and harness removed; the
deployed .ex5 was never touched.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The verdict recorded in this docstring - that volume, time and alt all land at +0.019-0.020,
identical to price alone, so none was being used - came from a run whose weekend clock was a
day early and keyed on a bar most feeds never trade. That left two feeds with 74-78% unresolved
trades and sd(R) of 0.21 against everyone else's 0.62, which handed them overwhelming weight in
the inverse-variance pooling.
With the clock fixed the ordering inverts. `geom` becomes the WORST row rather than the
equal-best one, and price+time nearly doubles it:
price+time +0.047 | price+vol +0.038 | price +0.037 | ALL +0.028 | price+alt +0.027 |
geom +0.026
So price and time DO add ranking power over the strategy's own entry arithmetic. What survives
both versions is the alt result: every set containing alt columns scores below the same set
without them, agreeing with the direction screens at H4 and D1.
`shuffle` permutes the training outcomes while leaving fold boundaries, purge, threshold rule,
kept fraction and scoring identical. A lift that survives that comes from the machinery, not
the data - and this session has already produced two results that did exactly that, so the
+0.047 does not get believed until this run comes back near zero.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The encoder was rewritten once to convert every value rather than the one field known to hold
an array. It still only looked one level deep, and `evaluate` attaches the ENTIRE built matrix
under a 'd' key - so it stepped past d['X'] and killed the last line of a second 40-minute run
with the same TypeError the first fix was meant to end.
plain() now recurses into dicts and lists. The matrix and its column index are dropped rather
than converted: they are working state, tens of MB per instrument, reconstructible from
features.build, and nothing re-analysing these rows needs them.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The portfolio case for this strategy is that more uncorrelated strategies alongside it beat
an index. That only holds if the members do not share a hidden common factor, and a backtest
correlation matrix cannot see one: it is dominated by the calm months that make up most of a
sample, while the shared exposure surfaces in the month that breaches a drawdown limit.
So measure it directly. Aggregate each strategy's R by month, correlate against the
underlying's own monthly return, and split into up-months and down-months where a long bias
actually shows.
The answer is not marginal: mean correlation +0.69, positive on 14/14 feeds across crypto,
indices, energy, FX and metals, R2 up to 0.62 on USDJPY. Mean monthly R is +1.94 in up months
against -1.61 in down months. Every vendor pair agrees to within 0.03. These are not seven
independent bets, they are one bet placed seven times.
The placebo decomposition says why: across eight markets the barrier term is close to the
negative of the drift term (BTCUSD +0.055/-0.041, SP500_d +0.083/-0.030, USDCAD_d
-0.056/+0.059). The stop and target are a trend-capping device - they clip the gain where the
asset rises and limit the loss where it falls - so the residual cannot be an edge. It is the
same exposure with both tails trimmed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The two-arm version answered only half the question. It showed the MACD entry beaten by random
timing on every feed, but left the residual +0.01 to +0.05 R unexplained, and the first story
built on it - that the strategy is a drift harvester - died on the second market.
The drift arm settles it by running the same random entries with the stop and target moved out
of reach, so every trade holds to the weekly close. Its R is then rebuilt by hand on the REAL
stop distance, because simulate divided by the widened one; leaving that alone would report
every drift trade as ~0 R and make the comparison vacuous. real-timing is what the signal is
worth, timing-drift is what the barriers are worth over just holding.
The driver moves into this module as a `placebo` subcommand instead of living as a loose script
beside it, and reports per-market statistics for both terms.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
It was reported in every table as the drift this strategy inherits, and as the column that
isolates what the RULES contribute. It is neither. buy_and_hold_R marks each trade at
exit_idx - the bar its own barrier fired on - so a trade that exits at its target is compared
against the close of the bar that touched the target. The two are nearly the same number by
construction.
Measured, the per-trade difference has sd 0.10 against R's own 0.62. That is where a
per-market t of 8.23 and p=0.00017 came from: a quantity that mostly cannot vary will always
look significant. Every "skill over buy-and-hold" figure quoted from this module is withdrawn.
What the column legitimately shows is exit slippage, and it now says so.
placebo() replaces it. Same number of entries, same previous-day-low level, same ATR-scaled
stop and target read at the entry bar, same weekly close, same non-overlap, same fill engine -
only WHEN the orders are placed moves, drawn from the bars the rule could have fired on so the
null inherits the same calendar exposure. Geometry then appears in both arms and cancels, and
only the MACD timing is on trial, which is the question that was being asked all along.
The run also now prints whether any vendor PAIR disagrees in sign. Before the weekend-clock
fix SP500_d and SP500_5 disagreed (-0.003 against +0.064) and so did the two FTSE feeds; they
are the same market at correlation >= 0.999986, so that disagreement was evidence of a
machinery fault and nothing noticed it. Now it cannot pass unremarked.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Two bugs in one function, and together they invalidate every number this module has printed.
1. `(days + 4) % 7` makes Friday 5 and THURSDAY 4, so every `dow == 4` test matched Thursday.
Trades were force-closed a day early and the no-entry window blocked Thursday night to
Saturday night. Epoch day 0 is a Thursday, so Monday=0 needs +3. Verified against a known
calendar week instead of re-derived by argument.
2. The close was keyed on a literal 23:45 stamp. That assumes every feed trades up to it and
they do not - FTSE_d has 175 such bars in its entire history against XAUUSD_d's 13,657,
because an index CFD session closes hours earlier. Most FTSE trades found no close ahead of
them, and a trade past the last close bar got a negative horizon that maximum(_, 1) turned
into a ONE-BAR hold: a silent instant exit indistinguishable from an ordinary unresolved
trade. The close is now the last bar of the trading week, which is feed-agnostic and is
what 'flat for the weekend' means.
The tell was in the diagnostics, not the result: SP500_d and FTSE_d showed 74-78% unresolved,
sd(R) of 0.19-0.21 and 1.1-hour holds while every other feed sat near 0.62 and 20 hours. Those
two carried a third of all trades and, having almost no variance, dominated the inverse-
variance pooling - which is where metafilter's implausible t_mkt of 9 came from.
Corrected, the five feeds agree: 22-34% unresolved, sd(R) 0.61-0.64, RR 0.25-0.36, win 67-74%,
and expR still positive on all of them at 1bp/side. The finding survives; its statistics do not
and are being re-run.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
It was introduced as 'no market information whatsoever'. That was wrong. With g = fill minus
the previous day's low, rr = (6.51*ATR25 - g) / (g + 3.23*ATR15), a monotone decreasing
function of g/ATR - so rr is a NORMALISED DISTANCE ABOVE YESTERDAY'S LOW, a price feature in
the same family as donch and smadist, reparameterised until it looked like bookkeeping.
What survives the correction is the part that matters: volume, time and alt add nothing, every
combination lands at +0.019 to +0.020, and price alone already reaches +0.020. What changes is
the explanation - the lift is one price relationship, not an absence of one.
And the relationship is not monotone, so 'prefer a better payoff ratio' is the wrong summary.
By decile on SP500, XAUUSD and USDJPY alike it is an inverted U: filling far above the low
pays ~0, the middle band (rr 0.13-0.40) pays +0.05 to +0.14, and filling AT the low is
negative on all three. Buying the level the strategy aims at is the losing case.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
First run returned +0.020 R lift at t_mkt 6.5 - an order of magnitude beyond anything else in
this project - and the tell was in the same table: price, price+alt and ALL returned the SAME
lift to three decimals. A model given more information that does exactly as well as one given
less is not using the extra information, so whatever it found was in something all five sets
shared.
`geom` is that something: two columns of the strategy's own entry arithmetic, the realised
reward-to-risk ratio and risk as a fraction of price, both known at entry and carrying no
market information at all. It scores +0.022 - the LARGEST lift in the table - and ALL+geom at
+0.020 is no better. Every data family contributed nothing, which is exactly why they all
agreed.
The cause is the entry. A buy stop at the previous day's low sits below the market, so it
fills a median 4.75 ATR from the level its stop was sized against and the reward:risk of each
trade is close to arbitrary. The model was ranking that, not the market.
Also adds cut_from='train'. The keep-threshold was taken from the TEST fold's own prediction
quantile, which keeps exactly q by construction but cannot be known in advance - so the filter
as first measured was not implementable. The training quantile is fixed before the fold is
seen and lets the kept fraction float, which is the version that could be traded.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
edge.py asks a question that is mostly closed in this project: can a model call direction on
a symmetric barrier. Filtering is a different and easier question - the rules have already
chosen the side, and the model only has to rank trades that were going to be taken. A series
that cannot say 'up or down' can still say 'not today', and nothing built here so far could
have detected that.
For each trade the strategy takes, the feature vector is read ONE BAR BEFORE the decision bar
- every column in features.py is a function of bars <= i including i's own close, and the
order goes in at bar i's open, so reading row i would hand the filter the outcome of the bar
it is deciding on. A gradient-boosted regressor predicts R under a purged walk-forward, the
top q of each test fold is kept, and the lift is measured against the mean R of ALL trades in
those same folds.
That baseline is the one that cannot be gamed by the strategy being good: a random subset of
the same size has expected mean equal to the fold mean, so the difference is exactly what the
ranking contributed, and a profitable strategy raises both columns together rather than the
gap. Families are ablated as in edge.py and the verdict is read from the per-market t.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>