forked from animatedread/Warrior_EA
942 commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
af59b41820 |
fix(features): a Wyckoff sentinel was 1137x wider than every other column
FEATURE HEALTH now reports SCALE, and the first thing it found was severe:
median column range 2, widest/median 1137x
wyckoffEvent[4]=2274 (x1137) wyckoffEvent[5]=2265 (x1133)
ZoneTop and ZoneBottom ARE ATR-normalised - the defect is that the indicator reports "no active
zone" as 0, and `(0 - close) / atr` then evaluates to -close/atr. That ratio is scale-free, so it is
about -580 on EURUSD and -580 on NAS100 alike, injected into a vector where every other column lives
in +/-2.
WHY IT MATTERS MORE THAN A COSMETIC OUTLIER:
* It owns the COVARIANCE MATRIX. This is the entire explanation for the redundancy report claiming
"1 of 92 columns carry 95% of the variance (top component alone 100%)". The data is not
one-dimensional; one column is three orders of magnitude wider than the rest. Every PCA or
correlation-prune decision taken on that report would have been taken on an artefact.
* It owns the FIRST LAYER GRADIENTS. A column that much larger dominates every weight update, so
the network was learning mostly from a sentinel and drowning the other 90 columns.
Same defect class as CArrayDouble::At returning DBL_MAX for a negative index, found earlier today: an
out-of-band marker that arithmetic consumes without complaint. Guarded via WarriorZoneDistAtr, which
returns a NEUTRAL 0 for an absent level - the honest encoding of "there is no zone here" - and also
rejects EMPTY_VALUE and non-finite input, because an indicator with no data must not be read as a
price of zero. Width contract unchanged.
The SCALE half of FEATURE HEALTH is the durable part. The report has always named CONSTANT and
mostly-zero columns and never said how WIDE they are, so a scale outlier was invisible to it while
being the most damaging thing a column can be. It now prints the median range, the widest/median
ratio and the five widest columns by name, and says plainly that columns must be comparable in
magnitude before PCA or a correlation prune means anything.
Still outstanding from the same report: ichimoku[3,4,6] at 17-21x the median. Plausible for real
(close-cloud)/atr on small-ATR bars rather than a sentinel, so it is left measured and unpatched.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
d2574f4c3a |
feat(label): measure the horizon frontier - and it says the horizon is NOT the lever
The label horizon has been called "the only lever that raises evidence" for weeks and was never measured. This measures it, from PRICE and the leg replica alone - no network, no training, no era. What it computes per ZigZag depth: legs (which IS the effective sample size, because every bar inside a leg shares that leg outcome), mean leg life (= the label overlap), the ORACLE ride a perfect caller takes, the round-turn cost over the same window, the break-even capture, and the total at 2/5/10% capture. EURUSD, 60k bars: depth 4: legs 8994 life 6.7 oracle 2.055 BE-capture 1.62% @2%=+71 @5%=+625 @10%=+1549 depth 6: legs 6780 life 8.8 oracle 2.556 BE-capture 1.33% @2%=+116 @5%=+635 @10%=+1502 depth 8: legs 5330 life 11.3 oracle 3.046 BE-capture 1.13% @2%=+141 @5%=+628 @10%=+1440 depth 12: legs 3763 life 15.9 oracle 3.865 BE-capture 0.90% @2%=+160 @5%=+597 @10%=+1324 depth 16: legs 2887 life 20.8 oracle 4.581 BE-capture 0.75% @2%=+165 @5%=+562 @10%=+1223 depth 24: legs 1662 life 36.1 oracle 6.347 BE-capture 0.52% @2%=+156 @5%=+473 @10%=+1000 THREE RESULTS, TWO OF WHICH KILL MY OWN FRAMING: 1. Cost never binds. Break-even capture is 0.39-1.62% at every horizon, against rides of 2-8 ATR. The "shorter horizon sells payoff to buy evidence" trade-off this report was designed around barely exists. 2. At the ~2% capture this fleet has demonstrated, the horizon is nearly FLAT: +141 to +165 across depths 8-36, with the SHIPPED depth 12 already within 3% of the peak. Shortening to depth 6 makes it WORSE (+116). The horizon is not the lever. 3. Capture rate is first-order - 2% -> 5% roughly quadruples the total at any depth - and shorter horizons only win once capture is high (at 10%, depth 4 is worth 2x depth 16). The first cut of this report ranked by the ORACLE and therefore picked depth 4 on every chart, which is simply wrong: it credits a horizon with money no model here has ever taken. Ranking now uses the demonstrated capture, and the capture columns are the ones to read. Also fixed while here: LEG_STATE_DEPTH was NOT in the model fingerprint, though it sets where every pivot falls and therefore which leg each bar belongs to, its direction, its ride, and whether it is Buy/Sell/Neutral at all. Two models at depth 12 and depth 6 train on completely different targets and were sharing a key. Now TGT:LEG1:<minride>:D<depth>. The ZigZag replica takes depth/deviation/backstep as arguments defaulting to the shipped #defines, so every existing caller - including the in-situ verification against the stock indicator - is byte-identical. WHERE THIS POINTS: capacity is dominated by WIDTH, not horizon. Depth 12 gives 3763 independent legs; the mask-off change took inputs 180 -> 552, so the first dense layer is ~17,600 weights against 3,763 observations. Halving the horizon buys 1.8x observations; tripling the width already spent 3x. The fix is the redundancy filter - correlation between INPUTS, no label involved - which is designed in project_pca_reduction_plan and never built: measured worst pair |r|=0.999, PCA takes obs/param 0.49 -> 11.9. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
2a82b0bd9e |
fix(ichimoku): read the Senkou spans RAW - stop depending on the wrapper shifted buffers
CiIchimoku serves Tenkan and Kijun (buffers 0/1, no offset) perfectly and always has. The two SPAN buffers carry Offset(kijun), which makes CIndicatorBuffer::Refresh issue CopyBuffer(handle, num, -m_offset, m_size, ...) - a NEGATIVE start_pos - and THAT is the only path that has ever failed here. Six hypotheses died to measurement before this was accepted as the fix: stdlib cannot serve the buffers - false, a probe read them fine full-history count at negative start - false, returns 179048 err=0 a second Create corrupts the buffers - false, after 2 Creates spanA still ok primer/tuner parameter mismatch - false, tunedKijun=26 matched the primer the primer was missing - false, it ran and the spans were still EMPTY buffer starvation from re-init - live state was Total=10 spanA(0)=EMPTY The live object fails in a way a constructed one will not reproduce, and I could not pin it. So the dependency is removed instead of diagnosed further: the spans are copied ONCE PER ERA from the raw handle at start_pos 0 - measured reliable on all three charts, 179022 values, err=0 - into our own series-indexed arrays, and the kijun shift is applied by the READER. Same semantics the wrapper Offset(kijun) provided, now stated where it can be read instead of inside a base class. The raw handle is created with the SAME TUNED parameters as the wrapper, because a different triple is a different indicator instance - reading 9/26/52 does nothing for a chart whose tuner moved to 9/30/52 - and the old handle is released first so ReInitADIndicators cannot leak one per call. VERIFIED LIVE at the full 552-input width: 0 rejections, 0 stalls, all three charts training (EURUSD era 14 in 18s, USDJPY era 3 in 11s, NAS100 era 6 in 10s), combined vote scoring at era 13. Also: compile staging now excludes .venv/.git/docs/pdb/obj, which cuts the per-compile copy from 505 MB to 7.4 MB. It does NOT speed the compile - MetaEditor genuinely takes ~83s on this codebase - it just stops copying a 386 MB Python virtualenv into the terminal folder every time. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
ceef283ead |
feat(features): mask OFF - stop doing the network job with a univariate screen
The keep-screen scored each column MARGINAL mutual information against the label and dropped what did not clear Benjamini-Hochberg at q=0.10. That is univariate feature selection, and it structurally cannot see a column that is worthless alone and valuable in combination - which is most of the interesting ones. sin(hour) is independent of "will this leg continue"; a breakout AT THE LONDON OPEN is not the same event as a breakout at 03:00. Volume alone is noise; volume crossed with range is not. The header of FeatureMask.mqh already stated this blind spot - "the screen tests each column MARGINAL information, so it cannot see a column that is useless alone and useful in combination" - and the mask was built on it anyway, so every pass deleted precisely the features a network exists to combine. We were using an algorithm to do the model job, in a single pass, and then recording the result as "this data is worthless". FEATURE_MASK_VERSION 12 -> 0, which restores the full column set of every enabled block exactly, as that knob was designed to do. Width 6x36 -> 6x92 = 552 inputs. Re-enabled: time (hour/day/month, cyclic), volume, ATR, fracdiff, alt data, spread, and the four AD/Wyckoff blocks. Block selection is now a toggle decision made on data value and engineering cost, never on a per-column significance test. THE LEGITIMATE FILTER IS REDUNDANCY, NOT SIGNIFICANCE. Two columns carrying the same information is a correlation question between INPUTS and needs no label at all - worst measured pair is |r|=0.999. That screen stays and is the next thing to build; the supervised one is gone. ONE BLOCK REMAINS OFF AND NOT ON A VERDICT: cross-asset loads FOUR EXTRA SYMBOLS of full history per chart, which was the largest contributor to the 27.8 GB of committed address space that was crashing the machine. It returns when there is headroom to measure it in. MEASURED COST OF THIS CHANGE, stated plainly: commit 3.7 GB -> 33.3 GB, headroom 26 GB -> 5.1 GB. The AD/Wyckoff iCustom instances are that cost. Porting them to direct source inclusion is the fix that buys back both the memory and the throughput. Baseline to judge the new features against (pre-time, mask v11): USDJPY book +0.042/+0.067/+0.062 with alpha -0.077/-0.052/-0.058; EURUSD -0.070/-0.155. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
9550f94997 |
fix(ichimoku): the throwaway probe turned out to be the fix - it is now a deliberate primer
The operator was right that Ichimoku is a stdlib indicator like every other, and it is. Three hypotheses of mine died to measurement today: 1. "MT5 will not serve the Senkou buffers" - false; the wrapper reads them fine 2. "the full-history count at a negative start_pos fails" - false; returns 179048, err=0 3. "a second Create corrupts the buffers" - false; after 2 Creates, Total=10, spanA still ok What actually correlates with the fix is the one thing the probe started DOING rather than reporting. While it copied 64 values the feature path still failed every bar with spanA/spanB EMPTY and no chart completed an era. The moment it began issuing FULL-HISTORY CopyBuffer calls on buffers 2 and 3 at init, the feature path started working on all three charts - 0 rejections, 0 stalls, eras in 21-34s, combined vote scoring - with no other functional change between those two builds. The reading: the Senkou plots are shifted kijun bars FORWARD, and that shifted region needs one full-range materialisation before CIndicatorBuffer::Refresh - which asks at start_pos = -offset for m_size values - will serve it. Priming costs two CopyBuffer calls per init. Not priming cost the fleet hours across two days. CAUSALITY IS INFERRED FROM SEQUENCE, NOT PROVEN BY ISOLATION, and the header says so. It is renamed from WarriorProbeIchimokuBuffers to WarriorPrimeIchimokuBuffers and marked DO NOT DELETE AS SCAFFOLDING, because I was one turn away from removing it as spent diagnostics - which would have re-broken the fleet and left no trace of why. Buffer 3 is primed alongside buffer 2 even though nothing reports on it: the feature block reads SenkouSpanB every bar exactly as it reads SenkouSpanA. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
ddb9fa0897 |
fix(features): the Ichimoku failure branch is PER BAR - a CopyBuffer there stalled the fleet
I put a raw CopyBuffer plus a ten-argument StringFormat on the Ichimoku EMPTY-read branch to
capture the one number that would have prevented an earlier misdiagnosis. That branch runs once per
FAILING BAR, and transient rejections are entirely normal during a cold pass-1 scan, so it became a
terminal API call per bar.
Measured cost, on a fleet that had been healthy minutes before:
pass 1 throughput ~15,000 bars/s -> ~1 bar/s
EURUSD LSTM bar 1536 of 179042 after 1180s
USDJPY CONV bar 1024 of 178952 after 1277s
combined-vote eras scored in 20 minutes: ZERO
Every chart stopped completing eras and the ensemble never scored, so nothing could be certified.
The branch was invisible before this only because the feature had been disabled - re-enabling
Ichimoku is what exposed a cost that had been sitting on a path nobody was walking.
The raw-handle comparison it was meant to capture already lives in the one-shot init probe, which
answers the same question once per start instead of once per rejected bar.
THE RULE: a diagnostic belongs where the diagnosis is READ, not where the failure is DETECTED. A
rejection path in a per-bar loop is a hot path, and it does not stop being one because the feature
that walks it happens to be off today.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
b724fc58a2 |
fix(ichimoku): back ON - the stdlib was never the problem, the double-shift was
I switched Ichimoku off earlier today on the reasoning that "MT5 will not reliably serve the Senkou
buffers", from the symptom that spanA/spanB read EMPTY at every index while tenkan/kijun read fine
on the same bar. That conclusion was wrong. The operator pushed back - it is a stdlib indicator like
every other - and an init-time probe settled it by measuring the thing I had only inferred:
wrapper: BufferResize(39502)=ok BarsCalculated=39502
at idx 5 -> tenkan=ok kijun=ok spanA=ok spanB=ok
raw: spanA CopyBuffer(start=0)=64 err=0
spanA CopyBuffer(start=-26)=64 err=0 spanA[0]=29432.895
On all three charts. CIndicatorBuffer::Refresh reads shifted buffers with
`CopyBuffer(handle, num, -m_offset, m_size, m_data)`, and that NEGATIVE start_pos is deliberate and
works - it is how the forward-plotted cloud region is addressed. The wrapper is coherent and the
indicator serves data.
What was actually broken is the double-shift fixed in
|
||
|
|
58982c8651 |
fix(measure): a zero-cost instrument has no cost axis to split on
Indices carry no commission, so costPx is 0, every row costR is 0, the median is 0, and `costR <= median` is true for EVERY row - the whole population lands in the cheap half and the other reports n=0. Measured on NAS100: "CHEAP n=595, DEAR n=0", with the cost multiple printing 0.00x because a book divided by a zero cost is undefined, not infinitely profitable. Refusing to form the split is what makes the silence honest. A degenerate split that still prints two halves invites the reading that one half "wins", when one half is the entire sample. The test is the TOP of the distribution: if the dearest row still costs nothing, there is nothing to compare. Found because the fleet result came in and the NAS100 lines were unreadable next to EURUSD and USDJPY. The measurement itself has now closed its own question - the two charts that CAN be split disagree in sign - but the guard matters for every future cost-denominated statistic. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
cf62fc22a8 |
perf(features): masked out means switched off - 27.8 GB of commit becomes 3.3 GB
The keep-mask decides which EMITTED columns survive. It does NOT stop a block from RUNNING. Eleven
blocks that FEATURE_MASK_VERSION 10 discards were still loading history, creating indicator handles
and computing values on every bar so the result could be thrown away.
MEASURED, on a restart with nothing else changed:
terminal64 committed 27,817 MB -> 3,309 MB (-88%)
system commit headroom ~2,600 MB -> 26,491 MB (10x)
That 23.6 GB was being committed in the FIRST SIXTY SECONDS after launch, before a single era
completed - which is what finally localised it, after the training loop had been the obvious suspect
all morning. It is also the direct cause of the operator's VS Code dying overnight: Windows charges
COMMITTED pages against the commit limit whether or not they are ever touched, and at 94.6% of a
49.7 GB limit any process asking for a few hundred MB fails. The working set was never more than
2.1 GB, which is exactly why Task Manager looked healthy the whole time.
Switched off, each one named by the mask decision that already discarded it: volume (v6), time, atr,
fracdiff (v5), sot, wyckoffEvent, wyckoffFail, wyckoffBarInv, crossasset (v6), spread, alt (v5).
crossasset was the expensive one - it loaded FOUR EXTRA SYMBOLS of history per chart - and the four
Wyckoff blocks are iCustom, which is the per-bar cost.
THE EMITTED VECTOR DOES NOT CHANGE. FeatureBlockTable already sizes a disabled block to width 0 and
the mask already dropped these columns, so what reaches the network after masking is byte-identical.
Only the work disappears. Verified live: input 132 (6 bars x 22), zero feature rejections, zero
stalls, all three charts training.
const, NOT input, for the seven that were inputs. They move m_neuronsCount and are therefore
retrain-forcing, and this project's rule is that retrain-forcing values are not inputs - MT5 stores
inputs PER CHART, so an already-attached EA ignores a changed default and would have kept every one
of these pinned true. That is the trap that cost a deploy cycle in
|
||
|
|
0b2cf77606 |
fix(features): Ichimoku OFF - MT5 will not reliably serve the Senkou buffers
The two Senkou plots are shifted kijun bars into the FUTURE, and CopyBuffer from position 0 across
a forward-shifted plot does not dependably fill. Measured on NAS100: `spanA=EMPTY spanB=EMPTY` at
series index 5 (raw index 31) while `tenkan=ok kijun=ok` at the SAME bar, same handle, same
refresh, BarsCalculated=39500. Buffers 0 and 1 healthy; buffers 2 and 3 essentially unpopulated.
WHAT IT COST. With every cloud read failing, not one of 39494 scanned bars produced a usable
feature window, so every era was discarded and restarted from scratch, forever. It idled charts
for 66 minutes on 09-03 and for 18 more on a FRESHLY STARTED terminal on 09-04 - so it is not the
revive path and a restart does not cure it. The tell is the member split: PAI reached era 4 while
CONV/LSTM/HYB sat at era 0, because only the sequence models carry a 6-bar lookback window and one
dead column kills the whole window.
THE TOGGLE IS WHAT HAD TO MOVE, NOT THE MASK, and getting this wrong would have shipped a no-op.
The keep-mask decides which EMITTED columns are kept; CFeatureBuilder's Ichimoku block is gated by
UseIchimoku() (Warrior_EA.mq5:806) and would still have RUN and still returned false on an EMPTY
span, failing the whole window before the mask was ever consulted. Both are moved here: the flag
stops the block executing, and FEATURE_MASK_VERSION 9 -> 10 drops the columns and forces the
retrain.
AND THE EVIDENCE FOR KEEPING IT WAS NEVER VALID. The 7/8 and 3/8 screen votes v8 cites were
measured while the block DOUBLE-APPLIED the cloud shift (fixed hours earlier,
|
||
|
|
97e55c3879 |
fix(pool): the reader kept both defects the writer was repaired for
|
||
|
|
827b7c59f7 |
feat(measure): state the cost-split book in money, exactly, by letting the ATR cancel
costAtr is built as costPx / atrBar, where costPx = WarriorRoundTurnCostPriceByClass() is a
per-SYMBOL CONSTANT (commission) and only atrBar varies. The ride is measured in that same
atrBar. So the per-row ratio ride / costR has the ATR cancel outright:
sum(ride in money) / sum(cost in money)
= costPx * sum(ride/costR) / (costPx * n)
= mean(ride / costR)
which is exact, needs no per-row ATR array, and answers the most decision-relevant question this
gate can pose: HOW MANY TIMES ITS OWN COST DOES THE AVERAGE CALL EARN. Subtract one for the net
multiple.
It exists because the first cut of this report could be answered with "ATR is doing the work".
Comparing two ATR-normalised halves invites exactly that objection, since a quiet bar's +0.250 ATR
is a smaller cash move than a loud bar's. This statistic prices that difference instead of hiding
it, and it is not a normalisation of the money answer - it IS the money answer.
The same identity settles what the halves are. With costPx constant, costR varies ONLY through the
bar's ATR, so "cheap half" is EXACTLY "high-volatility half". That reading was an interpretation
when the split shipped; it is an identity now.
Guarded on costR > 0 rather than >= 0, because this term divides by it and a zero-cost row would
enter as an infinite multiple rather than as a free trade.
Report-only: no feature, width, or fingerprint change, so it forces no retrain.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
5cd93bfb7b |
fix(ichimoku): the stdlib already applies the cloud shift - we were applying it twice
CiIchimoku::Initialize sets Offset(kijun_sen) on the two Senkou buffers (Include\Indicators\ Trend.mqh:678-680), and CIndicatorBuffer::At(i) returns CDoubleBuffer::At(i + m_offset). So SenkouSpanA(idx) ALREADY means "the cloud as plotted at bar idx" - raw buffer index idx+kijun holds (Tenkan+Kijun)/2 for a bar kijun back, which is exactly the cloud edge visible at idx. The feature block passed idx + kijunShift on top of that, reading raw index idx + 2*kijun. Every cloud column was therefore lagged one full kijun past what its own comment claimed, and the oldest 2*kijun bars fell off the end of a buffer sized to barIndex + kijun. Not a lookahead - the values were merely stale - but the block was not measuring the quantity it named, which means the keep-screen's Ichimoku vote (7/8 and 3/8) was cast on a feature nobody had specified. The projected cloud is now computed from its definition rather than read from a buffer. Senkou Span A is (Tenkan + Kijun) / 2 by construction, so the span that will be drawn kijun bars ahead is available from the two lines already read AT idx - no buffer access, no lookahead. That matters because the buffer read it replaces would have been SenkouSpanA(idx - kijunShift), and a NEGATIVE index is the trap on the other side of this one: only CDoubleBuffer::At guards index >= m_data_total and returns EMPTY_VALUE, while a negative index falls through to CArrayDouble::At, which returns DBL_MAX - straight past every == EMPTY_VALUE guard and into a feature. Span B has no closed form from the values in hand, so the pair-thickness "twist" is not reconstructed; the near edge's forward slope carries the same regime information at the same width, and the block stays 8 columns wide. FEATURE_MASK_VERSION 8 -> 9 with the kept set UNCHANGED. The bump is the point: the fingerprint is what stops a resumed model from loading weights fitted to the old semantics, and column count alone cannot tell the two masks apart - the exact case that knob's own comment was written for. NOT the cause of today's fleet stall, and the log rules it out: USDJPY ran 291 eras today on this same code. The stall was Ichimoku buffers never becoming readable on the two charts that had been REVIVED via ChartApplyTemplate, and a clean terminal restart cleared it. Two separate problems that happened to name the same indicator. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
c7ed4df786 |
feat(measure): split the book by what the bar COSTS - the one asymmetry direction cannot close
Commission is fixed in PRICE. The book is measured in ATR. So cost-in-ATR is inversely proportional to the bar's own volatility, and the same contractual fee is a smaller fraction of the move on a high-sigma bar. If the ATR-normalised edge is roughly sigma-invariant, then restricting the book to high-sigma bars raises the edge-to-cost ratio WITHOUT predicting direction any better than we do now. That matters because of what this project has already measured. Excursion SIZE is the one quantity that is genuinely predictable here - RANGE clears at ~4x its null on three instruments with a working positive control, and the LSTM member scores +8.6% to +11.5% disjoint Brier skill at 3.46-4.58 sigma against a trailing-quantile incumbent. DIRECTION is closed (project_direction_closed_verdict, best-of-999 p=1.0000). A lever that monetises predictable sigma is therefore the only lever our own evidence supports, and this is the cheapest possible test of it: a split of rows we already collect, costing one pre-pass and six arrays. The rows are cut at the MEDIAN measured per-bar cost, not at a fixed ATR threshold - the cost distribution spans three orders of magnitude across the asset classes this fleet trades, so any absolute cut would put one instrument entirely on one side. Each half carries its own cost into its own accumulator and is judged NET OF THE COST IT WOULD ACTUALLY HAVE PAID; comparing both halves against one pooled cost is what would make the cheap half look profitable for free. Unpriced rows (-1) are excluded from the median and from both halves, on the same reasoning as the sweep's existing >= 0 guard: binning an unpriced bar as cheap is the error that guard exists to prevent. MEASURED AND NOT GATED, on the same discipline the split-half report was held to. It prints at the certified rung on every era. A positive cheap-minus-dear gap sustained across the fleet makes a sigma filter worth building; a gap that is noise around zero closes the question. Replacing an unreachable bar with an unmeasured one is exactly how the precision gate closed, and this does not repeat that. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
08591e9c4e |
fix(fleet): three charts, because the terminal was committing 29.5 GB of address space
WHY VS CODE KEPT CRASHING, and why the terminal died at 03:51 - one cause: terminal64 working set 2,222 MB COMMIT 29,509 MB Against a commit limit of ~47 GB (32 GB RAM + a 15.4 GB pagefile) on a box also running six 1 GB tester agents and fourteen VS Code processes. At roughly 76% committed, every new allocation starts failing. That is the bad_alloc that killed the terminal inside WarriorCPU.dll ( |
||
|
|
70a3c9a719 |
fix(dll+mask): the overnight crash was an unhandled bad_alloc, and the screen verdict is in
THE CRASH. terminal64.exe died at 03:51 and took the night's training with it:
Faulting module name: WarriorCPU.dll
Exception code: 0xc0000409 (__fastfail)
Fault offset: 0x00000000000182ed
NOT a system OOM - 32 GB total, 25.4 GB free afterwards - and not a stack overrun in the
kernels, which are all REQUIRE-checked against real buffer sizes. 0xc0000409 is what the
CRT raises via __fastfail when C++ calls std::terminate, and the cause here is an
allocation that threw std::bad_alloc straight across a __stdcall DLL boundary into MT5's C
code, which cannot unwind it.
CPU_BufferCreate had no exception handling at all: AllocBufferSlot may push_back onto the
slot vector and assign() reserves the whole buffer, and either can throw. One try/catch in
the entire DLL against five throwing allocation sites, no nothrow, no new-handler. So an
allocation failure - a NORMAL outcome when forty members each hold a multi-hundred-MB
adopted pool - killed the terminal instead of returning an error.
It now returns -1 with CPU_ERR_ALLOC, which is the value the function already returns for
bad arguments, so every existing caller handles it. First such crash in 14 days of event
log, and it landed the night the feature width went 288 -> 360; the fix makes that class of
failure survivable rather than fatal.
THE SCREEN VERDICT (mask v8). The measurement pass DID complete before the crash - ten
keep-screen votes banked. Width 360 -> 180.
KEPT rsi 1/1 on both charts that voted at 60 columns, macd 3/3 and 2/3, ichimoku 7/8
and 3/8. Their August verdict was taken under the OLD label and does not survive
re-measurement. Two charts only - the weakest evidence in the mask, and the next
screen supersedes it.
KEPT cumdelta - 4 of 6 columns on ALL FIVE charts at 48 columns and 4/6 again on
NAS100 at 60. Six of seven chart-observations, as consistent as the ma block.
DROPPED sot (1 slot of 28), wyckoffEvent (0/16 on five charts, 0/16 and 1/16 on the two
at 60 - the largest block in the set and where every CONSTANT column appeared),
wyckoffFail and wyckoffBarInv (inconsistent across charts, which for a fleet-wide
mask is the same as absent).
The August verdict was right about the FAMILY and wrong about one member. Screen per
COLUMN, kill per column.
Halving the width also halves the pool footprint that provoked the allocation failure.
RETRAIN-FORCING: mask version and input width both ride the fingerprint.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
2679b02970 |
fix(features): the eight restored toggles are const - chart-stored inputs were pinning them false
The fleet ran a whole cycle at 48 columns while the source said 60, and every log line agreed with the source. MT5 STORES EA INPUTS PER CHART, and an already-attached EA ignores a changed DEFAULT entirely. The five AD toggles took their true default because they were genuinely NEW to the attached build. The three oscillators were introduced in the SAME commit ( |
||
|
|
0853f09dd9 |
fix(fleet): the expansion list named the four charts it ADDED, not the fleet
FLEET_EXPANSION_SYMBOLS was "BTCUSD,GBPUSD,NAS100,DAX40" - the four the list was created to open. Correct while opening was its only job, because the original six were already there. It is now also the list the REVIVE sweep walks, and a symbol missing from it cannot be revived. Measured tonight, on the very first sweep: BTCUSD, GBPUSD and NAS100 came back; SP500, XAUUSD and XTIUSD stayed dead, purely because they were not named. Same failure, same minute, opposite outcome, decided by an omission nobody would have looked at. The list now names all ten. It means "the fleet", which is what its own header already implied - "adding an instrument is a reviewable list in source control". Not retrain-forcing. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
c6d001d94f |
fix(fleet): a chart whose EA was killed is invisible to fleet expansion - it now revives it
Five of ten charts sat idle for over an hour tonight with no missing chart and no error to
see. Sequence: an unchecked ArrayResize overran (fixed in
|
||
|
|
a225dca7f6 |
feat(features): enable RSI, MACD and Ichimoku for the measurement pass
Completes
|
||
|
|
151f2bc1eb |
fix(pool+vote): unchecked ArrayResize wrote out of range once the feature width tripled
Nine live 'array out of range' errors within minutes of the width going 72 -> 288: TrainingPool.mqh (298,32) on five charts, ExpertSignalAIBase.mqh (3254,24) on a sixth. BOTH ARE THE SAME DEFECT AND NEITHER IS NEW - the wider vector only made them reachable. Each site calls ArrayResize and then writes at the index it asked for, without checking that the resize succeeded. ArrayResize returns -1 on failure; the write then lands past the end of an array that never grew. TrainingPool also asked for an absurd reserve. The hint was 16384 * m_width, i.e. sixteen thousand ROWS - 196k doubles at a width of 12, and 4.7 million (38 MB) at 288, per pool, on a 2013 Xeon running forty members. A reserve proportional to row width turns a wider feature vector into a quadratically larger allocation for nothing. It is now a flat 65536 ELEMENTS, and the result is checked: a pool that cannot grow refuses the row. The ensemble vote buffer checks the two arrays that actually bound the write - they are resized in one block, so any failure in it surfaces there - and drops the bar with one throttled line rather than corrupting memory. Losing a bar reads as low coverage; the alternative reads as anything at all. Found because the restored Wyckoff groups took the width to 288, which is the point of a measurement pass: it exercises the code at a size nothing had run at before. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
8866ff7c3b |
feat(features): restore RSI, MACD, Ichimoku and the five Wyckoff groups for re-measurement
Reverts
|
||
|
|
05401c9c64 |
fix(nn): linear output head + the missing chain factor - gradient check now agrees to 1.0000
Test_Backprop: SUMMARY 38 passed, 0 failed, 38 total.
LINEAR head (shipped): all 18 output-layer ratios 1.0000
THE DEFECT WAS THE HEAD, NOT THE BACKPROP. The classification head ran
z -> sigmoid -> a -> logit = CLASS_LOGIT_SCALE * a -> softmax, and set the output gradient
to (target - p). The softmax+cross-entropy cancellation that makes (target - p) correct
holds only when the softmax acts on z DIRECTLY. With a sigmoid and a temperature in the
way the chain keeps both factors, so the true gradient was
dL/da = CLASS_LOGIT_SCALE * sigmoid'(z) * (p - target)
and the code carried neither term. sigmoid'(z) = a(1-a) VARIES per neuron and per sample,
so Adam could not absorb it either - I claimed earlier that it could, and that was wrong.
MEASURED, IN THREE STEPS, EACH ONE PREDICTED BEFORE IT WAS RUN:
shipped SIGMOID head ratio 0.667-0.675, differing per neuron missing 6*a(1-a) = 1.5 at a=0.5
LINEAR head, no fix ratio 0.1667 flat missing 6 exactly
LINEAR head, fixed ratio 1.0000 on all 18 nothing missing
The middle row is what nails it: swapping the activation alone left a clean 1/6, proving
the two factors are separable and naming each.
THE FIX. Output activation SIGMOID -> NONE for every classification head, and the softmax
branches in BOTH backends multiply by CLASS_LOGIT_SCALE. With a linear head there is no
activation derivative to carry, so those two changes make the existing (target - p)
exactly right.
WHY KEEP THE TEMPERATURE AT ALL. A linear head could use scale 1 and drop the factor, but
ApplyLogitAdjustment expresses its cap as a fraction of CLASS_LOGIT_SCALE, so changing it
would silently retune the prior correction by 6x. Keeping the temperature and carrying it
in the gradient is the same maths with a smaller blast radius. The sigmoid was the only
reason the temperature had to exist in the first place - squashing z into (0,1) capped the
reachable softmax probability at 0.576, and the scale was there to win the range back.
THE TEST ALSO FOUND TWO BUGS IN ITSELF, both worth keeping in mind:
* connections point FORWARD (prevLayer.At(n).Connections.At(m_myIndex)), so reading them
off the output layer iterates an empty array. First run: 0 assertions, and TSummary
called it ALL TESTS PASSED. A vacuous suite now fails loudly - that false green would
have applied to every future test.
* CNet::backProp does not only compute gradients, it calls updateInputWeights and MOVES
the weights. Differencing around the moved point against a gradient measured at the
original one left a systematic 0.3-0.5% error that Richardson extrapolation could not
remove, because it was never truncation. Snapshot-and-restore around backProp took the
agreement from 1.004 to 1.0000.
Numeric reference is Richardson-extrapolated to O(delta^4); tolerance is the standard
relative-error criterion at 1e-4; the seed is fixed so the operating point does not move
between runs. The old SIGMOID head is kept as a case in the suite and asserts DISAGREEMENT
at ratio ~4.02 = 1/sigmoid'(z), so the test stays honest about what it is measuring and
turns red if either head drifts.
CNet gains one read-only accessor, Layers(), used by the test and nothing else.
RETRAIN-FORCING: every model's output activation changes. EnforceOutputActivation corrects
a loaded net in place, so existing weights self-repair rather than silently running the old
head.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
ec2d1ff113 |
test(backprop): the linear head gives ratio 1/6 exactly - diagnosis closed
The check now runs BOTH heads and the second one settles it. shipped SIGMOID head ratio 0.667 - 0.675, differing per output neuron proposed LINEAR head ratio 0.1667, 0.1697, 0.1669 - constant 0.1667 is 1/CLASS_LOGIT_SCALE to four figures. The residual third-decimal spread is float noise on a 1e-5 central difference, not structure. So the two heads separate the missing factor into its two parts: sigmoid contributes sigmoid'(z) = a(1-a) - VARIES per neuron and per sample temperature contributes CLASS_LOGIT_SCALE - a clean constant of 6 and the shipped head is missing their product, 6*a(1-a) = 1.5 at a=0.5, which is exactly the 1/0.6663 measured in the previous commit. THE FIX IS THE HEAD, NOT THE BACKPROP. With a LINEAR output and CLASS_LOGIT_SCALE = 1 the softmax acts on z directly, the textbook cancellation dL/dz = p - target holds, and the gradient this code ALREADY computes becomes exactly correct - no change to backProp at all. The defect was putting a sigmoid and a temperature between z and the softmax, which breaks the cancellation that makes softmax+cross-entropy convenient in the first place. ENUM_ACTIVATION already has NONE, so the head is expressible today. THE TEST ASSERTS DISAGREEMENT, NOT AGREEMENT, on both heads for now - a suite that went red for a defect it exists to document would just get muted. It turns red the day either head is corrected without updating it, which is the behaviour wanted from a regression test around a known bug. Still no production change. Correcting the head is retrain-forcing, touches OutputLayerActivation, EnforceOutputActivation and the persistence path, and wants its own commit. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
2202354c24 |
test(backprop): numeric gradient check - and it FAILS 18 of 18 on the output layer
The MQL5 backward pass trains the whole fleet and had no verification of any kind. DirectML/lstm_seq_gradcheck.cpp verifies the LSTM kernel to 2.3e-10; nothing verified CNet::backProp - and BOTH of this project's real maths bugs lived in exactly that path (the transpose bug, Adam's second moment). Both were found by reading. A third would not have been. Central differences at delta=1e-5 against the loss the code actually implements (cross-entropy on softmax of CLASS_LOGIT_SCALE * sigmoid output), comparing against gradient * prevNeuron.getOutputVal() - which is literally what updateInputWeights hands the optimizer. RESULT: 18 of 18 output-layer weight gradients disagree. out[0]<-prev[*] ratio 0.6996 (identical across all six weights) out[1]<-prev[*] ratio 0.7035 out[2]<-prev[*] ratio 0.6663 The ratio is CONSTANT WITHIN an output neuron and DIFFERENT BETWEEN them - a missing PER-NEURON factor, not a global scale. The value names it: 1/0.6663 = 1.501, and CLASS_LOGIT_SCALE * sigmoid'(z) = 6 * a(1-a) = 6 * 0.25 = 1.5 at a = 0.5, where a(1-a) is maximal. The output gradient is missing CLASS_LOGIT_SCALE * sigmoid'(z). WHY THE USUAL CANCELLATION DOES NOT APPLY HERE. softmax+cross-entropy gives dL/dz = p - target ONLY when the softmax acts on z directly. This head puts a SIGMOID and a temperature between them - z -> sigmoid -> a -> 6a -> softmax - so the chain keeps both factors and neither cancels. AND IT IS NOT ABSORBED BY ADAM. A constant scale would be; a(1-a) varies per neuron and per sample. The effect runs opposite to the intent recorded in calcOutputGradients, which omits the activation derivative to avoid vanishing gradients: here that leaves SATURATED output neurons with relatively TOO MUCH gradient rather than too little. SECOND DEFECT, in the shared harness. The first version of this test read connections off the OUTPUT layer. Connections point FORWARD - updateInputWeights reads prevLayer.At(n).Connections.At(m_myIndex) - so an output neuron owns none, the loop iterated an empty array, and TSummary printed ALL TESTS PASSED on 0 assertions. A suite that asserts nothing has proven nothing; it now fails loudly. That false green would have applied to every future test too. CNet gains one read-only accessor, Layers(), used by this test and nothing else - the check must perturb a weight, re-run the forward pass and read back the gradient, which is impossible from outside otherwise. NO FIX IN THIS COMMIT, deliberately. The measurement lands first and stays reproducible; correcting the gradient changes what every model learns and deserves its own commit and its own retrain. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
1016e52621 |
feat(features): mask v6 - drop crossasset and volume, and state what the capacity rule really says
Tag lean-2. Width 96 -> 72. crossasset scored 12 and 17 votes of 24, volume 12 and 13 - the four weakest survivors in the set, and this file already named volume "the block most likely to fall out of the set on the next screen". They were affordable when nobody was watching the budget. They are not now. THE RULE OF THUMB, APPLIED HONESTLY RATHER THAN QUOTED. Classical practice is 10-30 independent samples per parameter; 1 is the absolute floor. Measured live after this cut: chart member params indep obs obs/param SP500 Perceptron 2336 1139 0.49 SP500 Convolutional 2080 1139 0.55 SP500 LSTM/ConvLSTM 272 1139 4.19 BTCUSD Perceptron 2336 1231 0.53 NAS100 Convolutional 2080 867 0.42 Cutting 186 -> 96 -> 72 inputs moved the ratio from ~0.4 to ~0.6. It BARELY MOVED, and that is the finding: the binding term is not the input count, it is the 16-unit minimum first layer. Run it to the end - SP500 has 1139 independent observations, so 10 obs/param allows ~114 parameters, and at a 16-unit floor that is SIX INPUTS. Not 72. SO THE RULE DOES NOT SAY "make the network smaller". It says a network of any usable size is the wrong model at this sample count, and the only statistically supportable capacity here is roughly linear. The linear baseline has already been run on exactly this data: -0.1pp edge at 100% coverage over 116 independent calls. No edge. THREE MEASUREMENTS, ONE CAUSE. The in-sample book is 3-6x the out-of-sample one; capacity is 0.5 observations per parameter; a linear model finds nothing. All three are the same quantity: we have 1,100-3,600 independent observations, and the LABEL OVERLAP is what makes it so few. EURUSD turns 125,322 rows into 3,631 independent ones by dividing by the 34.5-bar leg-ride lifespan. The previous label's lifespan was 5 bars, so that one change cut the effective sample ~7x. The label horizon is the only lever that raises the evidence instead of shrinking the model to fit the lack of it. Nothing here fixes that; this commit just stops spending capacity on columns that were never earning it. RETRAIN-FORCING: FEATURE_MASK_VERSION 5 -> 6 rides the fingerprint, so every chart starts fresh. Verified live - 73 inputs (72 + bias) on all ten charts. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
03ecf11efd |
feat(features): mask v5 - drop alt and fracdiff, width 186 to 96
The mask exists to hold the first-layer budget effN/(width+1) near 60, which is what its
294 to 78 cut achieved. The LIVE keep-screen now reports that budget at 6.1 against a
16-wide floor. An order of magnitude below design, and inside the regime CTopology
already prints a warning for on nine charts.
THE WIDTH TRIPLED WHILE NOBODY WATCHED THE BUDGET. v3 re-added the alt block and v4 added
fracdiff, both KEPT BY CONSTRUCTION so the keep-screen could measure them - a fair deal,
because the screen only reports on columns that are actually emitted. Both measurements
have now come back, and both are empty:
alt -0.15pp ALT-only against price-only +5.08pp, and 60 of its 72 inputs are
mostly-zero repeats of the anchor bar
fracdiff CPCV ablation -0.47pp to +0.75pp against split standard deviations of 1.05
to 3.53pp, sign going both ways across six charts
The deal for a column kept by construction is that it goes when the measurement arrives.
They go. 90 of 186 inputs - 48 percent of the vector - were columns already measured to
carry nothing.
AND v3 GOT THE COST WRONG, which is how it got away from us. Its note claimed the alt
block cost twelve inputs rather than twelve per bar because the window layout zeroes
repeated readings. Zeroing does not remove an input: the first layer still carries and
fits a weight for every one of the 72. The live width proves it - 31 columns x 6 bars is
186, which only adds up if alt is 12 per bar.
Width 186 to 96, budget 6.1 to about 11.8. Still under the 16 floor, so this is a step
and not a cure - but every input removed was measured worthless first, so it costs
nothing to find out.
The alt fetch pipeline stays ENABLED and both blocks keep being BUILT, so re-enabling is
a one-line change and the finding stays falsifiable.
RETRAIN-FORCING BY DESIGN. FEATURE_MASK_VERSION rides in the model fingerprint, so 4 to 5
means no chart can load its old weights and every one starts fresh at era 0 - which is
exactly what the era cap in
|
||
|
|
ba8bc8e1c1 |
fix(training): stop at 100 eras - the run was picking the best of 1337 re-scorings of one window
Tag era-cap-1. MaxErasPerRun 10000 -> 100, and the era cap stops prompting. THE ERA CAP WAS LABELLED A "runaway backstop, not a training control" and left at 10000 because "the plateau ladder decides when a run ends". The ladder does not end runs. PLATEAU_PATIENCE_ERAS is 8, but ANY new best resets the counter, and on a noisy score a new best arrives by luck often enough that the ladder wanders indefinitely. Measured live today: SP500 era 1337, XTIUSD 1126, USDJPY 721, EURUSD 638. WHY THAT IS NOT FREE, and it is the mechanism behind the overfitting the operator has been pointing at all day. Successive eras RE-SCORE THE SAME OOS WINDOW - they add no independent observations at all (project_window_cut_verdict: "effective n at a rung is coverage x OOS_bars / lifespan, computed ONCE, not once per era"). So the checkpoint is the MAXIMUM of N draws of a noisy statistic, and the winner's OOS score is inflated purely by construction. It gets worse with every era the run survives: chart current era BEST era selection family SP500 1337 1323 best-of-1337, book now NEGATIVE XTIUSD 1126 1057 best-of-1126 USDJPY 721 666 best-of-721 EURUSD 638 624 best-of-638, book now NEGATIVE AND NOTHING IS GAINED PAST ~ERA 20. Measured 2026-08-27 across six charts and two runs: every precision-on-era slope under 2.5pp per 100 eras, signs disagreeing across charts AND across runs, and all six still holding their ERA-20 checkpoint at era 66-71. "In-sample error keeps falling; OOS does not follow." THE LIVE NATURAL EXPERIMENT ran itself today. The two charts that restarted sit at era 69 with their best checkpoints at era 20 (USDCAD) and era 24 (XAUUSD) - exactly where that measurement said they would be. XAUUSD books +0.66, second best on the fleet. 100 keeps the burn-in (ENSEMBLE_CHECKPOINT_MIN_ERA 20) plus ~80 eras of candidates, ample for a ladder needing ~24 eras of patience to reach PLATEAU_STAGE_DEPLOY, and cuts the selection family by 13x. It also makes a full retrain roughly an order of magnitude faster, which is what makes experimenting on topology or labels affordable at all. THE PROMPT HAD TO GO WITH IT, and this was caught before deploying rather than after. PromptContinuePastEraCap raises a MODAL MessageBox on any non-tester chart. Reaching a 10000-era cap was rare enough to be worth interrupting for; at 100 the cap is the NORMAL path, so every chart would raise one. Ten charts, ten modal dialogs, terminal frozen until each is dismissed - on a fleet that is not a prompt, it is an outage. It now prints the same text and deploys the best checkpoint, which is what the plateau ladder does anyway. NOT YET DEPLOYED. Every live chart is already past era 100, so a restart on this build would have them all hit the cap at once and ship their existing late-era checkpoints - the very ones this commit argues are noise-selected. The change only pays on FRESH runs, so it wants a fleet reset to go with it. Not retrain-forcing by itself; worthless without a retrain. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
1c329004c7 |
feat(gate): positive expectancy is the whole bar - stop charging spread, stop requiring alpha
Tag expectancy-1. Two operator decisions, both recorded with the reasoning so neither gets
quietly "fixed" back.
SPREAD IS NO LONGER CHARGED. The book being judged is a ZIGZAG LEG RIDE, and a leg is
dozens of times the bid-ask difference. Measured on this fleet, the round-turn spread is
0.008 ATR on BTCUSD against books of +0.17 to +0.52 - about 4%. The operator reports the
same result in SQX, where setting 0, 6 or 60 pips does not move the outcome, and that
their broker does not bill it as a separate line. Spread is still SAMPLED once per bar and
still PRINTED, so the number stays visible; it simply stops gating.
THE CONSEQUENCE, STATED UP FRONT RATHER THAN DISCOVERED LATER: indices pay no commission
either, so SP500, DAX40 and NAS100 now have a cost of EXACTLY ZERO and their test reduces
to "book > 0". That is literally positive expectancy, which is the stated objective.
Per-chart, what stops being charged (from the live init lines):
SP500 0.68 | DAX40 1.73 | NAS100 2.20 -> all become zero cost
XAUUSD 0.55 of 0.60 | XTIUSD 0.09 of 0.11 | BTCUSD 3.58 of 28.00
USDJPY 0.003 of 0.009 | EUR/GBP/CAD ~0.00003 of ~0.00007
ALPHA IS REPORTED, NOT GATED. It gated for roughly an hour. In the operator's words: "We
do not require beating buy and hold either. I already explained the goal : positive
expectancy, simple as that."
THE OBJECTION WAS RAISED AND OVERRULED, which is the right order and is recorded rather
than re-argued: a book paying less than the drift carries the DRAWDOWN PROFILE of the
underlying trend, and a prop account fails on drawdown. Against that - buy-and-hold is
not a strategy a prop account can run, nobody pays a prop trader for alpha, and DAX40
booking +0.17 against an always-long +0.234 still MAKES +0.17. Operator's account,
operator's call.
The refusal branch became unreachable (tradeableOK is now ratesOK AND bookPays, so a
paying book implies deployable) and was converted into a DRIFT WARNING that prints on the
deployable era instead of withholding it. Same information, no longer a block.
THE GATE IS NOW: measurable, covers the base rate, fires both ways, and the book beats
commission. Four conditions, down from six this morning, and every one of them is a thing
an operator asked for.
Not retrain-forcing.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
50f71f42a3 |
feat(gate): measure the IN-SAMPLE book and compare it to the out-of-sample one
Tag is-oos-2. The operator's ask, in their words: "the goal is to train neural networks to recognize the patterns as best they can, deploying once we are not making progress for x number of eras. performance of is and oos should be similar. just like SQX does." The plateau ladder already IS "deploy once no progress for X eras". The IS-vs-OOS comparison did NOT exist: this project scored the out-of-sample slice only. There is an in-sample MSE as a training diagnostic, but no in-sample BOOK to compare the OOS book against, so nothing in the gate could see a model that had memorised its training span. This adds it. SAME RUNG, SAME LABEL, SAME DECISION RULE - the only difference is which bars. The derivation is copied from the OOS branch deliberately (softmax with the as-of leg side, then the prior-corrected adjusted signal) so the two books differ by the DATA and not by the decision rule, which is the whole point of comparing them. ITS OWN ARRAYS, NOT A FLAG ON THE OOS ONES. The vote rows feed coverage, precision, the chance rate, the cost mean and the exact binomial bar - they ARE the deploy gate. Letting in-sample rows into them would corrupt every one of those silently. Seven arrays instead of twenty-two, because only the book is wanted. THE FIRST CUT OF THIS WAS DEAD CODE, and it shipped as is-oos-1 before the fault was found. The in-sample scoring was placed inside the pass-3 loop, which starts at oosCutoff-1 and counts DOWN to 2 - so it walks the OUT-OF-SAMPLE bars only, and the in-sample ones are the higher indices it never reaches. The guard could never be true and the measurement never printed. It now rides its own descending cursor, stepped once per OOS bar, which also means it INHERITS this loop's yielding and can never become the blocking pass that would livelock an era (project_era_slower_than_bar). SAFETY CHECKED BEFORE WIRING, not after: ApplyClassificationSoftmax and AdjustedSignalFromSoftmax were both read for member side effects. Both mutate TempData and nothing else, so scoring an in-sample bar cannot reach the operating-point fit or any OOS tally. The in-sample step runs BEFORE the OOS bar rebuilds its own window, since TempData is shared. No backprop anywhere in it - this is a scorer, and training on these bars again inside it would be a second unshuffled epoch. Also caught before it could mislead: the OOS path adds signedVote RAW, because LiveVoteContribution already carries m_weight x the tier's pattern weight and voteWeight only builds the divisor. The first draft multiplied them again, which would have squared the weight and made the two books measure different things - a false overfitting signal from the very comparison built to detect one. MEASURED AND NOT GATED, same discipline that let the split-half test be rejected on evidence within the hour rather than becoming another closed door. Bounded cost: at most one extra feedForward per OOS bar, strided to ~ENS_IS_TARGET_BARS samples. WHAT IT CANNOT SEE, printed in the line itself so the number is not overread: our other overfitting route is SELECTING the best era and rung ON the OOS slice, and this comparison is blind to it. Only a third slice that neither trains nor selects would speak to that. Not retrain-forcing. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
f1268a16d9 |
feat(cost): the broker charges by ASSET CLASS - commission schedule, detected not guessed
Tag commission-1. Cost_CommissionPerLotPerSide (one number, shipped at 0.0, never set)
is replaced by the operator's actual contract, applied per class:
indices none forex 4 USD per lot
crypto 0.03% notional metals 0.001% notional energy 0.03% notional
CLASS IS DETECTED FROM SYMBOL_PATH, which is what the broker itself organises its tree
by, with symbol-name and SYMBOL_TRADE_CALC_MODE as fallbacks for a flat Market Watch.
Verified in situ on all ten live charts rather than asserted - every one resolved
correctly from its own path (Indices\, Forex\, Crypto\, Energy\, Precious_Metals\).
A PERCENTAGE OF NOTIONAL NEEDS NO CONTRACT SIZE AND NO FX RATE. Commission in money is
pct x contractSize x price x (quote->account rate); the price-equivalent divides by
money-per-price-unit, which is tickValue/tickSize - and tickValue already carries the
same contractSize and the same rate. They cancel exactly, leaving
price_equiv = pct x price for ANY quote currency
so nothing stale or missing can be read. Worth stating because it looks too easy.
THE FACTOR-OF-TWO, and the two branches need it OPPOSITE ways round. The percentage
branch builds the round turn itself, so a per-side quote is MULTIPLIED by 2. The flat
branch hands a per-side figure to WarriorCommissionRoundTurnPrice, which does its own
doubling, so a round-turn quote is HALVED on the way in. Writing them the same way round
would have been a factor-of-four error between asset classes. The operator's figures are
read as the FULL ROUND TURN (Cost_CommissionIsRoundTurn, default true), which is how a
prop contract quotes it; false doubles every figure without editing any of them.
MEASURED EFFECT - COMMISSION DOMINATES SPREAD ON HALF THE FLEET, and the gate has been
charging spread alone until now:
BTCUSD comm 24.42 + spread 3.58 = 28.00 cost was UNDERSTATED 7.8x
USDJPY comm 0.006 + spread 0.003 = 0.009 3.0x
EURUSD comm 0.00004 + 0.00002 = 0.00006 3.0x
GBPUSD comm 0.00004 + 0.00003 = 0.00007 2.3x
XAUUSD comm 0.04 + spread 0.55 = 0.60 1.1x
indices comm 0.00 unchanged
In ATR terms BTCUSD goes 0.008 -> ~0.063 and still clears on a +0.52 book, but GBPUSD
goes ~0.014 -> ~0.033 against a +0.03 book and should now FAIL. That is the correct
outcome: it was the thinnest book on the fleet and it was being charged a third of its
true cost.
The resolved class, both cost halves and the spread sample count are PRINTED at init, so
a symbol filed in an unexpected folder shows up as a wrong class rather than as a
silently wrong number.
Not retrain-forcing.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
728bc647ba |
feat(gate): the book must beat the DRIFT, not just its cost - and the split-half test is rejected on measurement
Tag alpha-gate-1. Second condition added to the ensemble deploy gate:
tradeableOK = voteGate.ratesOK && bookPays && alphaPays
WHY book > cost WAS NOT ENOUGH. book-gate-1 admitted DAX40, which books +0.20 ATR per
call against a 0.042 cost - a healthy 4.8x - while the always-long book over the same
window pays +0.234. The vote earns LESS than passively holding: its alpha is NEGATIVE,
measured -0.057 +- 0.073 over 18 consecutive eras. That is beta sold as signal, and it
stops paying the day the trend turns, which is exactly when a prop account's drawdown
limit is being tested. Neither more training nor a higher rung fixes it - the model has
found the trend rather than the turns - so the refusal now says that in those words
rather than leaving an operator to conclude the model is merely undertrained.
AND THE SPLIT-HALF CONSISTENCY TEST IS REJECTED, ON ITS OWN EVIDENCE. It shipped in
|
||
|
|
ead2b12838 |
fix(persistence): ArrayCopy grows a destination but never shrinks it
LoadNetWithRetry copied op.indicatorParams into the caller's array without resizing it
first. ArrayCopy GROWS a dynamic destination and never SHRINKS it (mql5book.pdf p1151),
and this particular destination is REUSED: BuildOrLoadTopology calls LoadNetWithRetry
twice into the same loadedIndicatorParams - once on the pure-MQL5 path, then again on
the DLL fallback when the first fails.
So a second load carrying FEWER params would leave the first load's values sitting in
the tail, and the reader's guard is a SIZE check -
if(netLoaded && ArraySize(loadedIndicatorParams) == AD_TUNE_PARAM_COUNT)
AdoptIndicatorParams(loadedIndicatorParams, indicators);
- which a stale tail can satisfy while holding the wrong numbers.
STATED HONESTLY: this is the contract being honoured, not a fault observed. Both calls
read the same file, so in practice the two counts agree and the tail never diverges.
One ArrayResize closes it anyway.
From the API-contract list in the full read of docs/mql5book.pdf. The other items on
that list came back clean: OrderSend is never called directly (CTrade owns retcode
handling, p1208), StringToCharArray is never used (p1780), and the one CopyBuffer read
already guards with != span rather than treating a short read as success (p1473).
Not retrain-forcing. Compiled clean; NOT yet deployed - the fleet is running book-gate-1
and a redeploy would cost the era measurements it is currently gathering.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
7606517dcd |
feat(gate): deploy on the book, not on the significance of the win rate
Tag book-gate-1. The ensemble deploy condition was
tradeableOK = voteGate.tradeable && bookPays
and voteGate.tradeable ANDs in `precision > chance + EDGE_MIN_SIGMAS x SE`. That test is
now REPORTED instead of GATING. Measured on the live fleet, one era per chart:
chart verdict book cost zero-skill alpha precision vs its bar
NAS100 DEPLOY +0.870 n/a +0.163 +0.707 37.8% vs 33.6%
XTIUSD DEPLOY +0.590 n/a +0.087 +0.503 31.7% vs 30.4%
XAUUSD refused +0.660 n/a +0.256 +0.404 34.9% vs 36.4%
BTCUSD refused +0.370 0.008 +0.032 +0.338 24.6% vs 30.7%
USDJPY refused +0.290 0.014 +0.124 +0.166 29.9% vs 30.7%
GBPUSD refused -0.010 0.012 -0.030 +0.020 27.2% vs 27.1%
DAX40 refused +0.180 0.042 +0.234 -0.054 29.3% vs 37.5%
EURUSD refused -0.130 0.018 -0.053 -0.077 23.1% vs 24.1%
USDCAD refused -0.040 n/a +0.070 -0.110 24.4% vs 27.6%
SP500 refused -0.020 0.074 +0.248 -0.268 25.1% vs 32.2%
THREE BOOKS PAYING 20-50x THEIR OWN COST WERE BEING REFUSED, and not for any reason to do
with the model. effN deflates by the 34.5-bar label overlap, so the bar is chance +4.3pp at
USDJPY's 14,584 calls and chance +15.0pp at DAX40's 1,385. BTCUSD was asked for a 14pp
precision edge over chance on roughly 46 independent observations. Nothing produces that.
This was not a strict gate, it was a closed one, and it was closed by call COUNT rather than
by edge.
IT ALSO CONTRADICTED THE GATE NEXT TO IT. WarriorRungBookProfitable's own header argues that
demanding significance "would deploy NOTHING, ever, which is the prove-your-edge trap this
codebase has backed off twice", and gates the book on a POINT ESTIMATE for that reason.
Then tradeableOK demanded significance anyway and overrode it. Two gates, opposite
philosophies, and the closed one won every time.
AND PRECISION IS THE WRONG QUESTION, which this same file already said twelve lines away:
precision "says a turn was CALLED, not that the leg after it paid". BTCUSD calls 24.6% of
turns right and earns +0.37 ATR per call, because its winners are much larger than its
losers. A gate on the win rate cannot see that; a gate on the book can.
NEITHER SUPPLIED BOOK RECOMMENDS A SIGNIFICANCE GATE. mql5book.pdf pp. 1482-9 optimises and
then FORWARD-TESTS, ranking on a criterion built to reward a smooth equity curve - signed by
slope, weighted by sample size. It never asks a win rate to clear a sigma bar, and it is
honest about the yield (of its top 1000 in-sample passes, 323 were profitable forward).
neuronetworksbook.pdf has no cost model in 690 pages. We invented this bar ourselves.
WHAT CHANGED
* SDeployVerdict splits its verdict into ratesOK ("is this a strategy at all" - measurable,
covers the base rate, fires both ways; none of which needs statistical power to read) and
edgeOK (the significance test). tradeable = ratesOK && edgeOK, UNCHANGED, so the member
gate and its ranking are untouched. Only the ensemble reads ratesOK.
* The rung selection reads ratesOK too. Selecting the certified operating point through a
test the verdict no longer applies would have picked the rung by the wrong criterion -
and on a low-coverage chart would have picked none.
* The Sidak family-wise test no longer gates. It corrects the PRECISION z, and precision is
no longer the decision; leaving it in the conjunction would have kept the closed gate
closed through a side door. It is still computed and printed, because a best-of-N maximum
is still a maximum and an operator should see how selected the number is.
WHAT STILL GATES: ratesOK, and the book beating its own measured per-rung cost.
WHAT IS NOW MEASURED AND DELIBERATELY NOT GATED: the SPLIT-HALF BOOK. The window is split at
its median timestamp and the ride book reported for each half at the certified rung. This is
the CONSISTENCY question - did the book hold up across the window, or is the total one good
stretch carrying a bad one - and it is what the coding book ranks on instead of significance.
It prints on every era. When there is fleet evidence that it discriminates a real book from a
selected one, it becomes the deploy condition. Replacing an unreachable bar with an unmeasured
one is precisely how the precision gate closed, so it is not being done blind.
THE HONEST CAVEAT, stated rather than buried: the book is still a maximum taken over
eras x rungs, so it is still a selected number, and at effN ~46 with ~3.1 ATR per-call SD its
own standard error is near 0.46 ATR - BTCUSD's +0.338 is under one sigma. This is a
point-estimate gate, knowingly. The controls that remain are the plateau ladder (this branch
is not reached until the best vote survives PLATEAU_STAGE_DEPLOY-1 warm restarts with no
improvement), the coverage and both-sides checks, and the live expectancy stop - which is the
only one of the three that judges the book the account actually experiences.
Three stale comments corrected to match (the sweep's [DERIVED] explanation, the stage-3
refusal attribution note, and the refusal branch's own condition).
Not retrain-forcing.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
7fcd7b0427 |
fix(journal): a position's ticket is not its identifier, and history is keyed by the identifier
CTradeJournalManager keyed its whole open-trade tracker on the value PositionGetTicket()
returns, and passed that same value to HistorySelectByPosition(). That function matches
DEAL_POSITION_ID, which is the position IDENTIFIER. The two coincide for an ordinary
position that opens and closes once, which is why this has never been visible.
THEY STOP COINCIDING ON EXACTLY THE PATHS THAT RENUMBER A TICKET (mql5book.pdf pp. 1228,
1322): a netting reversal, clearing, and - the one that reaches this fleet - a symbol whose
SYMBOL_SWAP_MODE is one of the REOPEN_* variants, where the broker force-closes and reopens
the position at rollover. POSITION_IDENTIFIER never changes across any of them.
TWO DISTINCT FAILURES, both silent:
* IDENTIFIER != TICKET (netting reversal, clearing). HistorySelectByPosition(ticket)
selects an EMPTY history, ResolveClose returns false, and the caller is written to keep
tracking and retry next tick - by design, for the case where history has not caught up
yet. Here it never catches up. The trade is never journalled and the tracking slot is
never freed, so the retry runs every tick for the life of the chart.
* ROLLOVER REOPEN. The ticket changes while the position is genuinely still open, so the
old ticket vanishes from PositionsTotal() and the new one is unrecognised. One trade was
recorded as two: the first written at the reopen price with its MAE/MFE truncated at
rollover, the second opening at that same price with its excursions reset to zero.
This is not cosmetic, for the same reason the DEAL_FEE omission in
|
||
|
|
34649e215e |
fix(journal): realised P&L was missing DEAL_FEE, which is not DEAL_COMMISSION
ResolveClose summed DEAL_PROFIT + DEAL_SWAP + DEAL_COMMISSION. That is three of the
four components MQL5 charges.
DEAL_FEE IS A SEPARATE CHARGE, not a synonym for commission: MQL5 defines it as
"fee for the deal which is charged immediately after the deal", and a broker may
levy one, the other, or both. The canonical accounting is
DEAL_PROFIT + DEAL_SWAP + DEAL_COMMISSION + DEAL_FEE.
So every closed trade's realised P&L was overstated by exactly the fee. That
matters more here than the size of the number suggests, because this figure is not
cosmetic: it feeds the EXPECTANCY STOP (Variables\RiskBudget.mqh) and the tier
ranking that sets each member's vote weight. An overstated result makes a losing
book look break-even to the one guard that is supposed to halt it.
Found by reading docs/mql5book.pdf end to end (p1338 for the property, pp. 1486 and
1535 for the book's own accounting). Two comments claiming ResolveClose "sums all
three" corrected to four - a stale comment counts as a guess.
STILL OPEN, from the same read, and deliberately not bundled:
* DEFERRED-COMMISSION BROKERS READ AS ZERO. If commission is charged at period
end rather than per deal, DEAL_COMMISSION is 0 on the trading deals and the
charge arrives as separate DEAL_TYPE_COMMISSION / _DAILY / _MONTHLY balance
deals. ACCOUNT_COMMISSION_BLOCKED != 0 detects such an account (it is
permanently 0 on a per-deal broker). Our journal would silently report zero
commission there.
* SWAP IS FORWARD-READABLE AND THE GATE IGNORES IT. Unlike commission,
SYMBOL_SWAP_LONG/SHORT are real symbol properties. The measured mean hold is
17-19 H1 bars, ~0.75 of a day, so positions cross rollover and pay it; the
book's own measurements put XAUUSD at -12.60 points long and AUDUSD at -14.80
short, the same order as some of our spreads. ENUM_SYMBOL_SWAP_MODE has nine
variants, so this needs the same "handle what is measurable, report the rest as
unknown" discipline the commission input already uses.
Not retrain-forcing.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
cc3b0a219a |
fix(cost): measure the round turn from live Ask-Bid and an input, not from bar history
Corrects |
||
|
|
d70efc0f74 |
fix(gate): the book had to beat zero, not beat what it costs to collect
WarriorRungBookProfitable tested (hLong+hShort)/n > 0.0. The round-turn spread was
computed a few lines away in the same function, printed in the payoff line, and
carried an explicit comment saying it gates NOTHING - "spreadR is a live snapshot
to be judged by hand".
That objection was correct about the QUANTITY and wrong about the CONCLUSION.
SymbolInfoInteger(SYMBOL_SPREAD) is whatever the book looks like at the instant an
era happens to end, which is not what the historical trades in that window would
have paid - so it should not gate. But this project has already measured that "4 of
5 die to spread and the survivor dies on commission"
(project_cost_boundary_equilibrium), and EXPECTANCY IS THE BAR. The answer is to
measure cost properly, not to leave it out of the gate.
MEASURED, NOT SNAPSHOTTED. g_ensVoteCostR carries the round-turn spread at each
vote row, taken from the broker's own per-bar history (CopySpread, already
maintained as m_spreadSeries for the feature block) and divided by that bar's ATR,
so it lands in the SAME units as the ride book it has to beat. Filled once per ROW
in the block that already measures the ride and the excursions, for the reason
stated there: cost is a property of the CHART at that bar, identical for every
member.
PER RUNG, over the bars that rung actually fires on - not a chart-wide average.
Signals cluster, and a cluster can sit in a wider-spread regime than the window
mean, so a window average would understate the cost of exactly the bars being
traded. sweepCost/sweepCostN accumulate alongside sweepFired.
AN UNKNOWN COST DOES NOT WAIVE THE TEST. g_ensVoteCostR is -1 where the spread
series or the ATR was unavailable and those rows are DROPPED from the mean rather
than read as zero - a cost of zero is the one answer that can never be right. If a
rung has no measurable cost at all the gate falls back to the pre-existing `> 0`
bar, which is "cannot price this", not a free pass.
NOT a multiple of the cost. The "expected payoff should be at least double the
spread" rule is a common heuristic (and is what prompted this - MQL5 article 8410,
Ilin, sent by the operator) but it is not measured here, so it is not imposed.
Sampling error on the book is already carried by the exact binomial bar the same
verdict applies to precision.
WHAT IT CHANGES, on the fleet as it stands. Median book per call against the
census's measured spread/ATR:
SP500 +0.02 vs ~0.065 -> now correctly REFUSED (was certified)
GBPUSD +0.05 vs ~0.020 -> marginal
DAX40 +0.08 vs ~0.029 -> marginal
BTCUSD +0.23 vs ~0.006 -> clears
XTIUSD +0.57 vs ~0.083 -> clears
NAS100 +0.75 vs ~0.025 -> clears
Both charts currently DEPLOYABLE clear it comfortably, so this closes a hole rather
than reversing a live decision - but SP500's book was being certified while sitting
below its own spread.
The sweep line now prints cost beside the book at every rung, so BOOK-FAIL names a
visible reason instead of an invisible one, and its legend says what the number is
and what "n/a" falls back to.
Three call sites updated, not one - the rung walk, the derived-rung verdict and the
sweep report all ask this question (feedback_rename_leaves_readers_behind).
Not retrain-forcing: no fingerprint member moves, and g_ensVoteCostR is an in-memory
per-era buffer.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
b5d2bc1622 |
fix(ensemble): the vote's scale outgrew its rung grid and the fleet went silent
SYMPTOM: every chart's era verdict reported 0.0% coverage at EVERY rung of the
threshold sweep. Only DAX40 (6.1%) and USDJPY (0.9%) had any coverage at all. Seven
of ten charts could not fire a trade at any operating point the derivation was
allowed to choose.
IT IS NOT A DEFECT IN THE MODELS. LiveVoteContribution is
moduleWeight * (tierWinRate - measuredChancePct), so THE VOTE'S UNIT IS PERCENTAGE
POINTS OF EDGE OVER CHANCE. Measured from the terminal's own tier-ranking lines
("pooled X% raw -> Y% shrunk toward the Z% coin-flip rate"), medians:
2026-09-01 win 38.6% chance 13.4% EDGE 25.2pp
2026-09-02 win 20.5% chance 10.6% EDGE 9.9pp
2026-09-03 win 22.4% chance 19.6% EDGE 2.8pp
Two separate steps, both legitimate:
* 09-01 -> 09-02 the WIN RATE halved (38.6 -> 20.5) - the leg-ride label (
|
||
|
|
e13e317e54 |
feat(signal): one operating point per side - a single global cut had silenced the sell book
User report: "some charts are not showing sell signals at all, only buys". It is
real. By-side ride book and n on the fleet's most recent era per chart - the
drift-free test the era log already prints:
NAS100 160 long vs 5 short - 32 : 1
DAX40 30 long vs 0 short - no short book at all
(Earlier eras from the same charts read 744:1, 260:1 and 41:1, but those lines
predate today's restart and are quoted nowhere as current - the two above are the
post-restart measurement.)
THE LABEL IS NOT THE CAUSE. P(ride pays | up-leg) vs P(| down-leg) is 0.99x to
1.55x across the fleet and up-legs are ~50% of bars. A 1.2x rate asymmetry was
being amplified into 744:1.
THE MECHANISM is the operating point's depth. FitDirConfThreshold places one cut
where coverage matches the pooled directional label rate, and at 3-25% coverage
that cut sits deep in the tail. There, a small shift in one side's margin
distribution moves nearly ALL of that side across it: the threshold hits its
target POOLED and misses it per side by two orders of magnitude. The model was not
failing to see sells - BTCUSD's pre-threshold sell recall is 84% and its raw argmax
is B45/S40 - the single cut was discarding them.
Each side now clears the margin ITS OWN measured label rate asks for.
m_labelPrebuildBuyCount/SellCount already existed, so the two targets are
measurements, not choices.
THE PROPERTY THAT MAKES THIS SAFE, AND IT IS EXACT: the per-side targets are
100*buy/tot and 100*sell/tot against the same denominator the pooled rule uses, so
they SUM TO THE POOLED TARGET. Total coverage is preserved; only its split between
the books changes. This redistributes exposure rather than increasing it.
THE TRAP, NAMED:
|
||
|
|
3c0e2facb2 |
fix(chart): a signal mark's persisted time was the left edge of its line, not its bar
User report: "some labels are completly off (arrow not at the low/high, and entry
price way higher than candle body and much much away from what the spread would
add)". Not the training label, and not a fill - the only order attempt on the day
was refused by the client. Both halves are one display defect.
A mark is TWO objects since 2026-08-20: an OBJ_TREND segment from t-half to
t+half so it is wide enough to see, plus an OBJ_ARROW glyph. half is
PeriodSeconds * WARRIOR_SIG_LEVEL_HALF_SPAN = 4680 s on H1, so OBJPROP_TIME index
0 of the line is its LEFT EDGE. FIVE call sites read index 0 and every one of
them treated it as the bar time, because "which bar does this mark belong to" had
no single implementation and each site re-derived it from rendering geometry.
MEASURED on the live files, and it COMPOUNDS: Snapshot() wrote t-half, the restore
redrew a line centred on the stored value, and the next save read THAT line's left
edge. Predicted residues are the 10 multiples of 360 s, not the 60 possible
minutes - observed 704 of 704 persisted vote arrows across fifteen files on the
k*1080s ladder, ZERO off it, with k=1 where a mark had been saved once and k up to
6 where it had been carried through sessions. XAUUSD and EURUSD sat at k=6: 7.8 H1
bars adrift, which is why a mark's stored trigger price was nowhere near the candle
it was drawn over - the price was right for a bar 7.8 bars away. The arrow half
lost its low/high anchor for the same reason: a shifted time is not a bar open, so
iBarShift(exact) returned -1 and the glyph fell back to the trigger price, landing
inside the candle body.
WarriorSignalMarkBarTime() is now the one accessor, and it takes the MIDPOINT of
the two ends rather than t0+half: the midpoint recovers t exactly by construction
AND stays correct if the half-span is ever retuned, including for marks already
drawn under the old value. Its t1 <= t0 branch answers correctly for the
single-anchor arrow half too, which the rescan sweep needs. Routed through it:
CVoteArrowStore::Snapshot, CChartUI::SaveChartSignals,
WarriorReconcileVoteCooldown, WarriorLatestVoteArrowTime, and the rescan's
typed-blind scope sweep.
Two more defects the same read was hiding:
* the cooldown reconciliation deleted by a RECONSTRUCTED name built from the
shifted time it had just read, so it could not remove a freshly drawn arrow at
all and its log over-reported kills. The name now travels with the time
through an insertion sort over both arrays.
* the rescan's scope sweep gave the two halves of ONE mark two different times,
so at the window edge it deleted an arrow and left its line - exactly the
split WarriorDeleteSignalMark exists to prevent.
* WarriorLatestVoteArrowTime seeded the live cooldown clock 1.3 bars EARLY on
every restart, so the first signal after a restart could fire inside the
window the arrow on the chart was enforcing.
WarriorSignalMarkOnBarGrid() stops the drift surviving a restart. Arithmetic
rather than a history search, and the residue is READ off iTime(sym,period,1)
instead of assuming UTC alignment, because where the bar grid sits in epoch
seconds is the broker's day start. It KEEPS on an unknown - no history yet, or a
weekly/monthly frame that is not a modular grid - since deleting on an unknown is
the failure mode that cost this chart 272 of 273 arrows in
|
||
|
|
8280a7cd40 |
feat(fleet): open the charts the census says are worth having
The census answers which instruments carry deep history at a spread the book can cover; this is the half that acts on it, so adding an instrument is a reviewable list in source control rather than six manual chart operations nobody can reconstruct later. BTCUSD, GBPUSD, NAS100, DAX40 - every one with more than twelve years of server H1 and a sampled spread under 3% of ATR, against a measured ride book of +0.05 to +0.63 ATR per call. BTCUSD is the cheapest deep instrument the broker offers (0.63%) and the only one in a different asset class, which is worth the most to a pooled certificate that declines the diversification credit. DAX40 trades a different session, so its bars are not the same hours as everything else. IDEMPOTENT: a chart is opened only when none exists on that symbol and period, so a restart re-opens nothing and the list can be edited freely. That is what makes it safe to leave armed. THE TEMPLATE CARRIES THE EA. ChartSaveTemplate on the running chart captures this Expert Advisor and its inputs, so a new member comes up configured exactly like the one that spawned it - which is the point: a fleet whose members differ by attach order is how four charts ended up on a 10-bar cooldown and two on 30. Leased like the census and the alt-data fetch, because six instances would each try to open the same four charts - and the charts this opens initialise an EA that reaches this same code. COST, STATED IN THE FILE: four added charts are sixteen more models on a six-core box already at ~72% with six charts. Expect eras to slow across the whole fleet. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
0a68a0511b |
fix(census): an unsynced history depth is UNKNOWN, not zero
The same defect as the spread-reads-0 bug, one field along, and found the same way - by noticing a number that could not be true. SERIES_TERMINAL_FIRSTDATE needs the symbol synced with the server, which does not happen for an unquoted symbol within one sampling window of a terminal restart. So BTCUSD reported "0.0 years" of H1 history and would have been filtered out of the candidate list as having none. Measured across two runs twelve minutes apart: BTCUSD, NAS100, DAX40, UK100 and US30 all read n/a on the freshly restarted terminal and all five reported real first dates (2013, 2011, 2008, 2008, 2008) on the one that had been up longer. Written as UNKNOWN now, so it cannot be sorted or filtered as though it were shallow. Zero for an unmeasured quantity is the error this codebase already names in its rate gates - "-1 IS NOT MEASURABLE, never 0, because a gate reading 0 would treat an unmeasured quantity as a failure" - and a census is no different. WHAT THE SAMPLING ALREADY CHANGED, for the record: ten readings thirty seconds apart moved USDCAD from 4.76% of ATR to 6.23% (+31%), SP500 6.00 -> 6.49 and XTIUSD 11.87 -> 11.95. The single snapshot understated the cost of the very instruments whose books are thinnest, which is exactly where it mattered. Ships with the next deploy; the fleet is mid-era and this is a report, not a gate. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
7a5dddc61a |
fix(census): one chart, not six - and sample the spread instead of snapshotting it
TWO DEFECTS IN MY OWN CENSUS, both found by reading its output rather than by the compiler. IT RAN ON ALL SIX CHARTS. The guard was a plain global, and MQL5 globals are per PROGRAM INSTANCE - six charts are six instances, so every one of them walked all 64 symbols and overwrote the same file. Replaced with the atomic GlobalVariableSetOnCondition lease that System\AltDataFetch.mqh already uses for exactly this problem: six charts racing one shared file. Losing the lease is the correct outcome, not an error. THE SPREAD WAS ONE SNAPSHOT. Cost is the number this whole decision turns on - the measured ride book runs +0.05 to +0.63 ATR per call, so an instrument costing a tenth of an ATR a round turn has already spent most of what it could earn - and a single reading is one moment of one session. Spreads widen at rollover and around news. Now ten readings thirty seconds apart, reporting mean AND max, because the max is what says whether an instrument is quietly untradeable at the wrong hour. A reading only counts behind a real two-sided quote; a symbol with none reads NO-QUOTE rather than 0, since 0 would rank it as the cheapest instrument on offer. That is the same error the previous commit fixed one layer up, and the guard now sits at the sample rather than only at the report. WHAT THE FIRST GOOD RUN ALREADY SETTLED: 64 symbols openable both ways, and all 64 report a SERVER first-H1 date with server == local. So depth is real, not a sync artifact - the broker genuinely offers deep H1 on twelve instruments and added the other fifty-two in July/August 2026 with no history at all. The expansion universe is six symbols, not fifty-eight. Also promotes the derived-cooldown line out of PrintVerbose. It reports a change to the TRADING POLICY, and this codebase's rule is that a line reporting a state change never sits behind the verbosity gate - the tier-ladder restore learned that the expensive way. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
81efd64a7f |
feat(signal): derive the cooldown from the measured label lifespan, and census the broker's symbols
TWO CHANGES, ONE CAUSE: a per-chart input could not reach the fleet, and the fleet
had no data on which symbols were worth adding.
THE COOLDOWN IS NOW DERIVED, NOT SET. An input was the wrong shape for it twice:
* It is a measurable property of the LABEL, not a preference. The leg-ride label
resolves when the ZigZag leg flips, so the mean label lifespan IS the average
leg duration in bars. Two calls closer together than that concern the SAME
leg - the second pays a second spread for a move the first already owns. One
lifespan apart is where consecutive trades concern DISTINCT legs, which is
also what makes the deploy gate's independence assumption exact rather than
approximate: EffectiveSampleSizeDeclustered's divisor becomes 1.
* An input could not reach an attached EA. MT5 stores inputs per chart, so
"raising Signal_CooldownBars from 10 to 30 changed nothing on six live
charts" - and yesterday the fleet was found running 10 on four charts and 30
on two, two different trading policies inherited from attach order rather
than chosen. A derived value cannot drift that way.
The fraction is 1.0 because the argument picks it: less re-admits same-leg
duplicates, more declines distinct legs for no stated reason. At the measured
33-34.5 bar lifespan it lands within a few bars of the 30 the default intended.
Recomputed wherever the lifespan is measured - the label-cache build - so the
measurement and its consumer cannot drift apart. SignalCooldownOverrideBars still
wins, and switching declustering off entirely is still possible.
THE SYMBOL CENSUS answers "which symbols are worth adding" with data. The binding
constraint on this system is independent observations, and instruments are the
only lever that multiplies them, so it writes the three numbers that decide it:
history depth, spread against ATR, and whether both sides can be opened at all (a
close-only symbol can never satisfy the gate's two-sidedness test).
AND IT RUNS ON THE TIMER, NOT AT INIT, which the first version got wrong. At init
the terminal has just reconnected and nothing has a quote, so SYMBOL_SPREAD reads
0 everywhere - the first run duly ranked thirty untradeable-on-cost symbols as the
cheapest the broker offers. Caught by noticing GBPUSD reported a zero spread
beside EURUSD's 2 points. Now deferred three minutes, and a symbol with no tick is
written NO-QUOTE rather than a number, so the column cannot be sorted on by
mistake. It selects nothing: sixty symbols added to Market Watch is a change to
the operator's terminal, and this is a report.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
8892c9f17e |
docs(gate): correct the effect size - the live cooldown is 10 bars, not the 30 the default declares
The previous commit computed the independence gain from Signal_CooldownBars = SCB_30 and claimed a ~87%-independent traded stream. That is the DEFAULT, not the runtime value. MT5 stores inputs per chart in profiles\Charts\*\chart*.chr and an already-attached EA ignores a changed default - which is stated in this codebase's own comment beside that input, and is exactly why "raising Signal_CooldownBars from 10 to 30 changed nothing on six live charts". The fleet runs 10. The stream is ~30% independent, not ~87%. Read out of the live log instead of estimated, on a USDCAD era: before 8,669 signal calls -> effN 261, precision 22%, bar 24.9% FAIL after 1,385 traded calls -> effN 417, precision 23%, bar 23.7% fails by 0.7pp So the fix is worth about 1.6x the independent count and ~1pp of precision. It closes most of the gap rather than clearing it. Real, and smaller than advertised. AND RAISING THE COOLDOWN WOULD NOT HELP, written down so nobody tries it: effN = traded/(lifespan/gap), and traded itself falls as 1/gap, so the gap cancels. The independent observations in a window are bounded by window/lifespan whatever the spacing. Spacing trades cannot manufacture independence - only a longer window, more instruments, or a shorter label can. Comment-only; no behaviour change, no redeploy needed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
8fd2756aa5 |
fix(gate): the member gate certified a population the EA never trades
The deploy gate judged m_oos.DirCalls() - every bar the prior-corrected posterior
fired - and deflated it by the full 34.5-bar label lifespan. The EA does not trade
that set. It trades what survives declustering: live, a call the cooldown rejects
has its signal zeroed before the vote is published, so it produces no arrow, no
vote and no position.
The honest counts were already being computed, by the same CSignalDeclusterPolicy
object the live signal applies, on the same adjusted decision it feeds - printed
every era as "TRADED (declustered)". The gate simply never read them.
BOTH HALVES OF THE ERROR POINTED THE SAME WAY, which is what made it expensive:
* the traded set is measurably CLEANER. Live USDJPY eras: judged 24% where
traded was 26%, judged 26% where traded was 27%. The gate was reading a
precision the EA would never have realised.
* the traded set is far more INDEPENDENT. The cooldown is 30 bars against a
34.5-bar label, so consecutive traded labels overlap by at most 4.5 bars - the
stream is ~87% independent, not ~3%. Deflating it by the full lifespan applies
a correction the cooldown has already made.
On a live USDJPY era that is the whole verdict: judged 24% against a 25.4% bar
FAILS; traded 27% against a 24.7% bar PASSES.
AND THE SELECTION IS UNBIASED, which is what makes the traded precision usable at
all: the declustering keeps the chronologically FIRST bar of each run, never the
highest-confidence one, so this is not cherry-picking winners.
COVERAGE STAYS ON THE SIGNAL POPULATION, and that split is now explicit in the
signature rather than implied. Coverage asks "did this model call often enough to
be a strategy", which is a question about the signal; precision asks "were the
calls it took right", which is a question about the account. Coverage cannot move
to the traded set for a structural reason: a 30-bar cooldown caps traded coverage
at 1/30 = 3.3% while the floor is a quarter of a ~20% base rate, so every chart
would fail it forever for reasons unrelated to the model.
THE ENSEMBLE GATE STILL HAS THIS DEFECT and is passed through unchanged, on
purpose. Its OOS rows carry each member's raw adjusted vote - the decluster replay
is per-member state that never reaches the shared vote buffer - so there is no
traded population there to read yet. Fixing it needs the per-member decluster
decision carried on the vote row. One gate at a time, so the effect of this one
stays attributable.
Not retrain-forcing: gate arithmetic only, no fingerprint field moves.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
281f12300b |
docs(mask): the keep-screen does not vote on the run that adds a column
A correction to what the v4 comment claimed. It said the next screen would vote on the fracdiff columns "like everything else". It did not, and cannot: on a FRESH run the label cache is still filling while the screen's eight attempts are spent - "keep-screen DEFERRED ... usable sample holds 149 row(s) and 200 are required ... attempt 8 of 8 - GIVING UP for this run". The screen votes on a warm RESTART, never on the run that introduces a block. Anything new is unmeasured for at least one full run, and the comment now says so. Measured another way instead, and the answer was no: research/fracdiff_ablation.py in the Warrior_Research sidecar strips the 18 columns out of the published pool rows and re-runs the whole CPCV. The edge moves -0.47pp to +0.75pp against split standard deviations of 1.05 to 3.53pp, sign going both ways across six charts. LEFT IN ANYWAY, deliberately. That ablation is a linear model, so it shares this file's own caveat about the screen - it cannot see a column that is useless alone and useful in combination. Dropping the block on a linear null is exactly the over-reading the marginal-information warning exists to prevent. It costs 18 of 186 inputs; if the warm-restart screen also votes them out, drop them then. Comment-only: the mask predicate, the emitted width and the fingerprint are all untouched, so the deployed .ex5 is unaffected and no redeploy is needed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
6d266d4e62 |
feat(features): every price column here is memoryless - fractional differentiation
AFML ch.5. Every price-derived feature in this set is a FULL difference: (close-open)/atr, (close-MA)/atr, the 20-bar return, the leg extension. Differencing is what makes a series learnable - a model fitted on 2019 EURUSD levels cannot read 2026 ones - but a first difference is memoryless BY CONSTRUCTION. It keeps the last step and throws away the series. That is the trade this feature set has been making silently at every column. Fractional differentiation is the observation that the exponent need not be an integer. (1-B)^d for 0<d<1 interpolates between the raw level (all memory, not stationary, useless to a learner) and the return (stationary, no memory). Three orders are emitted - d = 0.3/0.5/0.7 - in ATR-relative units. A LADDER, NOT A CHOSEN d. AFML picks the minimum d passing an ADF test; there is no ADF test here, and adding one to tune a feature's parameter per instrument per era would fit the feature to the data before the model saw it. Three fixed orders go in and the FLEET KEEP-SCREEN votes on them, exactly as it has already deleted five ZigZag geometry columns and demoted the alt-data block. Kept by construction in mask v4 for one outing only: the screen reports only on columns that are EMITTED, so a block masked off at birth can never be measured. THE BUG THAT WOULD HAVE SHIPPED THREE DEAD COLUMNS, caught before deploy by checking the arithmetic at realistic price levels rather than on a series starting at zero: The weights of (1-B)^d sum to zero in the LIMIT - that is what makes it a differencing operator. TRUNCATED at 64 bars they do not: the residual is 0.222 at d=0.3. That residual multiplies the LOG PRICE LEVEL, so the raw sum carries 0.222 x log(2400) = 1.73 on XAUUSD against 0.222 x log(1.08) = 0.017 on EURUSD - an instrument-identity constant hundreds of times larger than the signal. Divided by the relative ATR it pins every bar to the clamp: measured 100% of XAUUSD and SP500 bars clipped, 68% of EURUSD. Three constant columns that FEATURE HEALTH would have reported only after a full retrain had been spent on them. Anchoring every term at this bar's log price removes exactly the level component and leaves a fracdiff-weighted combination of the multi-horizon RETURNS ending here - stationary, scale-free, same meaning on every instrument, and the long memory intact. Same series after: mean ~0, sd ~1.1-1.3, nothing clipped. The weight recursion is checked against two known values: at d=1 it terminates to [1,-1,0,...], the plain first difference, and at d=0.5 it reproduces the standard expansion of (1-B)^0.5 to six terms. REJECTS RATHER THAN DEGRADES at the oldest edge, unlike the swing and volume windows beside it. A 30-bar Donchian range is still a Donchian range; a fractional difference over a short window is a DIFFERENT OPERATOR - a different effective d - reported in the same column, which is a silently wrong number rather than a degraded one. Costs ~64 bars of 130,000. RETRAIN-FORCING twice over: the input width is field 4 of the fingerprint and FEATURE_MASK_VERSION rides it too. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
f063cfd6ba |
fix(calibration): the confidence rescale was blended with a per-SAMPLE time constant, once per era
m_confidenceCalScale has been frozen at its 1.0 constructor default for the life
of the mechanism, and the units error that froze it is one line:
m_confidenceCalScale += (eraScale - m_confidenceCalScale)
/ Net.recentAverageSmoothingFactor;
recentAverageSmoothingFactor is 10000. It is a PER-TRAINING-SAMPLE constant - it
exists to average the network error over ten thousand samples inside backprop.
Applied ONCE PER ERA it moves the value by one hundredth of one percent, so after
a hundred eras the scale has travelled 1% of the way to its target and after a
thousand it is still not halfway.
That reframes what was already known. ConfidenceBridge.mqh records that the
confidence is miscalibrated and that five trade-management modes were deleted
because of it. The mechanism meant to fix it was not merely shape-blind, as the
isotonic curve's commit message argued - it was NUMERICALLY INERT, and never had
the chance to correct anything at all. A quantity blended per era needs a per-era
time constant; borrowing one from a per-sample loop reads as a working mechanism
in every review, because the line is shaped exactly like a working EMA.
The new curve inherited the same divisor when it was written yesterday, which
would have frozen it at its first fit - visible in the log as a "carried" column
identical to "refit" to four decimals on every chart. Both now use CAL_BLEND_ERAS
(5 eras): slow enough that one noisy band cannot swing the live number, fast
enough to track a model whose output distribution moves every era.
VERIFY, DON'T ASSERT: the curve's log line now prints the scalar's current value,
so the first line after this deploy reports the number restored from .stats - the
value it reached over that model's entire history. 1.000 is the claim above,
measured rather than argued from the arithmetic.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|