forked from animatedread/Warrior_EA
492 commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
1cf4c57d57 |
fix(altdata): median-fill instead of zero-fill, and a one-shot feature-vector autopsy
ALT-DATA AUDIT. The files themselves are healthy - all six symbols, 6,073 daily
rows, 2010-01-01 to 2026-08-17, no constant or degenerate columns, sane tails
(mac_cpi/mac_unemp flat ~47d is monthly data behaving correctly). The problem is
not the data, it is what happens where the data ISN'T.
CAltDataPanel::Features() returned an all-ZERO vector for any bar older than the
file's first row, and left blank cells at 0 too. Both were deliberate ('the block
is additive context and must degrade, never reject the bar') and that reasoning
holds for the CHANGE columns - but half these features are LEVELS: vix, ivol,
mac_y10, mac_cpi, mac_unemp, eia_util. For a level, 0 is not a missing reading,
it is an impossible one far outside the series' range. VIX does not visit zero.
And the spike lands in exactly the wrong place. Every alt file starts 2010-01-01
while the charts run far deeper - USDJPY H4 reaches ~1994, roughly HALF its
history - so 'alt block is all zeros' is precisely the predicate 'this bar is
older than 2010'. The IS/OOS split is chronological, so that predicate covers
~half of IS and none of OOS: an in-sample feature guaranteed to be useless
out-of-sample, and a bimodal input for the first BatchNorm to normalise. Not a
lookahead leak - a distribution corruption, which is quieter and was never
reported anywhere.
Now filled with the column MEDIAN over the covered range. A constant cannot leak
whatever its source - it takes the same value on every pre-coverage bar, so it
carries no information about which of those bars won - which is what makes a
median computed over later data legitimate here. Median not mean because the
series are skewed. Blank cells get the same treatment (eia_stk_idx1y alone has
181 blanks in 6,073 rows) and the count is now logged at load.
THE BACKOFF WAS ALREADY THERE AND WAS DEAD. Training.mqh arms m_coldSweepTick on
m_featureFailTransient, but only the open/ATR guards ever set that flag, so
f0cf659's cold ADMovingAverage looked PERMANENT and the sweep re-ran at full
speed forever. Setting the flag in the indicator guards revives the mechanism
that was already designed for this; no second backoff was needed and the one I
first wrote has been removed in favour of it.
SELF-HEALING, as asked. ReportFeatureHealth() runs once, the first time pass 1
produces usable windows, samples ~400 bars spread across the whole training range
and names every feature slot that is CONSTANT or mostly-zero, tagging alt-block
slots as alt[i]. Both of today's failures were the same shape - a block silently
produces nothing while every downstream number stays plausible - and neither an
accuracy figure nor a model can tell 'this feature is always 0' from 'this
feature is genuinely 0 here'. Evenly spaced sampling so a block that dies only in
deep history is caught as surely as one dead everywhere. A report, not a gate:
a rare-flag feature can be legitimately constant, and refusing to train would
turn a diagnostic into an outage.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
f0cf659945 |
fix(features): a cold indicator is TRANSIENT, not a permanent miss - and back off instead of re-sweeping
Six fresh instances on USDJPY and XAUUSD swept 33,965-50,162 bars and produced
ZERO usable feature windows, repeatedly, for 40 minutes and 239 stall reports,
without ever completing era 0. The four instances already warmed up before those
charts were attached trained normally throughout.
THE STALL REPORT NAMED THE SPOT EXACTLY: 'lookback slot 0 REJECTED the bar
(window had 24 of 832 values)', and 24 is the core block to the value - 4 price +
5 swing + 4 range + 4 volume + 6 time + 1 ATR. So feature 25 was the wall, and
feature 25 is the first value of the MA block. The same 24 appeared on XAUUSD
against a 816-value window (51 features/bar vs 52), which is what ruled out any
symbol-specific data gap: the wall sits at a fixed feature index, not a date.
ADMovingAverage is a CUSTOM indicator, so MT5 fills its buffer asynchronously and
returns EMPTY_VALUE for EVERY index until it has calculated - not just the
warm-up tail. That guard did not set m_featureFailTransient, so every bar of the
sweep was cached as a PERMANENT miss. This is precisely the failure the ATR guard
twenty lines above it was fixed for on 2026-08-10; the fix was never propagated
to the indicator blocks that follow. RSI, MACD and Ichimoku had the same defect
and are fixed too. (The Donchian high/low guard is a break into a
degraded-but-usable path, not a rejection, and is deliberately left alone.)
IT ALSO SELF-SUSTAINED, which is why it never recovered. The ok=0 self-heal drops
the feature cache and re-sweeps immediately, so each stuck instance spent every
millisecond re-reading 30-50k bars - six of them at once, on a six-core box,
competing for CPU with the very indicator calculation they were all waiting on.
The recovery was preventing the recovery. A transient total failure now re-arms
m_warmupPassesRemaining, yielding the CPU for a few separately-scheduled Train()
calls - the same mechanism a fresh model already uses to let history sync finish,
pointed at indicator warm-up instead.
Verified in the terminal journal first: indicators load and unload in matched
counts and there is no OOM, so this is NOT the
|
||
|
|
d87f7d88ff |
feat(gate): cross-instrument pooled certification
The deploy bottleneck is CERTIFICATION, not training. A 4,738-bar OOS window at L=75.6 holds ~63 independent observations; certifying a 3pp edge at 2 sigma needs ~1,036. More bars of the same symbol barely help - they overlap. Other symbols do not. WHAT POOLS. Not win rates: symbols have different derived geometries, different break-evens and different drifts, so averaging raw rates across them is meaningless. What pools is each symbol's EXCESS OVER ITS OWN CHANCE RATE, combined by inverse-variance weighting (fixed-effects meta-analysis). Each symbol keeps its own model, geometry and chance rate; only the evidence is combined. THE CORRELATION PROBLEM, bracketed rather than assumed away. SP500 and NAS100 are ~0.9 correlated and pooling them as independent inflates the evidence. Nothing here can measure that without sharing return series, so instead of guessing a correction the gate reports both ends: SE_INDEP = sqrt(1/SUM(1/var_i)) all members independent SE_CORR = SUM(w_i * sqrt(var_i)) all members perfectly correlated The truth is always between. THE GATE USES SE_CORR, so a pass cannot be an artifact of correlated instruments - that bound already assumes the worst. The ratio is logged as the diversification credit the gate declines to claim, so the cost of that conservatism is visible instead of hidden. SCOPE, deliberately limited: the pooled result is REPORTED, never folded into tradeableOK. The local gate certifies the model that actually trades this symbol; the pool answers the different question of whether the strategy has an edge at all. Letting a cross-symbol result license a local deploy would ship a model that never cleared its own bar - so it cannot. Mechanics: one file per instrument (no concurrent-write path to get wrong), every FileOpen carrying FILE_SHARE_READ|FILE_SHARE_WRITE, records skipped rather than reinterpreted on a version mismatch, 12h staleness cutoff so a stopped chart cannot vote, and pooling refused below 3 instruments. Poolability requires matching timeframe and ratio; differing SYMBOL is the entire point. Publishing is unconditional - a pool that only hears from winners is a selection effect, not a meta-analysis. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
2b5d0f8355 |
feat(measurement): fix zero-skill denominator, publish the deploy bar, measure lifespan per rung, add a MEASURE scale objective
The last run could not have demonstrated an edge either way, and nothing in the
log said so. Four changes so it does.
1. THE ZERO-SKILL LINE DIVIDED BY THE WRONG DENOMINATOR. m_oosWinLongTotal resets
every era; m_oosSamples only resets on a full model reset. So 'always-long %'
decayed as ~1/era: a run whose true rate is 37% printed 1.2% at era 33 and
0.0% at era 2219. This is the SAME bug already found and fixed for
logBuyPredPct thirty lines above ('era-15 Buy:2% that was really ~30%'), left
in the one line whose whole job is to be the reference every other number is
read against. Correct at era 1, wrong everywhere after - including the '62%
zero-skill' figure in the 2026-08-16 notes. Now per-era, and always-short is
finally readable.
2. THE DEPLOY GATE STATES ITS OWN BAR. 'edge -1pp' era after era cannot separate
'short by a hair' from 'short by an amount no strategy could cover'. The era
line now prints the required win rate, the SE, the effective n and the
lifespan it was deflated by; above 100% it says UNREACHABLE. At 4,738 OOS bars
and L=75.6 there are ~63 independent observations, putting the bar near 66% at
typical coverage.
3. LIFESPAN MEASURED PER RUNG. The first-passage cache already stores touch ages
at every ladder level, so each candidate geometry's resolution time is
readable without training on it - L-vs-width becomes a measurement across the
whole ladder in ONE run rather than a second chart. Each rung reports L,
n_eff, min provable edge and min provable EV.
4. SCALE OBJECTIVE IS PHASE-AWARE, defaulting to MEASURE. Width and detectability
are opposed: labels overlap by L, L grows like m*k = width^2 at fixed ratio,
so min provable EV ~ width^2 while the cost saving from width is only linear.
Doubling width quadruples the smallest EV you can prove. DEPLOY (widest that
clears reachability) is right once an edge is known; MEASURE (narrowest that
keeps round-trip spread under BARRIER_MAX_COST_FRACTION_PCT) is right while it
still has to be shown. The direction does not depend on the exponent, and
item 3 makes the exponent checkable.
Fixed in review: m_lastRungLifespan is cleared on every LadderWinShare entry or a
rejected rung reports the previous rung's lifespan as its own; per-rung
detectability is labelled IS-sample based (the deriver may not see the holdout),
so absolute figures are optimistic while the ranking is unaffected.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
d1ac18ebdb |
fix(labels): correct EffectiveSampleSize clamp order, share the horizon ladder, retract a false justification
Self-review of |
||
|
|
1540ba8e64 |
fix(labels): overlapping-label sample correction + horizon cap on the scale ladder
Three defects, all surfaced by the 2026-08-17 SP500 H4 run that shipped
stop 4.86 / target 9.71 (width 14.57*ATR, horizon 384).
1. EVERY STANDARD ERROR ASSUMED INDEPENDENT SAMPLES. Triple-barrier labels
started one per bar overlap by the label's lifespan, so n calls are worth
~n/L independent observations (Lopez de Prado, AFML ch. 4 - sample
uniqueness). All three sqrt(p(1-p)/n) sites divided by the RAW count.
The tell: the operating point's null-of-the-maximum gate is family-wise and
should fire on ~5% of eras under the null. Measured fire rates - PAI 47/73
(64%), ConvLSTM 9/24, LSTM 8/21 (38%), CONV 4/62 (6%). CONV, the only model
whose margin distribution admits few bins, sat on the null; the rest cleared
a bar that was too low by ~sqrt(L). PAI's deployed threshold consequently
alternated between the ENDS of its own range era to era (0.10 -> 0.88 ->
0.86 -> 0.66; coverage 16% <-> 73%).
TripleBarrierLabel now records when each label became KNOWABLE - the first
winning touch, or both stops, or the timeout - and the prebuild accumulates
the mean. EffectiveSampleSize() feeds the operating point, the member deploy
gate and the ensemble vote gate. Conservative by construction (n/L is an
upper bound on the damage); gates get harder, never easier.
2. THE SCALE LADDER RAN AWAY, again. Horizon scales as swingMedian*sl*tp, and
since
|
||
|
|
4d8cb08501 |
fix(geometry): the reachability floor measured the WRONG WINDOW - my bug from bc57aca, and it cost real width
RECONCILED: the derivation reported "target reached on 17.7% of bars" while the
label cache reported Buy on 35.9%. Nothing was broken. They measure different
windows, and both are correct:
EXCURSION window ~12 bars (the SWING MEDIAN) - what m_excUpCache accumulates
over. Deliberately short: sizing a barrier off travel
measured over a horizon that itself scales with the barrier
is circular, and it ran away to 14-31*ATR on EURUSD/USDCAD
in 2026-08-07. That guard is correct and stays.
BARRIER horizon 64 bars - what the LABEL walk and the first-passage ladder
run over, and how long the EA actually holds the trade.
So `up >= target` is a 12-bar question and `label == Buy` is a 64-bar one, and
the second can freely exceed the first. TripleBarrierLabel gates the excursion
accumulation on `idx - t <= excWindow` while the barrier walk and the ladder run
the full horizon - the split is explicit and intentional.
THE BUG IS MINE. bc57aca's scale ladder tested reachability with `up[i] >= tp`,
i.e. it asked the 12-bar question about a 64-bar trade. That understates
reachability by ~2x, which is why EVERY wide rung was rejected and the geometry
fell back to the tightest rung at 1.61/3.21. The data supported considerably
wider; the test was just asking the wrong question.
FIX: LadderWinShare() reads the answer off the first-passage ladder - target
touched strictly before the stop, over the full horizon, tie to the stop. That
is the identical question the label walk asks, so the ladder share and the Buy
rate should now agree to within rung discretisation. Both legs snap to the
SMALLEST rung at or above the requested multiple (harder target, harder stop) so
the floor stays conservative.
Expect the scale ladder to select a WIDER rung on the next relabel. On this
data the excursion test read 17.7% at q50 where the true full-horizon share is
35.9%, so rungs that scored 8.1% and 2.8% were likely well above the floor.
ALSO:
- Window reconciliation now PRINTED every derivation: excursion travel share,
ladder win share, and the label cache's Buy share side by side, with the
ladder-vs-label gap flagged if it exceeds rung discretisation. Those two must
agree; if they ever stop agreeing, one of them is wrong and the line says so.
- Renamed tpReach/slReach -> tpTravel/slTravel and relabelled the log line. They
describe the EXCURSION window and are near-tautological there (a q50 stop is
exceeded by ~50% of bars); calling them "reached within the horizon" is what
made the two quantities look like one.
- BARRIER_MIN_TP_REACH_PCT is now BARRIER_MIN_REACH_FRACTION_OF_BE (0.60) x
break-even instead of a hardcoded 20.0. Break-even for 1:RR is 100/(1+RR), so
the absolute floor silently tightened as RR rose - 0.60x at RR=2 but 0.80x at
RR=3, penalising the user for asking for a bigger target. Evaluates to exactly
20.0% at the shipped RR=2, so this is a no-op today and correct if the knob
moves.
NOT COMPILED - user compiles.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
bc57aca15d |
fix(geometry): the target was small BY CONSTRUCTION - ratio is now policy, scale is measured, ladder ceiling removed
The derivation read the stop from q75 of ADVERSE travel and the target from q50
of FAVOURABLE travel. Over one horizon those distributions are broadly the same
shape, so q75 > q50 MECHANICALLY - the target came out smaller than the stop no
matter what the market did. SP500 H4 shipped stop 3.07 / target 1.70: a 0.55:1
payoff needing 64.3%. That was never a measurement, it was two mismatched
constants.
The reachability line printed beside it - "target on 50.0% of bars, stop on
25.0%" - is exactly 1-q50 and 1-q75. Tautological. It cannot disconfirm
anything, and it read as validation.
WIDTH AND RATIO ARE INDEPENDENT AND ONLY ONE PAYS. EV = edge x width;
ratio is EV-neutral (a driftless walk reaches +m before -k with probability
k/(k+m), which IS break-even). Width is what buys cost efficiency: the spread
is a fixed 0.047*ATR here, so the shipped 4.77*ATR width paid it 21 times per
unit of travel. So:
RATIO = policy. BARRIER_TARGET_RR = 2.0 (user's 1:2). Break-even 33.3%.
SCALE = measured. The stop quantile is chosen from a ladder, WIDEST FIRST,
taking the first rung whose implied 2x target is still reached often
enough to be a trainable class.
That last clause is the difference from the min-reward:risk raise removed in
2026-08-09, which forced target = 2 x stop with NO reachability test, landed on
6.66*ATR reachable on 3.3% of bars, and trained the model to predict something
that essentially never happened. Same ratio; the scale now retreats until the
data says the target is attainable. Every rung is logged.
LADDER CEILING REMOVED. BARRIER_LADDER stopped at 5.00 and the expectancy scan's
"best resolvable pair on width alone" came back as stop 5.05 / target 4.95 - it
pinned to the top rung. A recommendation landing exactly on the edge of its own
search space is a boundary, not a finding: it cannot tell "5 ATR is optimal"
from "5 ATR is all we allowed". Extended to 20*ATR (8 -> 14 rungs). Nothing else
needs editing - every consumer is parameterised by BARRIER_LADDER_COUNT - and
the horizon constraints (decided >= 60%, reachability floor) now bind instead of
a constant.
THE SCAN COULD NOT SEE THE SHIPPED GEOMETRY. ReportBarrierGeometryScan looked
the configured pair up in its integer grid, and DeriveBarrierGeometry produces
CONTINUOUS multiples (3.07/1.70) that can never equal a grid point - so
cfgExcess stayed at its -1.0 sentinel and the report printed "configured 3:2
scores -1.00000", which reads as a catastrophic score and actually means "never
evaluated". Worse, the grid skipped target<stop entirely because it "inverts the
trade's whole premise" - while the derivation was shipping exactly that. The
incumbent is now always scored as a peer (never crowned; it is already in force
and is not an enum pairing the scan could adopt).
BREAK-EVEN NOW INCLUDES THE SPREAD. Every report quoted the frictionless
SL/(SL+TP). On SP500 H4 that read 64.3% while the MEASURED zero-skill rate was
62.1% - a 2.2pp gap that IS the cost, and that made every model look 2.2pp
better than it was. CostAdjustedBreakEvenPct() prices a win at (TP - spread) and
a loss at (SL + spread), matching the expectancy scan's convention exactly so
the two reports cannot disagree.
It also feeds FitDirConfThreshold, which is the correctness half: the operating
point subtracts break-even from precision, so the frictionless figure made every
candidate threshold look better by the width of the spread - 2.2pp against a
measured edge of 2.3pp, i.e. very nearly all of it.
Era line now carries both: "break-even 64.3% frictionless, 66.6% AFTER SPREAD".
Forces a full relabel and retrain. Requested.
NOT COMPILED - user compiles.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
d30420e3f2 |
fix(batchnorm): bound the normalized value - a constant input feature was amplified 1e4x and pinned PAI's head to its rails
BN_MIN_STD = 1e-4 caps the per-unit gain at 1/1e-4 = 1e4, and the comment above
it states that as though it were a safety property. It is not. A unit whose
running variance is ~0 is a CONSTANT feature carrying no information, and
dividing its rounding noise by 1e-4 hands the next layer an activation of
several hundred. BN's contract is "output has ~unit variance"; a unit that
cannot supply that must contribute nothing, not the largest signal in the layer.
MEASURED, 2026-08-17 SP500 H4, four topologies on identical separate charts:
model spread Neutral CHOSE Neutral TIED rail
CONV 0.386 0.68% 0.10% 0.48%
LSTM 0.392 0.63% 0.00% 0.00%
HYB 0.376 1.79% 0.00% 0.01%
PAI 0.192 0.09% 80.63% 99.99%
bn1's cached nx normed 1.38e4 over 800 units. PAI's SIGMOID head was on its
rails on 99.99% of bars, with Buy and Sell landing on the SAME rail so they
compared exactly equal, and ApplyClassificationSoftmax()'s strict-majority rule
reported that tie as Neutral on ~80% of bars.
So the long-running "PAI is heavily biased toward Neutral" was never a
class-prior problem: the net CHOSE Neutral on 0.09% of bars. It was float
equality on a saturated head. The
|
||
|
|
331ab29c56 |
feat(diagnostics): split a reported "Neutral" into CHOSE vs TIED - they need opposite fixes
ApplyClassificationSoftmax() requires a STRICT majority over both rivals and
sends every tie, 2-way or 3-way, to Neutral. So "OOS recall Neutral:100%" is
two completely different events sharing one label:
CHOSE - the net genuinely ranks Neutral highest. A class-prior/label problem.
TIED - the top two are EXACTLY equal, so the net expressed no preference and
the tie-break reported Neutral. A SATURATION problem: the head is
SIGMOID, and a saturated sigmoid returns exactly 0.0f or 1.0f in the
DLL's float32, so two classes pinned to the same rail compare equal
and the bar is silently discarded.
Nothing in the logs could tell them apart, and the fixes point opposite ways.
Eras 1-25 of the 2026-08-17 solo PAI run read "Neutral 100%" at spread avg 0.99
- fully saturated - and broke out at era 27 as the spread fell to 0.75. That is
consistent with EITHER story. The user reports the Neutral phase on most runs,
so it is worth four longs to stop guessing.
Four per-era counters on the pass 3 OOS walk, reported as:
| Neutral CHOSE 12.4% / TIED 38.1% (of which B=S 1204) | rail 61.2%
m_oosNeutralStrict - Neutral strictly highest
m_oosNeutralTie - no strict winner; the tie-break produced Neutral
m_oosTieBuySell - the costly subset: Buy and Sell tied AT the top, i.e. a
DIRECTIONAL reading thrown away by float equality
m_oosRailBars - any raw output sitting on a sigmoid asymptote, the
saturation that makes exact ties possible at all
Read on the RAW logits, before ApplyClassificationSoftmax() overwrites TempData
in place. Legitimate because softmax is strictly monotone: it cannot change the
ordering and cannot break a tie either, so the raw reading and the decision
always agree. Placed alongside the existing min/max/spread capture so all the
output diagnostics describe the same values.
Measurement only - no decision path reads these.
NOT COMPILED - user compiles.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
17f808e90d |
fix(ensemble): MAX_AI_SIGNALS was 3 - the ensemble creates 4, so CONVLSTM was silently dropped
AI_HYBRID enables PAI + CONV + LSTM + CONVLSTM and RegisterAISignal registers
them in exactly that order. MAX_AI_SIGNALS was 3, and the guard returned
silently, so the FOURTH - CONVLSTM - never entered g_aiSignals[].
Reported as "convlstm is not listening to the control panel buttons", which is
the visible tip. Everything in Warrior_EA.mq5 that reaches a model does so by
looping g_aiSignals[], so the dropped member also lost:
- every control panel button (pause/resume, stop/start, retrain, deploy,
save, load, reset weights)
- PollTraining() in OnTimer - no wall-clock training progress, so it only
advanced on ticks
- AutosaveWeightsIfDue() -> SaveWeightsNow()
- AltDataReload() on both the mapping-dialog and hourly-upkeep paths
- StartChartSignalRescan()/RescanPending() - the Show Signals sequence
- the All*/Any* aggregates (deployed/paused/stopped/complete), which were
therefore computed over 3 of 4 members and could report the ensemble
finished while CONVLSTM was still training
- OnDeinit's MarkShutdown(), ShutdownChartCleanup() and FlushTrainRun() -
so its arrows were stranded on the chart and its training run was never
flushed on shutdown
It stayed hidden because the model still trains and still votes: it lives in
the signal's own filter array, and it registers itself with the status panel
(ENSEMBLE_PANEL_MAX_MEMBERS is 6) rather than through g_aiSignals[]. So it
appeared on the panel, drew arrows and moved the vote while being unreachable
from every action and unsaveable on exit.
MAX_AI_SIGNALS 3 -> 5 (4 is today's true maximum; the spare slot means adding
META to a preset cannot reintroduce this - the array holds borrowed pointers,
so unused slots cost nothing).
RegisterAISignal now PRINTS on overflow instead of returning silently. A cap
that discards a model without saying so is a trapdoor, not a guard.
NOT COMPILED - user compiles.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
34fa0045e3 | feat: extend MAX_ERAS_PRESET enum and update MaxErasPerRun to ME_10000 | ||
|
|
7414570d9d |
fix(diagnostics+calibration): the frozen-layer reading was a broken ruler; gate the operating point on a null of the maximum
Two defects behind the "training is highly unstable" report, from 101 eras of
SP500 H4 PAI logs. Neither was the optimizer.
1) LayerLearningReport's dW/W for BN layers divided by the WHOLE packed block.
getWeightsBN concatenates the outgoing dense matrix, gamma/beta, the running
mean/variance, the Adam moments AND BN_OPT_NX - the forward-pass scratch copy
of the normalized input. At era 101 bn1's dense matrix normed 15.1 against a
block norm of 15430.3, of which NX alone was 15429.3: the weights were 0.098%
of their own denominator, a 1022x inflation. NX is also near-constant between
era-end reports (same last forward pass), which pins the numerator down too,
so the layer read "bn1:0.000%" for 101 consecutive eras and was diagnosed as a
frozen first layer. It was the ruler that was broken. The ratio now covers
trainable parameters only (dense matrix + gamma + beta); mean/var/NX/Adam are
excluded. NX is reported separately because it is a health signal in its own
right - bn5 read nx 6.8e6 over 16 neurons, ~1.7e6 per unit against a healthy
~1.0, which is what a near-zero running variance in the denominator looks like.
NO historical dW/W reading on a bn* layer is admissible evidence that a layer
did or did not train. That includes every such claim in this repo's notes.
2) FitDirConfThreshold took a bare argmax of coverage x (precision - breakEven)
over 50 bins. Measured across 98 consecutive fits:
correlation(chosen threshold, win rate at it) = -0.056 over 0.00..0.74
win rate stdev across fits = 1.32pp
binomial SE of that win rate at ~1430 calls = 1.25pp
The correlation is zero - the margin does not rank trades - and the era-to-era
spread IS its own sampling error to within 0.07pp. So the objective was
coverage x (3.4 +/- 1.3) and the argmax over ~37 eligible bins returned
whichever bin drew the luckiest sample. The threshold teleported
0.42 -> 0.04 -> 0.74 in three eras, swinging OOS coverage 0% -> 39%, leaving
the era win rate measured on 1-5 calls and swinging 0% <-> 100%. That is the
entire reported instability.
The argmax is now adopted only if it beats a DETERMINISTIC fallback - the most
selective bin still clearing the coverage floor, chosen from the margin
distribution alone and never from a win rate - by more than a best-of-N
maximum could manage on noise, sqrt(2 ln N) standard errors. Same null-of-the-
maximum correction the deploy gate already applies to model selection.
A plain one-standard-error band was tried first and is NOT sufficient: its
edge is bestScore - bestSE, and with a 2.3pp edge against a 1.25pp SE that
edge is itself +/-50%, so the admitted set would still wander by half its own
width every era. The fallback has to be independent of the noisy quantity.
Simulated on the observed numbers: falls back every era at the current 2.3pp
edge (stable), adopts the argmax once a real edge reaches ~5pp.
NOT COMPILED - user compiles.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
208da4cbaa |
fix: drop the ranking slice for the calibration band; un-collapse the tiers
NOT COMPILED - user compiles. (1) THE RANKING SLICE IS GONE. It reserved 20% of the OOS window so the pattern-DB backfill would read bars the deployed checkpoint was not SELECTED on. That objection stands; carving a new region to answer it did not. The calibration band already has every property the slice was buying: never trained on | never graded by pass 3 (which walks [0, oosCutoff) and so never reaches it) | never seen by the deploy gate | purged by a full label horizon on BOTH sides | and larger besides - 1,684 bars vs the ~970 carved So the backfill now walks [calibLo, calibHi) and pass 3 goes back to grading the entire OOS window, exactly as before any of this. The gate gets its full sample back (~10% of a sigma), the split loses a region, and the failure mode found an hour ago - a reserved region silently blanking ~10 months of chart arrows, because arrows are only drawn on bars pass 3 grades - becomes impossible. One impurity, stated in the completion log rather than hidden: m_dirConfThreshold is FITTED on that band and the walk applies it to decide which bars fired, so coverage there is mildly optimistic. One scalar under a coverage floor, against checkpoint selection over hundreds of eras. This backfill IS the deploy-time warm-up: it runs right after FinalizeTrainRun() restores the deployed weights, so it scores with exactly what is about to trade. (2) EVERY CALL WAS TIER 0, AND IT WAS ARITHMETIC. ConfidenceTier() quartiles [floorConf, 1] where floorConf = 1/3 - the lowest magnitude a 3-way softmax winner can hold. But it was fed CalibratedConfidenceMagnitude(), which multiplies by m_confidenceCalScale, clamped to [0.3, 1.5]. That lower clamp is BELOW 1/3. Whenever calibration bottoms out, t goes negative and MathMax(0, ...) pins every call to tier 0. Which is what the live run does. m_confidenceCalScale is EMA'd toward empiricalAccuracy / avgClaimedConfidence; with the model over-calling Neutral, 3-class agreement sits near 10% against a claimed confidence near 0.9, so the ratio is ~0.11 and clamps to 0.3 every era. Logged: tier prec T0:72%(828) T1:n/a(0) T2:n/a(0) T3:n/a(0) 828 calls, one bucket - the four tier weights and the entire per-tier pattern-DB ranking reduced to a single number. The backfill was feeding a mechanism that structurally could not rank. Tiering now reads the RAW head magnitude, which genuinely lives on the [1/3, 1] range these bounds were written for. Calibration keeps its real jobs - AIConfidence() for MM sizing and SignedAIConfidence() for the vote are unchanged. STILL OPEN, deliberately not touched here: the calibration TARGET itself. empiricalAccuracy is 3-class agreement, which is the wrong quantity to scale a DIRECTIONAL confidence against - it counts a Neutral class that is 0.19% of labels. The honest target is the win rate on the calls the confidence describes (directional precision), with the claimed-confidence average taken over those same called bars. That needs a new accumulator and it interacts with the Neutral over-calling being fixed elsewhere, so it wants one clean run first. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
75d23e9b82 |
fix(gate): move the ranking slice to the OLD end - it walled off the recent chart
NOT COMPILED - user compiles. User: "there is quite some trading going on, but absolutely nothing on the recent area of the chart, like there is a hard wall starting around november 2025." That wall is 7caf2f6's ranking slice, and it was placed at the wrong end. Chart arrows are only ever drawn on bars pass 3 GRADES, and the slice reserved the NEWEST 20% of the OOS window plus a label-horizon purge. At the live sizing - ~4,860 OOS bars, 128-bar horizon - that is ~1,100 H4 bars withheld from grading, about ten months back from today, exactly where the wall appears. The invisible cost was worse than the visible one: it handed the deploy gate the OLDEST 80% of the OOS window and withheld the most recent regime from the single decision that has to generalise forward. Both fixed by putting the reserve at the oldest end instead: [0, oosScoreHi) OOS - graded by pass 3 (NEWEST, arrows restored) [oosScoreHi, rankLo) purge - one label horizon [rankLo, oosCutoff) RANKING - backfill only, graded by nobody [oosCutoff, calibLo) purge [calibLo, calibHi) CALIBRATION ... IS Of the three consumers competing for those bars, recency is worth least to the ranking: it is an ORDERING of confidence tiers, far less regime-sensitive than an absolute win rate, while the gate's power and the operator's read of the chart both want the newest data. The slice keeps every property that made it worth carving - never graded, never selected on, never seen by the gate, purged on both sides - so the backfilled rows are still honestly out-of-sample. RankSliceHiIndex is replaced by RankSliceLoIndex + OosScoreHiIndex; pass 3 now excludes the slice at the TOP of its walk and descends to 2 as it always did. The backfill walks [RankSliceLoIndex, oosCutoff) via a new m_dbBackfillStopIndex, clamped at both ends so a degenerate slice yields an empty walk rather than one that wanders into graded bars. Verified no reference to the old helper survives. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
c0c9f4a285 |
feat(target): withdraw the TrainingTarget option - barrier is the only live one
NOT COMPILED - user compiles.
The private build still DEFAULTED to TARGET_FRACTAL, so every fresh attach was
training the target adjudicated dead that morning (5,700 model-eras flat at -2pp,
best-of-243 p=0.17). The campaign closed; the default was never flipped back.
Rather than re-default it, the input is withdrawn entirely (user: "remove the
option if there is only one choice for now"). An input offering a single live
choice is worse than no input - it presents a dead option as supported, and an
operator picking it silently trains a model already known to carry nothing.
Direction models are now unconditionally triple-barrier.
Removed: the input, the TrainTargetFractal() call in the signal setup, and the
HoldToBarrier() exit-policy block (which existed only because the fractal vote
flips at swing-marker cadence, ~3-5 bars, far inside the barrier's travel time -
barrier-target models keep vote exits and always did, their label IS the vote's
horizon). Verified no code reference to TrainingTarget survives; the four
remaining mentions are comments.
Kept deliberately, so a rerun is a re-enable and not a rebuild: the TRAINING_TARGET
enum, the fractal label itself, its |TGT:FRA1 fingerprint token, its conditional
barrier-geometry derivation, HoldToBarrier()/m_holdToBarrier, and the campaign's
trained models on disk. Three lines bring it back; Inputs.mqh names them.
ALSO CORRECTS THE RECORD from
|
||
|
|
1b5a412946 |
fix(imbalance): the class-imbalance correction was subsidising the abstain class
NOT COMPILED - user compiles.
Root cause of the Neutral collapse. Logit adjustment (Menon et al. 2020) makes a
classifier Bayes-optimal for BALANCED error by subsidising rare classes. It was
wired here when Neutral was the DOMINANT class - the "big move up / big move down
/ nothing much" era, where the correction pulled the model off the majority.
The triple-barrier relabel (
|
||
|
|
1eeed3ac06 |
revert(ui): restore the unconditional era-end arrow repaint
|
||
|
|
b19799910d |
fix(ui): chart arrows follow the BEST checkpoint, not the latest era
User report: "as soon as the next era training begins the chart signals are
erased, they should persist for as long as they are accurate."
Cause: PruneDirectionalClusters ran unconditionally at the end of pass 3, and it
DELETEs the arrow on any bar the CURRENT era scored Neutral. A model exploring
away from its best therefore wipes the chart every era even though the best
checkpoint still calls those turns. On the run that prompted this the model sat
at OOS recall Neutral:100% for 40+ consecutive eras, so essentially every arrow
was deleted at every era boundary.
The render is now deferred to the era-end block - the first point that knows
whether the era beat the best checkpoint - and only a new best repaints. On eras
that did not improve, the previous best's arrows stay untouched. Two exceptions
keep the chart from ever showing nothing: before the first checkpoint exists
there is no best to preserve, so early eras still paint; and a finishing run
repaints unconditionally, because FinalizeTrainRun is about to restore the
deployed weights and the chart must describe THOSE.
Recorded for ensemble members too. A member's own best era is not the deployable
one (the joint checkpoint decides that), but it is still the most accurate thing
that member has drawn, and the alternative is a chart that empties itself.
Also verified against the log, since two other symptoms were reported alongside:
era cadence PAI-b6b5 (before these changes) 3.76 s/era
PAI-17ae (after) 3.53 s/era
topology both "2 dense from 16 units | input 800 (16 bars x 50)"
calibration both fitted on 1699 held-out bars
So training speed is unchanged - the ~6% is the ranking slice removing 20% of
pass 3's bars. It only FEELS fast because this is a single PAI chart taking the
whole 120 ms budget, not four ensemble members sharing it behind an era barrier.
The Neutral collapse is also pre-existing, not new: b6b5 ran at Neutral 94-100%
with 1-5% directional calls for all 723 of its eras, before any of this work.
That is the known neutral-collapse/recall-gate failure mode, and it is what
"barely drawing signals" actually is. Worth watching, separately: b6b5 reached
best-bal 34.2% by era 723 while 17ae is at 13.6% after 77 - too early to read,
but it is the number to check once 17ae has run comparable eras.
Compile-verified: 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
ee3682d949 |
fix(features): collapse only the anchor's own run - leave lagged readings put
User's call before deploy: "I would rather avoid lagging so the NN finds
accurate patterns." Correct instinct, and it picks the conservative variant.
|
||
|
|
110b38470a |
fix(features): dedup the alt block BY VALUE - aba9bd2 broke D1 charts
|
||
|
|
aba9bd2bea |
perf(features): the external block enters the window once, not once per bar
Measured on the live SP500 D1 export (6073 rows, 13 features, 5888 simulated 16-bar windows): distinct values per feature per window : 1.7 - 2.7 of 16 slots variance in the first 13 PCs : 96.5 - 97.0% components for 95% / 99% : 12 / 17-19 effective rank (entropy) : ~11.5 208 inputs carrying about 12 dimensions. Only 6 of the 13 features move daily (VIX complex, USD, the rates trio); 5 are weekly (COT, EIA, output gap) and 2 monthly (CPI, unemployment). The lookup is as-of by bar open time into a DAILY file, so bars sharing a calendar day are byte-identical by construction. The cost is NOT overfitting capacity - collinear copies span ~12 directions, not 208, so an earlier claim that this wasted 26% of the model overstated it. It is GRADIENT WEIGHTING. Batch norm standardizes each of the 208 coordinates independently; that rescales the copies without decorrelating them, so one factor arrives on 16 unit-variance coordinates, each weight takes a full-size step, and the factor's aggregate coefficient moves ~16x faster than a per-bar price feature's. The network was biased toward the external block by a factor of the window length - and pointing the wrong way, since these features cleared only a marginal incremental screen while price is the base signal. Zeroed at WINDOW ASSEMBLY, not in BufferTempData: that output is cached PER BAR and a bar sits at slot 15 of one window and slot 0 of the next, so a slot-dependent value there would poison the cache or force a recompute per slot. The cache keeps true values; only this window's copies are cleared. Width contract untouched - same count, same positions - so conv/LSTM/HYBRID keep their bar-major rectangle unchanged and the block arrives at the newest bar, which for the LSTM is the final timestep. Zero-variance coordinates are safe through batch norm (divisor is MathMax(MathSqrt(var + BN_EPSILON), BN_MIN_STD)). Fingerprint gains |ALTW:1 when alt data is on. Same width and same .cfg, so nothing else would have caught a model trained under the replicated layout resuming under this one. Conditional append per the existing rule: configs without alt data keep their fingerprints and their trained models. NOT the concat branch. CNet is a strictly linear stack (CLayerDescription has no input-source field; NetBuild wires i to i+1 and stores layer L's weights on L-1), so a real two-tower model needs a new multi-input layer type across WarriorCPU, WarriorDML and the OpenCL kernels plus an .nnw format change - the highest-risk change in this repo, in the code that produced the transposed dense gradient, the Adam second-moment bug and the reversed LSTM window. This captures the part of that idea the measurement actually supports, at no engine risk. Compile-verified: 0 errors, 0 warnings. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
7caf2f626e |
feat: derived taper restored; DB ranking reads a reserved slice, shrunk
TOPOLOGY - reverts the two constants and drops CausalHiddenLayerFloor. The MQL5 article's 30%-per-layer cut and floor of 20 are load-bearing on ITS first-layer width of 1000 (1000->300->90->27 needs a floor to stop). This codebase MEASURES that width, and on the live SP500 H4 config it is 16 units - already floored, with the budget printing "11360 estimated in-sample bars cannot support a 800-wide input ... roughly 1.1 weights per training bar - expect overfitting". At 16 units a floor of 20 makes lastHidden >= m_initialNeuronsCount, so ComputeHiddenLayerCount returns on its first branch and the width taper - the only part derived from this symbol's data - became dead code on all four ensemble members, with depth (2 -> 4) set entirely by counting feature domains. ComputeLayerWidths had already rejected this exact pair of constants in its own comment. The causal floor's premise does not hold either: layers are not inference steps. The "1 layer linear / 2 nonlinear / 3 multi-connected" result is Lippmann 1987 and is about hard-threshold units; with sigmoid/ReLU, Cybenko 1989 and Hornik 1991 give universal approximation from a single hidden layer. Depth buys parameter efficiency for compositional functions, not reasoning hops. ForceHiddenLayers remains for measuring depth directly. RANKING SLICE - the backfill no longer reads the window it is judged on. The deployed checkpoint is CHOSEN as the best-scoring era on the OOS window, so win rates measured back over it are selection-inflated, and the backfill was writing exactly those into the table filter weights rank on: the selection set consumed twice, beside a deploy gate that applies a Sidak correction for that effect. The newest RANK_SLICE_PCT_OF_OOS (20%) of the OOS window, plus a label-horizon purge, is now reserved and graded by nothing - not pass 3, not checkpoint selection, not the gate. The backfill reads only that. The gate keeps ~80% of its measurement (power goes as the square root, so ~10% of a sigma), and the slice is the newest data, which is the regime about to be traded. RankSliceBars returns 0 when no honest slice fits and the backfill then REFUSES and says so, rather than falling back to the scoring window and looking like a success. SHRINKAGE - per-tier win rates are shrunk toward the filter's own pooled rate by MIN_TRADES_FOR_WIN_RATE pseudo-trades before becoming weights. The raw ratio at the minimum sample count carries a ~15pp standard error, so a tier that went 8-2 was handed weight 80 and outranked a tier measured over hundreds of calls at 55 - the ranking was being driven by which small tier got lucky. Opt-in per call site (priorWeight 0 keeps the raw behaviour). Compile-verified: 0 errors, 0 warnings. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
b6736fd40b |
fix: the DB backfill could never run, and HEAD did not compile
Four defects in 64c5dd5/1a05e63, found by review + a baseline compile. Goals 1-8 of that session are unchanged; this makes 6 and 8 actually reachable. 1. HEAD DID NOT COMPILE - 6 errors. CControlPanel::Minimize/Maximize were declared `virtual bool ... override`, but CAppDialog declares both as `virtual void` (Controls\Dialog.mqh). errors 265 + 404 on each, plus 151 on `bool ok = CAppDialog::Minimize()`. Return type is void now; there was never a success flag to forward. Verified: 0 errors, 0 warnings. 2. THE BACKFILL COULD NEVER ADVANCE, and neither could the OOS continual simulation (that one has been dead since it was written). Both are armed at the instant convergence is declared, and both advance only from inside Train(), one chunk per call. But ScheduleTrainingIfNeeded's only per-tick ArmStudyEvent site sits in the `else` of a branch taken whenever m_trainingComplete is set and m_trainRunActive is clear - which is exactly the state FinalizeTrainRun() leaves behind one line before they are armed. Train() was never called again, so the walks sat at their start index forever: no "simulation complete" line, and not one row written to the DB this feature exists to fill. Only a manual Resume/Retrain unstuck them. Both flags now keep the model schedulable. 3. IN AI_HYBRID - the mode this ships in - the backfill was never even armed. Ensemble members deploy at Train() ENTRY and return immediately (so no era is wasted), which skips the era-end block the backfill was started from. All four members were a no-op for a second, independent reason. Armed on the ensemble deploy path too, from m_resumeBars/m_resumeOosCutoff. 4. RE-RUNS DUPLICATED ROWS. RegisterSignal inserts unconditionally - no key, no duplicate check - and m_dbBackfillDone is in-memory, so every later attach that retrained to convergence wrote a second full set of rows for the same bars. The ranking would count one bar once per model that ever deployed, weighting superseded opinions as heavily as the live one. A .dbfill marker stamps the deployed era; written only on completion (an interrupted walk redoes itself rather than ranking a partial window) and deleted with the other sidecars on reset-weights. Also: WarmBlocking's timeout was silent, which restored the exact silent pin failure it was added to prevent - it now says so in the journal, and returns true for "no reference pairs to wait for" so the warning stays rare enough to be read. Not addressed, needs a decision: the backfill scores the OOS window with the checkpoint that was SELECTED as best on that same window, then writes those win rates into the table filter weights rank on - the selection set consumed twice, undiscounted, while the deploy gate right next to it applies a family-wise correction for exactly that effect. The rows are also simulated triple-barrier outcomes at today's spread sharing a table with realised fills. The completion log line now states both plainly. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
64c5dd55d3 | feat: implement one-shot pattern-database backfill and enhance accuracy tracking for ensemble models | ||
|
|
1a05e632eb | fix(altdata): ensure alt-data is available before model initialization to prevent undersized models | ||
|
|
5a5be8999e | fix(altdata): add late warning for alt data arrival after model build | ||
|
|
e049b624ba |
feat(ensemble): deploy gate on the COMBINED VOTE, with a joint checkpoint
The unit of evaluation in ensemble mode becomes the vote, because the
vote is what trades (user: "at the end of the day they will vote
together during live trading so that would make sense").
Four decisions move from the member to the ensemble:
* which era is "best" -> the era whose COMBINED VOTE scored best
* what is checkpointed -> a JOINT snapshot: every member's weights
at that one era
* when the run gives up -> one shared plateau ladder
* whether it may deploy -> family-wise gate on the vote
WHY THE JOINT CHECKPOINT IS THE POINT: per-member selection picks each
net's own best era, and those eras differ. The resulting quartet was
never measured together at any instant, so the vote it casts live is a
configuration no OOS number ever described. Capturing all four at the
era whose vote won makes the deployed ensemble exactly the measured one.
Correct because of the era barrier (
|
||
|
|
b77e7b4766 |
fix(ensemble): responsive panel + synchronized eras + combined-vote accuracy
Four user-reported/requested items, one root cause chain:
1) DEAD CONTROL PANEL in AI_HYBRID mode. All members posted custom event
id 1 and handled id 1001, and CExpertCustom broadcasts every chart
event to every filter - so each posted event ran a train chunk in ALL
N members (N*N chunks per round) and the chart thread never idled
long enough to deliver clicks/drags. profiling.csv: 99.45% of time in
OnChartEventHandler. Fix: per-instance study-event ids
(STUDY_EVENT_ID_BASE + construction order, offset above the Controls
library's ON_* codes - id 1 was also ON_DBL_CLICK, so panel
double-clicks fired training chunks). ArmStudyEvent() is the single
post site; lost-event watchdog replaces the accidental
sibling-clears-my-flag rescue.
2) WARM-UP DUPLICATION. The auto-tune sweep is deterministic over
identical features/labels, and it ends in the full MI diagnostic
suite, which the MI-share gate never intercepted on the sweep path -
four members ran four identical ~36s sweep+report blocks. First
member publishes outcome (g_ensembleChartTuneDone/Installed/Settings);
the rest apply it and skip both.
3) DEINIT STRANDED PANEL+ARROWS (user repro 18:52). Root cause from the
log: the 4,500ms budget runs from MetaTrader's stop REQUEST - a heavy
autosave in flight ate it, OnDeinit got ~430ms and died in the first
member's arrow persist ("Abnormal termination" 432ms in). Fix: early
visible-UI sweep (native prefix deletes for status/panel/dialog)
right after ClearStatusLabel, and a fast path for still-training
models - their arrows are re-rendered every era, so they get one bulk
purge instead of scan+atomic-write in the death window.
4) ENSEMBLE FEATURES (user requests): era BARRIER - members advance era
by era together; a member ahead of the slowest still-training member
declines Train() calls and its chunk budget is donated
(TRAIN_TIME_BUDGET_MS = 120/activeTrainers, UI headroom constant).
COMBINED-VOTE OOS SCORE - each member's pass-3 scan contributes its
adjusted per-bar decision (0.0 on abstain) to a shared row buffer;
the last member to finish the era scores the averaged vote vs the
mirrored Min_Vote_Open against the same target-before-stop outcomes
members grade themselves on, publishing an "Ensemble vote" line on
the aggregated panel. Member headlines now carry their lifetime win
rate with break-even.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
65c4b1dce7 |
fix(ensemble): per-member arrow namespaces; ConvLSTM rename; dialog in purge list
The ensemble chart UI had a shared-namespace defect that answered the user question "what do the arrows represent?" with "a bug": all four members drew arrows under the same WarSig_<bartime> object names, so the chart showed whichever member rendered LAST, one member Neutral deleted another member Buy at the same bar, each member init sweep wiped the arrows the previous member had just restored, and SaveChartSignals - which rebuilds the sidecar by SCANNING the chart - persisted every other member arrows into its own history (the exact cross-model laundering its own header warns about, now happening BETWEEN ensemble members). Arrows are now namespaced per member (WarSig_PAI_, WarSig_CONV_, WarSig_LSTM_, WarSig_HYB_): draw, delete, restore, prune, member init sweep, destructor purge and the sidecar scan are all member-scoped, and the tooltip names the model. Global purges keep matching the bare WarSig_ prefix, which covers all member namespaces plus old-format leftovers from earlier builds. Labels: the ensemble panel header no longer says "HYBRID ensemble" (HYBRID is one member; the header is the ensemble) and the CONVLSTM member displays as ConvLSTM instead of Hybrid. Its SHORT id stays HYB deliberately - it names the model folder and changing it would orphan every model trained under that path. Deinit: the alt-data mapping dialog namespace (WarriorAltMap_) joins WarriorChartPrefixes, so both the OnInit purge and the deinit final sweep now cover it - it was in neither list, so a dialog starved of its own Destroy() left its controls on the chart permanently. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
ed15711e8c |
fix(altdata): first live fetch findings - key leak masked, EIA UA, bounded FRED backfill
Log review of the 18:12 attach. The wiring works: VIX, dollar index, COT, all seven macro series fetched and SP500_D1.csv rebuilt with its full 13 features on the first pass. Three findings from the same log, fixed: KEY LEAK: the EIA failure line echoed the first 80 chars of the URL, which included most of the api_key. Every URL-echoing error path now goes through MaskUrl(). The key itself is unchanged - it was printed to a local journal, not transmitted - but rotate it if that log ever leaves the machine. EIA HTTP 1003: an MT5 transport-layer code, not a server response. Requests now carry a User-Agent (gateways reject empty-UA at the edge; the CBOE probe showed no-UA is fine THERE, but EIA fronts differ) and 1xxx codes are explained in the log line. Retries were already hourly. UNBOUNDED BACKFILL: an empty cache fetched full series history - CPIAUCNS goes back to 1913, whose pre-1970 dates are outside MQL5 datetime range and whose 1913-era levels sat below the plausibility band, producing 157 scary-but-meaningless REJECTED lines. All FRED fetches now start at 2005 (5y of lookback margin ahead of the 2010 grid). DTWEXBGS staleness horizon raised to 10 days to match its weekly H.10 publication lag. Also confirmed from the log: the running build predates the H4 fallback, so the H4 panels still show 0 features - resolved by the recompile this commit requires anyway. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
3e9b46452b |
fix(altdata): robustness pass for arbitrary symbol/timeframe - validation + error handling
Systematic audit of the alt-data stack against "any symbol, any timeframe",
prompted by the H4 surprise. Findings, each fixed:
CAPACITY: ALTDATA_MAX_FEATURES was 16 with FX symbols already at 15 - the next
added column would have been silently truncated by a MathMin. Raised to 32,
pin-chars 512 -> 1024.
SYMBOL NAMES: the panel reload path re-derived symbol/timeframe by splitting
the file path on its FIRST underscore - mis-parsing every symbol containing
one (OANDA-style EUR_USD and US_500 are in our own alias lists) and knowing
only three timeframes. It now stores the (symbol, period) Load() was called
with and reuses them verbatim. Path-hostile characters in broker symbols
("EUR/USD") are sanitized by a shared AltDataFileSymbol() used by the panel,
the fetcher and TunedPeriods, so a slash cannot route a write into an
unintended subfolder.
DOWNLOAD VALIDATION: every FRED-family fetch now enforces a per-series
plausibility band (VIX 1-200, yields -5..30, CPI index 20-1000, ...) -
StringToDouble on transport garbage returns 0.0, and one absurd value poisons
every change/percentile feature computed across it (the BatchNorm NaN-latch
incident came from exactly one huge-but-finite input). Rejected rows are
counted and reported, never dropped silently.
LOUD EMPTINESS: a successful response with zero observations on an empty
cache now says so - naming the series (wrong id / format drift) or the COT
predicate (the unverified like-clauses) instead of leaving 0-filled features
unexplained. GEX gains a truncation guard: a day-over-day contract-count
collapse >50% is the fingerprint of a partial 13 MB download, not of markets,
and is skipped rather than recorded as a plausible-but-wrong number.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
862d0f909f |
fix(altdata): H4 charts read the daily alt file; multi-chart write races hardened
The user attached H4 charts and the panel found no {SYM}_H4.csv - the fetcher
only writes _D1 files - so the run trained with ZERO alt features, silently.
The daily file is timeframe-agnostic by construction (rows are as-of daily
values, and Features() joins published <= bar open per bar), so the panel now
falls back to {SYM}_D1.csv on any timeframe, logging the substitution. A
per-TF file still takes precedence if one ever exists.
Per-symbol subfolders (the user suggestion) are NOT the fix for multi-chart
concerns: filenames are already symbol-keyed, and the shared caches
(raw_VIXCLS etc.) are shared deliberately - one download serves every chart.
The REAL races were: (1) two charts of one symbol (D1+H4) each caching their
own last-GEX date and double-appending the same day - UpdateGex now re-reads
the file date before spending the download; (2) whole-file rewrites were
truncate-then-write, so a concurrent reader could parse a torn file - SaveRaw
and RebuildFeatures now write a temp and FileMove-swap it into place.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
120afde2a3 |
research(altdata): H4 timeframe screen - information survives, diluted; costs 2.6x worse
Same harness as the D1 screens, forward 30 H4 bars (~5 days), 199 perms, on the surviving htf mid bars (19.8k-36.4k bars per symbol). Question: does the wired alt/volume information carry to H4, for the chart-timeframe decision. Answer: the signal survives but is roughly halved, and the cost side worsens 2.6x per step down. SP500 vix_chg5 clears the family bar with MI|vol 0.0112 (vs 0.031 at D1); volLevel50 (the EA activity feature) is incremental on all four symbols at H4 - USDJPY family-clean, and it is EURUSD strongest non-control signal there too. Gold gvz_chg5 stays incremental (0.0038 vs 0.0197 at D1 - a fifth of the strength). COT is null at H4 on FX (weekly cadence pasted across 30 bars/week dilutes it below detection on EURUSD/JPY; survives conditionally on SP500). Cost table (median spread/ATR; the 1.74xATR geometry in spread units): SP500 138->52->25, EURUSD 282->113->57, USDJPY 230->87->44, XAUUSD 92->34->16 for D1->H4->H1. Every step down multiplies the cost share ~2.6x. Also fixes the disaggregated-COT column name for gold/WTI in the screen (M_Money vs Lev_Money - the same crash the EA-side catalog documents). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
1f277c254a |
fix(altdata): 4014 alert now NAMES the blocked host; backoff is per-host
The user whitelisted the hosts and still got the alert - because the alert never said WHICH request failed. Two hosts (api.eia.gov, cdn.cboe.com) were added to the EA after the original whitelist instruction, so any build newer than the whitelist raises 4014 on the new hosts while the message implied the old ones were the problem. Three defects fixed: the popup and journal now print the exact blocked host as a copy-paste whitelist line; the stale "three URLs" text is gone (the full four-line reference prints once per session); and the backoff is per-host instead of global - one missing entry no longer silences the whitelisted sources for an hour per miss. Popup fires once per host per session; hourly retries log one quiet line naming the host. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
40ddf06ce6 |
docs(altdata): adjudicate the NASA API trio - POWER queued behind NATGAS, GIBS unconsumable
POWER is the real find of the three: daily temperature -> degree days -> natural-gas demand is the textbook gas fundamental, numeric and daily. But it is point data needing construction into a national series (NOAA CPC ships that ready-made), and its target symbol is not traded yet - fetch code written for a chart nobody attaches first runs months later, unobserved, which is the silent-FRED failure shape. Queued for the AvaTrade expansion, not refused. FIRMS: re-raised, nothing changed since it was parked - point fire detections behind the same unproven proxy chain. GIBS: imagery tiles, not numbers; our CONV is 1D and NASA already sells the extracted products. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
626b591ca6 |
feat(altdata): wire everything the sources serve - screens become priors, not gates
Owner decision (stated twice): available data gets wired; the networks judge
usefulness; the deploy gate remains the arbiter of what trades. Implemented:
MACRO block (6) on every symbol: 10y yield 20d change, curve slope, 5y
breakeven 20d change, Fed-ECB policy gap, CPI yoy, unemployment 12m change.
Screened null vs forward range on all four research symbols - recorded as
the honest prior in the catalog comment, wired regardless.
RISK block (3) extended to every symbol (FX majors, metals, energy, BTC all
now carry vix/vix_chg5/usd_chg5).
IVOL pair extended with the level alongside the change.
Vintage integrity kept where it is free: CPI is fetched as CPIAUCNS (NSA,
essentially never revised) so the plain-FRED backfill stays first-print-clean;
yields/curve/breakevens/policy rates are unrevised by nature. UNRATE is the
one exception (seasonal refits, ~0.1-0.2pp) - the EA cannot run the ALFRED
protocol, accepted and documented at the declaration site.
UpdateFred gains a staleDays parameter so the monthly series do not fire a
pointless fetch attempt every hour for three weeks after each print.
FeatureValue now takes the day and does its own as-of lookups - adding a
source no longer widens a parameter list. Feature counts: 12-15 per symbol;
symbol feature-order changed, safe only because no models exist yet.
export.py mirrors the new catalog for the five research symbols (13-15
features), smoke-tested: all five CSVs written, 6,072 daily rows each.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
3b77500726 |
research(altdata): macro/rates/country data is NULL vs forward range on all four symbols
Screened yields (DGS2/DGS10), curve slope, inflation breakevens, Fed policy,
the Fed-ECB policy differential, and monthly US unemployment and CPI - all on
ALFRED first prints, 499 permutations, against forward 5-day range.
NOT ONE macro feature clears the family-wise bar on any symbol. The only thing
that clears anywhere is the trailing-range positive control, which is what it
is there to do. Best a-priori candidate, the Fed-ECB differential on EURUSD,
came in at MI 0.00170 p=0.088 - nothing. The two features flagged INCREMENTAL
(dgs2_chg5 on SP500) have null marginal MI and are isolated conditional cells
at the expected false-positive rate, not findings.
The `distinct` column quantifies the power argument instead of asserting it:
unemployment takes 51-66 distinct values across 3,745-6,159 bars, CPI 174-277,
against 6,159 for a continuous feature. A monthly series pasted onto daily bars
carries about 1% of the resolution, and it showed - the monthly features were
among the weakest in every table.
The contrast with the implied-vol screen is the useful part: the options
market FORWARD-LOOKING view of an instrument (gvz_chg5 on gold, MI|vol 0.0197)
carries real information about its range, while the economy BACKWARD-LOOKING
state carries none. Mismatched timescales - rate levels move over months,
5-day range moves daily.
Also makes load_bars fall back to htf/{SYM}_D1_mid.npz when the tick-derived
build is absent (the 2026-08-16 disk cleanup removed bars/ but htf/ survived),
with need_ticks=True turning that fallback into a loud failure for the
order-flow screen rather than silently testing flow features on OHLC data.
No EA change: nothing survived to wire.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
d83cecc011 |
feat(altdata): wire instrument-specific implied vol; fix FRED vintage path
Wires the screen_ivol survivors (
|
||
|
|
41d726c746 |
research(altdata): instrument-specific implied-vol screen - GVZ is a major find on gold
No free historical GEX exists: probed the CBOE chain endpoint with date/dt query params (both silently ignored, returned today) and dated/historical paths (403), and the CBOE index-history CSVs are 403 too. The forward recorder stays the only path to GEX history. But the options market publishes its per-instrument view of future range as the CBOE vol indices, and FRED carries the whole family free with 15-25 years of history - screenable today with the existing collector and harness. Fetched GVZ (gold), OVX (oil), VXN, VXD, RVX, VIX3M. HEADLINE - XAUUSD: gvz_chg5 (gold IV 5-day change) MI 0.02103, MI|vol 0.01971, p=0.002. That is 3.6x the trailing-range positive control and 4.6x the vix_chg5 this project currently ships on gold - the second-largest incremental MI of the whole campaign, on a symbol that carries exactly one screened feature today. Vol-change is incremental on all four symbols: SP500 (known), USDJPY vxd_chg5 0.00492, and EURUSD vxd_chg5 0.00412 / vix_chg5 0.00379 - notable because EURUSD has no screened features at all and its own trailing range is a weak control there, so external vol carries information its own history does not. Caveats recorded in the script and memory: SP500 within-family ordering (VXN > VIX3M > VIX, all ~0.031-0.038 conditional) is a best-of-N artifact and must not be cherry-picked; XAUUSD noise control misbehaved this run (MI|vol 0.00271 p=0.002), so anything under ~0.003 conditional on gold is unresolved - gvz_chg5 at 7x that floor is unaffected; EVZ (euro IV) is DISCONTINUED since 2025-03 and must never be wired. Nothing wired - the EA is mid-deploy and this would re-key every model again. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
1bee06c945 |
docs(altdata): FlashAlpha free tier probed - no validation possible, verdict hardens
Spent 3 of 5 daily requests. All three were informative:
ETF data (SPY/QQQ/IWM) requires Basic - free tier is single stocks only.
Full-chain GEX (all expirations) requires Growth - free and Basic must
query one expiration per request, so even with history a full-chain
backfill would be 24-54 requests per day of history.
AAPL?expiration=2026-09-18 returned 200 with the right schema but a nearly
empty payload: 13 of 93 strikes carried any open interest, total call OI
4,296 against CBOE 373,253 for the same expiry, put OI zero, and every
near-the-money strike blank.
So the construction could not be validated - not because the math disagreed
but because there was nothing to compare against. From outside it is not
possible to tell free-tier degradation from their flow-signed methodology,
and finding out costs $1,499/month.
Verdict hardens: the free CBOE CDN is strictly better than Basic for this
project - complete chains, every expiry and strike, gamma and open interest
populated, unlimited, $0. Our own AAPL figures were internally coherent
(+0.929 Bn/1% total, Sep-18 expiry +0.154 Bn, near-money gammas 0.013-0.019).
GEX stays externally unvalidated; if that ever matters, use a different vendor.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
9e49b2aeae |
docs(altdata): FlashAlpha backfill is dead - historical API is Alpha-tier only
Pricing checked: Free $0 (5/day), Basic $79 (250/day), Growth $299 (2,500/day), Alpha $1,499 (unlimited) - and the Historical API is ALPHA-EXCLUSIVE. Basic and Growth serve live data only. The archive was the only thing worth buying from this vendor, so nothing in budget helps: Basic would spend $79/month to make a once-a-day snapshot 15 seconds fresh instead of 15 minutes. Not subscribing. The free key keeps one genuine use: a single live call to compare their GEX against our CBOE-computed number, validating the recorder formula against a commercial implementation (sign and magnitude only - they sign strikes from classified tape, we use the standard open-interest assumption). Recorded the EV argument for future sessions: the recorder banks this history for free in ~12 months, and on this project base rate most alt-data families die at the incremental gate. Paying four figures to test GEX a year early is a poor trade. If revisited, price bulk ARCHIVE sellers, not analytics APIs. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
1693afbd89 |
docs(altdata): register the FlashAlpha key - stored, zero calls made
Free tier is 5 requests per DAY, so the quota is reserved rather than spent: one request answers whether a historical timestamp returns the whole chain or one expiration, and that decides whether backfilling 2018->now is ~2,100 requests (GEX screenable now) or infeasible (compare bulk vendors instead). Deliberately NOT an EA input, unlike FRED/EIA: the EA must never depend on a paid, rate-limited vendor in its live path. Research-side backfill only, so the key lives in the gitignored Market Data\altdata\keys.json plus this backup. Auth is an X-Api-Key HEADER, not a query parameter - worth noting because the EA's HttpGet currently sends no headers. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
3abb2b1a7f |
feat(altdata): GEX forward recorder (CBOE delayed-quotes CDN, no key)
Option open interest is a snapshot source - no free history exists anywhere -
so the series only accrues from the day recording starts. That is why this
ships BEFORE the redeploy: every day the EA is not running is a day of history
that cannot be recovered later.
Records one row per weekday after 21:00 UTC to gex_{CANONICAL}.csv: net/call/put
dollar GEX per 1% move, call and put OI, the three nearest expiries and the
front expiry code. Feeds NOTHING - wiring a feature that is missing across ~100%
of the training sample would waste input width and hand batch-norm a constant.
It becomes a screening candidate at ~250 rows, gated like every other feature.
Thesis: dealer gamma is a RANGE mechanism (long gamma -> hedging sells rallies
and buys dips, range compresses; short gamma amplifies both ways), and range is
this project's one proven channel.
Verified in situ against the live SPX chain before writing any MQL5: 29,362
contracts, 20,993 with nonzero gamma, 54 expiries, total +90.7 Bn/1% (calls
+305.7, puts -215.0), and 100% of net GEX inside 5% of spot. The CDN publishes
per-contract gamma directly, so no pricing model - and no model risk - enters
the recorded data. Also verified the CDN does NOT gate on User-Agent (the old
"CBOE is UA-gated" note in DESIGN.md was a different CBOE path), so plain
WebRequest reaches it.
Dropped a zero-gamma "flip level" field: the probe returned a crossing above
spot while total GEX was strongly positive, which is incoherent - a static
gamma snapshot cannot give a flip level without repricing. Recording a
plausible-looking wrong number is worse than recording nothing.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
0a198fb95f |
feat(altdata): EIA wired, 24-instrument symbol catalog, mapping dialog for unknown symbols
EIA (user directive: "the NN might find patterns in it for both oil and regular symbols"). Weekly Petroleum Status Report via the v2 API - crude stocks ex-SPR, field production, refinery utilization - three features (1y percentile, 4w change, utilization) on EVERY catalog symbol, not just oil. EIA screened NULL on WTI's short 7y sample, so these ship as EXPLORATORY inputs: the deploy gate, not the screen, decides whether a model trained on them trades. Publication stamp observed+6d mirrors research/altdata/eia.py. Symbol handling was hardcoded to three if-blocks; it is now a catalog of 24 instruments x alias lists covering The5ers/FTMO/AvaTrade/Dukascopy/OANDA/IC Markets naming, with prefix matching for the broker suffix zoo (US500.cash, XAUUSDm, EURUSD.r). Adding an instrument is one AddSpec row. COT caches are named by CANONICAL so two brokers' names for one contract share a download. Unrecognised symbol -> a chart dialog (Panel\AltDataMapDialog.mqh, CAppDialog + dropdown) asks which instrument it is; the answer persists in symbol_map.cfg and "No alternative data" is a recorded choice, not a nag. Non-blocking by design: an unmapped symbol contributes 0 features and must never hold up a chart. Also: UrlEncodePart now escapes '%' - SoQL like-predicates use it as the wildcard and an unescaped one corrupts the query; docs/ gains the whitelist URLs, an API-key backup, and the catalog reference. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
e2e6d61855 |
feat(altdata): API keys are now EA inputs - defaults survive the folder wipe
The AltData folder in Common\Files gets wiped before every fresh test, and keys.txt died with it (2026-08-16 silent-FRED incident). The credential now travels with the EA: FredApiKey input, owner key as default; keys.txt demoted to a fallback consulted only when the input is blanked. EiaApiKey stored the same way - reserved, nothing consumes it since the WTI screen came back null. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
3248a16638 |
fix(altdata): missing keys.txt failed SILENTLY - FRED never ran, SP500 rebuild never fired
The 2026-08-16 first live fetch looked complete but was not: COT (keyless) downloaded 1053 reports, then FredKey() hit the absent keys.txt, returned "" with no log line, UpdateFred bailed, and the SP500_D1.csv rebuild - gated on all three raw series - never happened. The panel stayed at 0 features with nothing in the journal explaining why. FredKey() now logs loudly when keys.txt is missing, and only latches once a key is actually FOUND: the file is re-read on each hourly-throttled attempt, so dropping keys.txt in after attach recovers without a restart. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
2f901994fa |
test(altdata): flow/activity screen - tick ACTIVITY clears family-wise on all 4 symbols for range
VPIN-style toxicity (|imbalance|): null-to-marginal everywhere. But tick ACTIVITY (count vs 20d mean) clears the family bar on ALL FOUR symbols for forward range AND survives conditioning on trailing realized range (SP500 MI|vol 0.025, XAUUSD 0.0088, USDJPY 0.0069, EURUSD 0.0049, all p<=0.006). On EURUSD it beats the trailing-range positive control itself - resolving the void-control anomaly: EURUSD D1 range IS predictable, just not by its own trailing range. USDJPY spread_stress (max/mean) also family-clean + incremental. Direction: nothing beyond the known SP500 leverage effect. Validates the EA's volume feature block for the RANGE objective the tuner now optimizes. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
0788238c00 |
feat(inputs): unify ALL indicator periods under the tuner; EnableAltData input; AI-first defaults
- PeriodMA/MA_Type/PeriodRSI: input -> const seeds (closing the set: every
indicator parameter is now tuner-owned)
- Variables\TunedPeriods.mqh: chart-level tuned-period state. A gated
install writes TunedPeriods_{SYM}_{TF}.cfg; next attach reads it BEFORE
the DB fingerprint and classic-signal config, so classic votes, DB key,
and tuner seeds always describe the same indicators regardless of
classic/AI/hybrid use. Restart-grained adoption by design (no mid-run
handle churn); new periods re-key the signal DB (semantics rule).
- EnableAltData input in AI Input Features (consumption gate only;
collection keeps running); |ALT DB-fingerprint token; opt-out on an
alt-trained model correctly starts fresh via the width compare.
- Defaults: all four classic votes OFF (AI-first; WARRIOR_MARKET_BUILD
branches collapsed with the marketplace pivot), order-flow/Wyckoff NN
features OFF (alt data is the default information diet; toggles stay).
Compiles 0 errors / 0 warnings.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
ffeb136537 |
refactor(inputs): prune 18 AD/Wyckoff menu inputs; auto-tuner defaults ON
The 18 inputs added 2026-08-08 (when the tuner defaulted off and the values needed an operator path) become compile-time aliases of their own defaults - same names, zero consumer churn, byte-identical values. The tuner is now the only path by which these values move: it defaults ON (the 08-08 off-flip was measured against the direction target's flat landscape; the objective is now RANGE, which has signal), searches from the seeds under the Sidak family-wise gate, and persists winners in the .nnw beside the weights. ADP fingerprint token retired (deviation now impossible by construction; tuned values were never its job). Menu shrinks 102 -> 84 inputs. Compiles 0 errors / 0 warnings. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |