forked from animatedread/Warrior_EA
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
75d7362161 |
feat(warrior): the vol-gated dip-buy book, ported into Warrior_EA
Warrior's defaults are now the validated book: DIP_ZSCORE alone, long only, H4, risk 0.25%, one chart per index with a shared Magic. - System/BarCache.mqh: whole-history closed bars, Wilder ATR, GK sigma and the expanding vol percentile (no 1024-bar stdlib ceiling) - System/AccountGuard.mqh: open-risk cap, kill switch, cross-chart lock and Friday flat, shared through terminal globals by Magic - CWarriorExpert: guard on every tick; a transient open failure retries the bar - CWarriorSignal::SetupStop: the dip owns its 3 x Wilder ATR stop from the bid - SignalDipBuy: no entry vote while holding (a still-dipping time exit never closed, and Processing re-entered on the exit bar); no entry on a stop bar - WarriorMoney sizes on equity; WARRIOR_RISK allows fractional risk - TradeLog + research/compare_ea.py: trade-for-trade check vs WarriorDipZ - SP500/US30/DAX40 identical to the cent, NAS100 96.9% (stale-quote timer fills) - research/nn_cross_index.py: pre-registered cross-index NN meta-label - FAIL (AUC 0.564, CI [0.498, 0.630]); DipMetaCut stays off Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
47a5ef338b |
Refactor Warrior EA: Integrate custom signal modules, enhance voting mechanism, and improve management features
- Replaced standard library signal modules with custom implementations to allow for named patterns and improved voting. - Added new input parameters for module weights, allowing for optimization of individual signal contributions. - Enhanced the management of trades with new options for breakeven and management cut. - Introduced a mechanism for dynamic ranking of signal weights based on historical performance. - Improved initialization logic to ensure proper registration of filters and handling of trading conditions. - Added detailed logging for trading permissions and account status during initialization. |
||
|
|
5f8a2b1df8 | Remove obsolete log and data files: deleted cpu_directml.log, opencl.log, and profiling.csv to clean up the repository. | ||
|
|
ac57a720e0 |
wip: snapshot before the KISS restructure
Everything from tonight, committed so the restructure that follows is recoverable: the graded stdlib vote, the Wyckoff modules and feed, the ALGLIB serializer workaround, the restored DB queue, and the Simple/ prototype that is about to be folded into the real filetree. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
e37f15edb3 |
build: tag books-d1-census-1 for the first deploy of the swing-book and census fixes
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
fa0d0fccc3 |
perf(journal): in the tester, read the journal once at deinit, not daily
The daily adaptive pass reads the whole table; with the shadow journal that table holds every firing so far, and a decade of daily full reads of a growing table is quadratic - 2% per five minutes on the first fill run. Live keeps the daily cadence; the tester learns once, at the end, and the next run adopts it. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
10463eba2f |
feat(meta): the meta-labeller - a random forest in pure MQL5, learned from the shadow journal
Operator: "could we use mql5's alglib random forest and mlp instead of
relying on python? very quick training could reopen the door to selling
the bot."
Database\MetaLabel.mqh. Trained from the Virtual:<setup> rows (every
firing, not the few the gate traded): 14 features parsed from the
notebook line by ONE parser shared with the live gate (setup ordinal,
side, agreement count, opposed, armed count, headroom, minutes to the
forced close, risk in bp, day, hour, order type, valid test, window,
target); label = the firing paid after cost under its own setup's exit,
re-priced from the ladder as research/pricing.py prices it; cost = the
symbol's spread plus the class's commission, printed with every run.
The honest number is walk-forward: for every year from the third, a
forest trained on the years before scores that year, and the take
threshold is the one whose out-of-sample rows paid best - adopted only
if it beats taking everything, else no gate is written and a stale one
is deleted. Final model on every row, ALGLIB CDecisionForest via the
builder (100 trees, 0.66 subsample, Gini importance), serialised to
Adapt\{SYM}_{PERIOD}_meta.rf with the feature contract in the header; a
file whose contract differs is refused.
Gate: in the agreement block, last, as the mean P(pay) over the setups
armed on that side; a refusal is counted as its own entry gate. Trained
at deinit in the tester (fill, then learn, then trade gated - the 2024
loop) and daily when live. Pure MQL5, no DLL, no Python.
Also: virtual firings advance once per minute, not per tick - the
per-tick walk made a decade run four times slower.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
|
||
|
|
af3db90353 |
feat(journal): the shadow journal - every firing, independently, next to the trades taken
Operator: "meta labeling should drastically change the outcome anyways. unless the journal logs every patterns independently (I think it should and also voted trades outcomes for comparison)." Until now the journal recorded POSITIONS - what the agreement gate let through - so a meta-labeller would have learned from a few dozen trades per setup. Every arming is now a virtual trade: a pending order at the setup's own entry and stop, filled when price reaches it inside the setup's own window, tracked through the same first-passage ladder as a real position (one shared AdvanceTrack), closed at the ladder's last horizon, written to the same table as filterID="Virtual:<setup>". The parent annotates each bar's firings with agree=N, opposed and the armed combination. Real trades stay filterID="Book", so the two populations sit on one table. Closed virtual rows are written 200 per commit. Context now also carries headroom, minutes to the forced close and the order type. The `journal` global lives in the header so setups and the signal base reach the one instance the expert feeds. Also: SQL identifiers are quoted - "2WD_Pattern_0_Sell" starts with a digit and every Second Wind pattern table failed to create the moment the database was on in the tester; an index on the natural key so the replace-on-key scales to a decade of firings; "no closed trades yet" is verbose again. Smoke, BTCUSD H1 one month: 52 firings, 26 filled, 26 rows, annotated, no database errors. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
ad327735db |
feat(journal): the tester fills the database, and the journal can skip a categorical loser
Operator: "is there not a meta labeling neural network in the EA? why did it keep taking a losing pattern? The journal is there for that." There is no meta-label net (CNNFilter is a loader nobody instantiates; the offline net was never built), the adaptive layer learns management only by design, and it could not read anyway: SignalDatabaseActive() switched the DB off for every tester run on the argument that "a tester run's DB is written and never read" - a premise AdaptiveExitWriter had already deleted. Every run ended with 93 verbose-level "journal read failed - database not initialized" lines and learned nothing. - The DB is ON in a single tester pass; optimisation and forward stay off (12 agents on one FILE_COMMON SQLite file finished zero passes). - A closed trade REPLACES its own row on the natural key (symbol, open minute, side, entry price): tester tickets restart at 1 on every run, so a re-run over the same history no longer double-counts. - AdaptiveExitWriter writes skip_long / skip_short when a setup-side is a categorical loser on this chart: n >= 60 and t <= -2.5 on realised R x risk in bp, gross of commission. Not ranking - the measurement that made selection anti-predictive ranked marginal cells against each other; this is the shape Second Wind showed on forex (t -4.0). A skipped side never arms, so it never counts toward agreement. - The reader adopts the flags with the exit, once per server day, and prints the verdict when it changes. - "Could not read" is printed once at normal level with the reason. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
7abad65c80 |
feat(book): port the seven missing triggers, and the classic crowding veto
Operator: "port everything from research... there cannot be too much
analysis." Two things, and the first is the cause of the second.
SEVEN SETUPS PORTED - 14 to 21
The journal carries 25 distinct triggers where the EA had 14, and the gap
was not random: EVERYTHING the EA implemented was a BAR SHAPE. An inside bar
is by definition not an engulfing bar, a pin bar is a third shape, and only
ma_cross was shape-independent. Four mutually exclusive patterns cannot
agree, so the agreement count - the only measured edge this project has -
had almost nothing to count. Measured on EURUSD M5 over 74,678 bars, peak
simultaneous arming was 0:68761 1:5333 2:517 3:63 4:4, which is agreement 2
on 0.78% of bars and 3+ on 0.09%.
sr_retest forex p89-91 broken resistance becomes support
tl_bounce forex p92-95 a line through two KNOWN pivots
tl_break forex p92-95 the same line, broken
cci_div forex p237-243 price/oscillator divergence
pin_bar_c forex p206-209 the book's actual 3-bar pin, option 2
cons1234 stocks p185-186 four small bars at the extreme
fvg SMC a three-bar imbalance
NONE OF THESE IS A BAR SHAPE - levels, lines, an oscillator, a gap - so each
can fire on the same bar as an inside bar. On SP500 M5 that took agreement 2
from 2.40% to 4.70% of bars and 3+ from 0.23% to 0.64%, and the plan audit
now reports 19 setups eligible together where it reported 4.
Trend lines read pivots only k bars after they print, which is where that
study usually leaks the future. cons1234 carries no window because the
stocks book explicitly calls it session-agnostic - which also makes it one
of the few that can agree across hours. fvg is labelled book="smc" and is
reachable ONLY through BOOK_ROSTER_ALL: AUTO must never hand a chart a setup
no book endorses, and no result it contributes may be called a book result.
THE CLASSIC MODULES EARN THEIR PLACE - AS A VETO, INVERTED
Their fate, asked properly: given a book setup has fired, does a classic
module agreeing change what it is worth? Measured on ~1.3M journal rows
across 8 instruments, sparse (0-2 of 26 patterns) against crowded (3+):
agreement 1 +1.074 vs +0.407
agreement 2 +2.283 vs +0.696
agreement 3+ +3.922 vs +2.016
Positive in all three periods at every level, and sparse beats crowded on
7 of 7 instruments. ONE binary comparison, not the best of a grid. The
classics are not confirmation - they measure how OBVIOUS the move already
is, and the book setups are largely reversals and breakouts, which do worse
when everything already agrees. That mechanism predicts the sign, which is
why this is trusted where their standalone edge (at chance) is not.
Scale does not transfer: research counted 26 PATTERNS, each module here
reports ONE direction, so the EA counts modules and tops out near 15.
Calibrated on the EA's own distribution instead - 14,338 armed firings, and
vetoing at 6 keeps 19% of them against the journal's 20.75% sparse cell.
AND THE EA DOES NOT REPRODUCE THE SIZE. SP500 M5 2024, agreement 1:
no veto 1780 round turns +0.354 bp t = 1.02
veto at 6 465 round turns +0.414 bp t = 0.63
Right direction, 1.17x where the journal measured 2.6x, on samples too small
to tell either from zero. Two decisive sources disagreeing about magnitude
is what an input is for, so Signal_CrowdVeto is one, defaulting to 0.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
9e4b03ad44 |
feat(audit): report what the quorum actually did, not only what it could
The init audit says whether a quorum COULD form from the declared windows. EURUSD M5 proved that is not enough: the audit correctly reported "agreement 2 of 4 setup(s) is reachable" and the year traded NOTHING. Both were right. The four forex setups share sessions and are MUTUALLY EXCLUSIVE BAR SHAPES - an inside bar is by definition not an engulfing bar, a pin bar is a third shape, and only ma_cross is shape-independent. They can be eligible together and never arm together, and no init-time check can see that. It is the second structural reason within-book agreement cannot work, after the Stocks book's disjoint windows. So the run now reports what actually happened. ReportArmingHistogram() samples the peak simultaneous arming ONCE PER BAR, taken where the counts already exist so "armed" keeps a single definition, and prints at deinit: ARMING over 74678 bar(s) - peak simultaneous setups per bar: 0:68761 1:5333 2:517 3:63 4:4 | the threshold of 2 was reached on 584 bar(s) (0.78%). ...with an explicit "NOTHING COULD HAVE TRADED" when the threshold was never reached. Init asks could it; deinit asks did it; the pair is the diagnosis. That measurement is also a finding in its own right. Agreement 2+ occurs on 0.78% of bars and 3+ on 0.09% with ALL FOURTEEN setups enabled. Research's 3+ cell held 416,366 rows because its journal carries 36 triggers where the EA implements 14, so co-firing is far rarer here and the threshold optimal there is structurally too strict. The default of 2 is right for this EA. CExpertCustom exposes the report as a BEHAVIOUR rather than widening GetCustomSignal to public: OnDeinit needs this one line, not the signal pointer, and handing out the pointer to get it would widen the class for a single caller. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
c961f5c8d7 |
feat(audit): the EA states its resolved trading plan at init
Operator: "the EA should be dynamic and work on any timeframes and symbols.
it needs its own internal research framework."
Every failure this EA had today was a SCALE failure - a rule meaning one
thing where it was measured and something else on the chart it ran on:
* the opening reversal is a 20-minute construct; on H1 that is ZERO bars
* a 120-minute time stop is 24 bars at M5 and TWO at H1 - a different rule
* the Stocks book's three windows are disjoint, so a quorum of 2 is
unreachable and the run trades nothing
* m_entry was shared by both sides, so an armed setup shaped from a zero
None of those raise anything. The EA trades nothing, or trades a different
strategy, and looks exactly like a quiet market doing it. Three cost a full
test run to find; the fourth had been live since the setups were written.
The fix is disclosure, not cleverness. At init every setup now states what
it will actually do on this symbol and period, and what cannot be true here:
===== TRADING PLAN, resolved for EURUSD PERIOD_M5 =====
IB inside_bar forex win 10:00-12:00,15:00-17:00 entry 1 bar
stop 1.00R target none trail 1.00/1.00R be 1.00R timestop 24 bars
...
agreement 2 of 4 setup(s) is reachable - up to 4 eligible together
===== 4 setup(s), no scale problems =====
DESIGN, and it is the reason this cannot rot. Each filter writes its OWN
line (PlanLine) and answers for its OWN coherence (ScaleIssues); PlanAudit
only collects and orders them. Adding a setup, or a rule to one, cannot
leave the audit behind because the audit knows nothing about either.
ValidationSettings now asks ScaleIssues rather than carrying its own copy of
the timeframe test, so a rule is checked in exactly one place.
ScaleIssues catches, among others, the case that cost today's run: a
wall-clock time stop that resolves to under three bars on this chart is a
value measured on a faster one, and says so.
It warns, never refuses - one roster can hold setups from books of different
granularity, and aborting because one cannot fire here is a worse answer
than running the rest and saying so. WarnIfAgreementUnreachable is called
from the audit so there is a single init entry point.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
0c58e55173 |
feat(adapt): build the writer half of the live-learning loop
Operator: "I just want it to keep learning during live trade. similar to
2024 versions of the ea where I would run a backtest to fill the database
and then it would keep filling in real time, averaging on the whole sample."
Most of this already existed. CBookSetupSignal::LoadAdaptiveExit() has
always been the READER - once per server day it re-reads
Common\Files\Warrior_EA\Adapt\{SYM}_{PERIOD}_{setup}.cfg and adopts the exit
in it, refusing anything outside the research grid. FILE_COMMON is the whole
point: the tester and the live chart share one file and one journal table,
so a backtest fills the sample and live trading carries on filling it.
Nothing ever WROTE that file.
And the journal already records what is needed to re-price an exit without
re-running anything: `passages` carries the first-touch minute of 9 stop
levels and 10 target levels plus the signed R at 6 horizons, in exactly the
field names research/pricing.py reads.
So this is the estimator: read every recorded trade for this symbol, group
by setup, re-price each one across the (target x horizon) grid with the same
first-passage rule pricing.py applies, and publish the winning cell.
WHY THIS IS NOT THE FEEDBACK LOOP THAT WAS DELETED. Two differences, both
measured rather than asserted:
1. IT LEARNS MANAGEMENT, NEVER SELECTION. Selection by a cell's own past
P&L measured ANTI-predictive on this journal - a cell gets WORSE as its
evidence accumulates - while management measured positive. So it may move
a target or a time stop. It may never decide which setups fire, what the
agreement threshold is, or what any vote weighs. DB_RankingFeedsWeights
stays false.
2. IT IS A CUMULATIVE MEAN OVER THE WHOLE SAMPLE, NOT A ROLLING RE-FIT. The
harm in the old loop was a MOVING RULER - a pattern's contribution changed
as the DB re-scored it, so the same setup voted differently at different
times and nothing could be evaluated against anything. An average over
everything ever recorded converges instead of chasing.
THE ASYMMETRY THAT WOULD OTHERWISE POISON IT, and TradeJournalManager's own
header stated it before this was written: a live trade is closed by its own
exit, so levels beyond the one it took are never touched and record as -1.
From live rows an exit can honestly be re-selected TIGHTER, never WIDER, and
a layer ignoring that would "learn" that wide targets never pay because it
never saw one reached. That is not hypothetical - today's journal sweep found
the best exit is the WIDEST (6R on a 1R stop is never touched, so it is
really "no target", with a long horizon: +2.311 bp at 3+ agreement, positive
in all three periods). So the writer is hard-capped at "never wider than what
produced the evidence".
Other guards: 120 trades before a setup is considered at all, 60 re-priced
rows before a cell is, a positive mean, and at least one neighbouring cell
also populated and positive - a plateau rather than a spike, because picking
the argmax off a grid is what this journal punishes. The stop is never tuned:
a book setup's stop is its structure and the books state it.
FetchClosedTrades is now public - it is the read side of this loop and the
writer only reads.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
5418a5e326 |
feat(book): port the session findings research had and the EA did not
The operator asked whether everything from research was ported. It was not,
and the gap was systematic: every TIME-BASED finding was missing.
THE FOREX BOOK'S THREE ENTRY WINDOWS
BookForex.mqh contained no SessionWindow call at all, so all four forex
setups fired around the clock - while every research number that measured
them ran only inside the book's own windows. The book names them (p145):
London 03:00-05:00 ET, New York 08:00-10:00 ET, Tokyo 19:00-21:00 ET, "the
first one or two hours after a session open".
This is the same mismatch that made research/add_session.py apply the
STOCKS cash session to forex. Correcting it there TRIPLED gross edge on M5;
every pair went positive gross on both timeframes, EURUSD M60 net
-0.856 -> -0.271 bp against a 0.80 round turn. Raw hourly range confirms it
independently of any trade definition: a 3.5x spread across the day, and
the dead hours are exactly the ones an all-day EA was trading.
The three windows are DISJOINT, so CBookSetupSignal now holds a SET of
windows rather than one pair. That is also why the forex windows could not
simply have been added before.
NO NEW ENTRY INSIDE THE LAST 60 MINUTES
The most robust result of the campaign, and it was absent entirely - nothing
in the EA knew how much session was left when it placed a trade. On 951,919
session trades from 2005: enter with 30 minutes left and 87.4% never resolve
at all, so the forced close decides them, and that close is worth +2.8 bp
early against -1.0 bp late. Late setups are not worse patterns; they are
never given room.
cutoff kept net bp vs all years improved
15 96.2% -0.961 +0.039 22/22
30 92.1% -0.936 +0.064 22/22
60 83.4% -0.895 +0.105 22/22
120 67.7% -0.867 +0.134 22/22
Monotone, saturating near 90-120, improving EVERY one of 22 years. Trusted
where mined rules are not because the mechanism predicted it before it was
measured. 60 rather than the 120 that measured best: 60 carries 78% of the
improvement while keeping 83.4% of trades against 67.7%, and trade COUNT is
the binding constraint right now. One #define to raise it.
Measured against WarriorScheduledFlatMinuteOfDay - the SAME forced close the
expert actually applies, now extracted so both callers share one owner
rather than each deriving a schedule that can drift from the other.
Context - BTCUSD H1 2024, equal weight, risk guard off, 0.10 lots:
agreement 1, no clock 391 RT -5.65 bp
agreement 1, clock 113 RT -3.15 bp
agreement 2, no clock 116 RT -2.11 bp
agreement 2, clock 31 RT +18.93 bp <- first positive, but n=31
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
534d622ae5 |
feat(book): give crypto a measured clock, and say when a quorum is impossible
Two fixes for the same class of failure: a configuration that trades
nothing while looking exactly like a quiet market.
THE CRYPTO CLOCK - OURS, AGAINST THE BOOK, AND MEASURED
The Bitcoin book says crypto has no session and no hour effect: "as active
at 3AM on a Sunday morning as it is at 9AM on Monday" (p38, p77). The
second half is FALSE, on 868,560 BTCUSD M5 bars from 2017-05 to 2026-09:
quiet 05:00-14:00 server 51.0 - 60.7 bp
middle 00:00-04:00 63.4 - 71.8 bp
ACTIVE 15:00-23:00 66.3 - 91.8 bp peak 17:00 = 91.8
quietest 07:00 = 51.0, busiest 17:00 = 91.8, ratio 1.80x
Sat 0.59x, Sun 0.63x against weekdays 0.99-1.02x
WHY IT IS A COST RULE AND NOT A SESSION CLAIM. The book is right that there
is no open, no close and no meaningful "day", and nothing is gated on one.
But cost in R is cost_bp / dist_bp, so at 51 bp an hour instead of 92 the
same pattern gives a stop about half as wide and pays close to TWICE the
relative cost for the identical trade - on the class that already pays the
highest commission we trade.
HOW THE WINDOW WAS CHOSEN, so it is not a fitted parameter: the median of
the 24 hourly medians is 64.25 bp, and 15:00-23:00 is the ONE CONTIGUOUS
BLOCK entirely above it. It lands on the US cash session plus the hour
into it. One statistic, one threshold, one contiguous run, no search.
The weekend skip is cruder and needs no threshold at all: 0.59x and 0.63x.
AN UNREACHABLE AGREEMENT THRESHOLD NOW SAYS SO AT INIT
SP500 M10 under AUTO ran a full year and traded nothing at agreement 2.
Nothing was broken. The Stocks book's three setups are SESSION-DISJOINT BY
DESIGN - opening reversal 16:30-19:00 server, ERBO 17:00-19:00, PDH/PDL
21:00-23:00 - so PDH/PDL can never be armed on the same bar as either
morning setup, at most two of three are ever eligible at once, and a quorum
of two needs a gap fill and a range break on one bar and one side.
WarnIfAgreementUnreachable() walks all 1,440 minutes, counts how many book
setups are ELIGIBLE at each, and takes the maximum. Firing is rarer than
eligibility and can only be rarer, so that maximum is a hard ceiling: a
threshold above it is unreachable with certainty, not merely unlikely. It
warns rather than refuses, so a sweep across rosters still scores the pass.
Measured today on BTCUSD H1 2024, equal weight, risk guard off:
AUTO agreement 1 391 RT -5.65 bp (research: -6.47 for one alone)
AUTO agreement 2 116 RT -2.11 bp
ALL agreement 2 269 RT -2.18 bp
ALL agreement 3 64 RT -5.13 bp
The 1 -> 2 improvement is the confluence mechanism reaching the EA for the
first time. It does not continue to 3, and nothing is positive yet.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
5f1c459043 |
refactor(inputs): commission out of the EA, 63 inputs down to 22
Operator: "commission is automatically charged on mt5 during backtest, so
make sure to not include anything related to commissions in the EA. have as
little inputs variables as possible for a clean menu."
COMMISSION. Expert\CostModel.mqh deleted, the six Cost_Comm* inputs with
it, along with the per-bar spread sampler in OnTick and the init COST MODEL
printout. This removes a DUPLICATE, not the cost: MetaTrader applies the
broker's own schedule to every deal in the tester and on the account, so a
hand-typed second copy inside the EA could only ever disagree with it - and
a schedule that drifts from the broker's is worse than none, because it
looks authoritative. Nothing consumed the numbers any more in any case; the
gate that read them went with the era verdict.
Only the asset-class detector survived, moved to Expert\AssetClass.mqh. It
answers a question MT5 does not: which book applies.
THE ROSTER IS ONE INPUT NOW. Fourteen per-setup constants stood in
ClassicSignals.mqh, all false, none reachable without a recompile. They are
replaced by Book_Roster, chosen a BOOK at a time - which is the only
grouping that respects the accountability rule (one book per asset class, a
setup is only accountable on the class its own book covers) and the only
one the agreement count can use, since a roster of one cannot express it.
AUTO reads the class off the instrument. Eight tester values instead of
16,384 mostly meaningless combinations.
INPUTS 63 -> 22 (17 parameters and 5 section headers). Cut or made const:
everything dead after the training deletion, the two API keys (a credential
is not a strategy parameter and must not travel in a .set), and knobs that
cannot change an outcome - Signal_ThresholdOpen above all, since Direction()
now returns exactly 0 or +/-100, so every threshold in (0,100] behaves
identically.
TWO OF THOSE ARE FIXES, NOT TIDYING.
- The 30-bar signal cooldown was still live. Its floor came from the
leg-ride label ("a trade is held 5 + the median ZigZag leg = 18-19 bars").
That label is deleted, and the thing being spaced now is a book setup that
owns its trade end to end. Worse, it would have silently thinned the very
population the agreement count counts. Off, and const.
- Signal_MinAgreement defaults to 2 rather than 1. One trigger alone is the
policy measured NEGATIVE on every asset class.
The session filter is const off for the same class of reason: every setup
carries its own window from its own book, and a global one layered on top
applies one book's clock to another book's market.
SweepGuard's table was rebuilt - it was almost entirely names from the
deleted training layer. It now refuses Asset_Class (a statement about the
instrument, not a strategy choice) and pins the log-volume switches.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
338e2358be |
feat(book): the agreement count - the only measured edge, now enforced in the EA
An armed setup used to be decisive on its own. That is exactly the policy the research measured, and it is negative: across 942,918 journal trades, one trigger alone is -2.15 bp out of sample and up on 2 of 8 instruments, two agreeing +5.65, three or more +7.59 at 240 trades a year and 8 of 8, monotone in all three periods. Gross moves as much as net, so it is not a cost artefact. Direction() now counts DISTINCT setups armed on the same bar and side, over the same population ArmedSetupOwner() picks the order-shaper from - so the count and the shaper can never disagree about who fired - and takes the trade only at Signal_MinAgreement or more. OPPOSED IS TESTED FIRST, because it is a veto and not a tie. Setups firing both ways on one bar is the worst cell measured anywhere in this project, -5.80 bp pooled and -21.31 on BTCUSD, negative gross as well, and worth more than the agreement bonus is. It stands aside and says so once per bar. Firings that arm but fall short of the threshold are traced rather than dropped: those declined signals are the population the offline meta-label net has to learn take-or-skip from, so they belong in the journal. Signal_MinAgreement is an INPUT because choosing it is a Strategy Tester question - 2 buys far more trades, 3 a better rate, and the walk-forward decides. It defaults to 1, which is the measured-negative policy, only because that is also what a chart with a single setup enabled must do; a silent no-trade would be worse than an honest bad default. All three Walsh books state the principle in words, so this is the books' own rule rather than an overlay on them. Compiles 0 errors / 0 warnings. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
f7b12413a4 |
feat(book): fourteen book setups as signal modules, each owning its own management
The books are now implemented in the EA rather than only in the research repo. Fourteen setups
across four books, one class each, all on CBookSetupSignal:
FOREX inside bar (mother-bar entry), pin bar, engulfing, 5/30 MA cross
BITCOIN inside bar (own-bar entry, next bar only), pin-after-inside-bar, EMA10/21/100 + MFI(2),
second wind, triangle on rising volume
STOCKS opening reversal / gap fill, early range breakout, previous day's high-low
WYCKOFF spring, sign-of-strength bar
EACH BOOK'S OWN RULES, NOT AN AVERAGE OF THREE. The forex and bitcoin inside bars are separate
classes because they are separate setups: one enters at the MOTHER bar's extreme and leaves the
entry window unfixed, the other at the INSIDE BAR's own extreme on the next bar only. Collapsing
them measured the average of two books nobody wrote. The forex book's qualifiers are applied
rather than emitted raw - angled-and-not-flat average, inside bar smaller than its mother and at
the correct extreme, the engulfing bar's one-or-two-bar pullback taken with the trend.
MANAGEMENT IS PER SETUP because the books disagree: stocks breakeven at 1.25R (the corrected
constant), bitcoin's ratcheted trail, forex's ~1R breakeven and trail. Two setups carry MEASURED
targets set per firing rather than a fixed multiple - the triangle's mouth and second wind's leg
projection - which is what those books actually specify.
The session window moved from the stocks class into the base: a setup several books carry has a
window in one and none in another (stocks confine triangles to the afternoon, the Bitcoin book
states crypto has no session at all), so whoever constructs it says which reading is traded. The
base gates LongCondition/ShortCondition on it so no subclass can forget.
TWO THINGS DELIBERATELY NOT BUILT, and named rather than faked: zero bouncing, because the
round-number premise under it measured null against matched controls; and ERBO's "clear bias"
precondition, which is a discretionary read of the pre-market with no honest mechanical stand-in.
Second wind's book-stated stop is a round number for the same reason and is ours instead.
All default OFF and individually switchable - which roster to run is a Strategy Tester question,
not an EA one. Note they are useless alone: one trigger measured NEGATIVE on every class, and it
is three or more agreeing that pays, so the agreement count needs candidates to count.
Compiles 0 errors / 0 warnings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
c6f70d8ee1 |
refactor(inputs): remove the Neural Networks section - seven dialog switches that did nothing
Follow-up to the training deletion. The whole "Neural Networks" input section survived the file
deletions and every one of its knobs was inert: Use_MLP / Use_CONV / Use_LSTM / Use_CONVLSTM,
Run_Alglib_Baselines, Use_Training_Pool and OOSSplit are inputs the operator sees in the dialog,
and after the cut none of them reached anything. Exit_On_Leg_Flip was the same - a switch for a
label that no longer exists. 132 lines, and Inputs.mqh drops 1,022 -> 891.
An inert switch in the Inputs tab is worse than a deleted one: it invites the operator to change
a setting and conclude the EA ignores them, which is close to what the complaint that started
this refactor actually was.
TWO SLOTS DELIBERATELY KEPT AS LITERALS so no existing database is re-keyed. DbLegacyAiSlot()
encoded which architectures were enabled and now returns the legacy AI_NONE value; the optimizer
slot is written as 1, ADAM's ordinal, which is what TrainingOptimizerDefault always carried
(ENUM_OPTIMIZATION pins SGD/ADAM at 0/1 precisely so this cannot drift). I first wrote 0 there,
which would have silently changed the fingerprint on every database this EA has written.
The trade journal's run label was the enabled NN roster ("MLP+LSTM"); it is now "Book".
Compiles 0 errors / 0 warnings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
4e5e6fcae3 |
refactor(ea): delete on-chart training - the EA stops learning and starts executing
Operator, 2026-09-07: "we don't want to train on the legs anymore." That removes the reason the whole training stack existed, so it goes. 77 files deleted, 68,078 -> 29,113 lines: 57% of the codebase. A full build drops from 66s to 24s. Compiles 0 errors / 0 warnings against a 0/0 baseline taken before the first cut. THE SEAM. The four AI signal modules (PAI/CONV/LSTM/HYBRID) were the only consumers of ExpertSignalAIBase -> AI/Network -> AI/Impl/*, AIBase/*, Training/*, Persistence/*, Topology/*, Labeling/*, Features/*, OnlineLearning/*, ConfigLock/* and the training half of Chart/*. Cutting those four dropped all of it. Nothing else reached in. WHAT SURVIVES, and it is the part that matters: System\NNFilter.mqh - 225 self-contained lines with their own forward pass, reading a plain ASCII model written by research/export_nn_filter.py, with the feature-name contract that REFUSES a file whose feature list does not match rather than approximating it. The offline meta-label net's entire runtime already existed; it never needed any of what was deleted. THE LEG-RIDE LABEL AND ITS EXIT. LiveLegDirection() replicated the stock ZigZag so the exit could fire on the same event the label's ride ended on. No label, nothing to agree with - the replica, the exit and Expert\Labeling\LegState.mqh are gone, and the take-profit is unconditional again (it was suppressed only to avoid capping the tail the leg label selected for). ONE THING THIS NEARLY DID SILENTLY. ClassicVotesMoveMoney() returned "no while an AI member is present, yes when none is registered". Deleting the networks made the second clause true everywhere - which would have reinstated the worst defect this codebase has had: fifteen classic modules at weight 1.0 as the live money vote, which is what "the EA is not profitable" turned out to mean. It now returns false unconditionally. The only thing that can open a trade is an ARMED BOOK SETUP through the +/-100 override in Direction(). The classics stay wired because they are silent and free, and because they are the raw material for the agreement count. Also extracted System\ChartObjects.mqh - the chart-object namespace list and its sweep, which had lived inside ExpertSignalAIBase.mqh and were never about training. NOTE THE CONSEQUENCE, PLAINLY: the only book setup that exists in MQL5 is CSignalInsideBarGap and it is still disabled, so this EA now trades nothing until the book triggers are wired. That is slice 3 in REFACTOR_PLAN.md and it is a deliberate state, not an oversight. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
6df9132274 |
fix(gate+label): certify the trade the EA places - traded population, real fills, no survivorship
The second half of the sweep. Everything the first pass reported and left. THE ENSEMBLE GATE JUDGED A POPULATION THE ACCOUNT NEVER SEES. It scored every OOS bar whose vote cleared the rung; live, a vote that clears the threshold still has to pass VoteCooldownAccept, and a suppressed bar produces no arrow, no order and no position. On this fleet that window is 30 bars against a mean ride of ~34, so the certificate counted roughly an order of magnitude more trades than the account could take, each overlapping its neighbours. The member gate was fixed for exactly this on 2026-09-03 and both sites carried a comment saying the ensemble still had the defect. The sweep now replays the vote cooldown per rung: rows arrive in time order, so one kept-timestamp per rung reproduces it exactly. sweepTraded[] carries precision, the book, the cost and the by-side counts; sweepFired[] stays the signal population and only coverage reads it, because a cooldown caps traded coverage by construction. The effN uses the declustered helper - the cooldown has already spaced that stream. BOTH FAMILY-WISE SELECTION GATES FORMED THEIR SE ON RAW CALL COUNTS, the last SEs in the project still undeflated for label overlap, which made the correction guarding the deploy decision the most permissive test here. Now effective. This TIGHTENS both bars, which is why it was left standing until asked for; the ensemble gate's own note already recorded that every chart clears it by 6.5-12 sigma on effective calls, so the measured cost is nothing. THE RIDE WAS PRICED AT BAR CLOSES NO ORDER CAN FILL AT. Entry was the close of the bar the vote was formed on. Direction() runs on the first tick of the NEXT bar and the market order fills there, which is that bar's open; the exit is read from closed bars and acted on one bar after the flip. So the book credited every ride with two bar gaps and called it the trade the EA places. Entry is now the next bar's open and the exit the open after the flip - the same event, at the price the account gets. RIDES LONGER THAN THE 200-BAR CAP WERE DROPPED, not capped: one-way survivorship against the longest winners a trend-riding label has. A ride whose bars existed and simply had not flipped is now marked to market at the cap. Only one that ran off the leading edge of loaded history stays unresolved, which is the one case genuinely not knowable. Both label changes re-key: TGT:LEG1 -> TGT:LEG2. Every model retrains from era 0. Also: the minimum-stop floor wrote an un-normalised price (TCAdjustStops normalises only what it widens, so a legal floored stop reached the trade layer off-tick); and the journal's context window was 180 seconds, sized for the dead sixty-second order expiry, while an entry window is counted in BARS - so any fill later than three minutes silently lost the context column the per-pattern adaptation is built from. Compiles 0 errors, 0 warnings. Smoke-tested on EURUSD H1 over 2026-08-24..09-01: runs clean, no runtime errors, fingerprint reads TGT:LEG2:10:D12 in situ, and nothing trades - which is the re-keyed label refusing the stale models, as intended. Not deployed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
84d0af6e91 |
fix(vote): the classic modules were the live book - inputs, not voters; certify with the live stop
The sweep after "still not profitable" found the EA was not trading the strategy it certifies. THE VOTE. Direction() aggregated all children in one accumulator: 15 classic modules at the stdlib weight of 1.0 each, the four AI members at trust weights of 0.17-0.22, and the derived threshold (1% on most charts) had been certified on the AI members' vote alone. Four unanimous networks netted 1.01; two classic modules confirming a state (10 + 10) netted 1.27 and opened the trade with the networks silent. Proven in the tester on the old build: EURUSD H1 from 2025-09-01, threshold 1% published at 00:05, market buy 3.5 lots on the first bar, then vote magnitudes of 4-6 that the networks cannot produce. The book that traded was the classic consensus this project measured at chance; the certified book could barely open. - ClassicVotesMoveMoney(): classic modules stay in pass 1 (journal, raw arrows, the vote vector the networks read) and leave the money sum and the overlay whenever an AI member exists. With no AI member they remain the book. Announced once. - CExpertSignalAIBase::LiveVote(): the parent sums the AI member's certified contribution (module weight x (tier - chance), clamped) instead of its raw tier weight. - VoteCapableWeight() uses LongCondition's readiness test, m_deployedLive included: a deployed member had numerator and no divisor share, which with the classics gone would have been a division by zero live while the inference-only tester looked fine. THE BOOK. The ride book the gate judged carried no stop; every live position carries one at the published mean adverse excursion. The verdict now gates the ride under that stop (sideG*: a ride whose adverse excursion reached it pays -stop) and prints both. Sweep line: [book|stopped|cost]. THE REST OF THE SWEEP. - Scheduled close-all: a +-1 minute window with no catch-up, and Processing() ran on the same tick with the cached vote, so it could re-open five minutes before the weekend. Now a per-day latch from target-1 min, retried every tick and timer, and OpenPosition refuses while latched. - News filter: an empty CalendarCountries() answer (the base still synchronising) was cached for the session, leaving the filter inert with EnableNewsFilter true. Not cached any more. - NF_MinImpact default HOLIDAYS vetoed 50% of EURUSD weekday hours (36% GBPUSD, 38% USDJPY, 27% USD-only), measured on the terminal's own calendar export; HIGH vetoes 14/12/11/9%. Default -> HIGH. Charts attached before keep their stored value. - One m_tradeOwner for both books delegated the short book's exit to a long-only setup, whose default CheckCloseShort fell into the base vote exit on the child (m_direction EMPTY_VALUE = DBL_MAX >= any threshold): every short died the tick after the inside bar armed. Per-side owners, and the stdlib's EMPTY_VALUE guard restored in CheckClosePosition. - One expiry clock for two books -> per-book m_bookExpiration[2]. - The inside bar's time stop selected the lowest-ticket position on the symbol -> by book magic. Build tag inputs-not-voters-1. Compiles 0 errors, 0 warnings. Not deployed. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
9cfcb5f42d |
fix(news): export the calendar on the upkeep tick, not only at attach
NEWS_REFRESH_SEC documents an hourly refresh, but g_newsExport.Update()
was called only from WarmExternalData() in OnInit. The exporter would
therefore have run once per attach and never again, and the calendar is
exactly the file that keeps changing after startup - a release lands and
the terminal's base gets its actual. OnTimer's alt-data upkeep block was
already calling the alt-data fetcher under the same tester guards, so the
exporter joins it there. Its own hourly throttle makes the call cheap and
means the two cadences do not have to agree.
Also lands the per-tick order expiry clock (CExpertCustom::ExpirePendingOrders)
that was still sitting uncommitted, with its reasoning corrected: it argued
from a "sixty-second order", and the entry window became a count of BARS in
|
||
|
|
0dcad40f38 |
feat(news): export the terminal's own economic calendar - the two news setups become buildable
The operator found a paid MQL5 product that exports historical news, and drew the right
conclusion: MT5 carries the calendar itself and exposes it to MQL5, so the EA can export it the
same way it already maintains alt data. Verified before writing anything - a probe script
compiled clean against CalendarValueHistory / CalendarEventById / CalendarCountryById and the
full MqlCalendarValue field set on this terminal.
THE CONSTRAINT THAT SHAPES IT: the calendar API returns nothing inside the Strategy Tester. So
this follows the alt-data contract exactly - a LIVE chart writes
Common\Files\Warrior_EA\News\calendar.csv, and backtests, research scripts and models all read
the FILE. In the tester it says so once and leaves the existing file alone, because a backtest
must never truncate what a live session wrote.
WHAT IT UNBLOCKS, and it is three things rather than one:
1. The two NEWS SETUPS - two of the Forex book's six, unbuildable until now. The News Straddle
carries the only explicit R-multiple in the whole book (~8:1).
2. The news AVOIDANCE rule both Walsh books state and we have never honoured: flat five minutes
before a release, nothing new until fifteen after (forex s13). Every backtest so far has
been holding through releases it should have been flat for.
3. News as FEATURES for the networks - time to the next release, its importance, and the
surprise itself (actual minus forecast), which is an axis price data does not carry.
SCALING, STATED HONESTLY. MqlCalendarValue carries actual/forecast/previous as longs and
MetaQuotes documents them as the real value times a million, with LONG_MIN meaning "no value".
That is documentation, not something this code has observed, so every row carries BOTH the
scaled reading and the raw long, plus digits/unit/multiplier - the first live export settles the
convention and nothing downstream has to trust it in the meantime.
Written atomically through AtomicFile, refreshed hourly beside the alt-data fetch.
Compile-verified in _claude_stage: 0 errors, 0 warnings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
b647b745ec |
feat(inputs): Asset_Class - the operator declares which book applies to this chart
There is one day-trading book PER ASSET CLASS: the Stocks book (applied to indices), the Bitcoin book for crypto, and the Forex book. A pattern is only accountable on the class its book was written for. If it also works on another class, keep it there; if it does not, that is not a failure of the pattern and must not count against it - it simply is not traded there. DECLARED, NOT INFERRED. WarriorAssetClass() reads SYMBOL_PATH and falls back to SYMBOL_CALC_MODE, which is a good guess and still a guess: a broker filing US500 under "CFD" would hand the chart the wrong book silently. ASSET_AUTO keeps the detection and any other value overrides it, set when the EA is attached. WarriorAssetClassResolved() is the single place that resolution happens, so the cost model, the logs and any future book gate cannot disagree about what this symbol is. OnInit now states it either way - auto-detected (with a note to set the input if the broker files the symbol oddly) or declared, printing what the detector would have said. Compile-verified in _claude_stage: 0 errors, 0 warnings. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
358c8b74e0 |
feat(signals): CSignalInsideBarGap - the first research-validated standalone book
M60 inside bar, buy stop at the MOTHER bar's high, risk down to the inside bar's
low, stop 1.25R, target 8.0R, 15-minute time stop, and a SIXTY-SECOND ORDER
EXPIRY that carries the edge.
Validated in Warrior_Research (RESULTS.md, tag promising-results-20260906) on
2003-2026 across eight instruments, selected pre-2019 and measured on 2019+:
unfiltered +0.116 R/trade (t 4.57, n 1,203)
with the context rule +0.285 R (t 6.45, n 412), and 2022+ +0.309 - stronger
recently than in sample
The expiry is the finding, not the pattern. By time-to-fill, orders filling in
the first minute paid +0.398 R in sample and +0.116 out; every later fill paid
-0.004 and -0.015. Only ~4% fill that fast. Being an ORDER parameter rather than
a model feature, it is causal by construction and needs no run-time inference.
Maps onto the existing architecture without changing it: LongCondition() fires on
the closed bar, and OpenLongParams() returns a price ABOVE the market, which is
what makes CExpert place a pending buy stop - the order type the research
measured, and one that fills at its own level instead of crossing the spread.
The expiration is set absolutely to 60 s rather than the base class's whole
chart period, because that is the whole point.
Definitions match the research exactly and are pinned in the class comment, since
they are easy to "tidy up" into something that no longer matches what was
measured: entry is the MOTHER bar's high, and valid_test is Wyckoff p118 - volume
below EACH of the two preceding bars, a two-bar lookback and not a moving average.
EnableInsideBarGap DEFAULTS TO FALSE. Unlike the vote modules it sits beside, this
fires at weight 100 and names its own entry, stop and target, so enabling it
changes what the expert trades rather than how it weighs an opinion. Shorts return
0 deliberately: the mirror setup measured +0.054 R against the long side's +0.285.
Compiles clean - the expert built with exactly the same single pre-existing error
(Network.cl resource path, unrelated) and zero warnings, before and after.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
3a925fa1d9 |
feat(features): the classic vote vector becomes model INPUT - meta-labelling wired end to end
The fifteen classic modules now feed the neural nets as features instead of only voting beside them. One signed column per pattern-bearing module, replayed AT EACH HISTORICAL BAR so the network sees what the rule-based primaries said there, not what they say now. WHY FIFTEEN COLUMNS AND NOT SIXTY. Every module OVERWRITES rather than sums (verified across all fifteen), so exactly one pattern weight is returned per call - and the weights are DISTINCT within every module (verified: no collisions anywhere). The signed net vote is therefore a lossless code for which pattern fired and in which direction: RSI +10 rising, -10 falling, +20 oversold reversal, -20 overbought, +-60 divergence, +-90 double divergence, 0 nothing. And it works as a SCALAR, which a categorical code normally does not. Feeding codes to a network as a number line is usually a modelling error - it interpolates between values that have no order - and the fix is one-hot at ~60 columns of width. This escapes that because the ladder was built ordered BY CONVICTION: 10 is a weak confirming state, 95 is rare and strong. Magnitude carries meaning independently of identity, so the net can read signed strength before it ever learns the codebook. Both properties exist ONLY while the weights are frozen priors. If the database is ever allowed to re-score them this column becomes an unreadable moving quantity. SCALED /100 on the way in. The raw ladder runs to +-95 while the rest of the vector sits within a few units, and one column a hundred times wider than its neighbours owns the first layers gradients and the covariance matrix with it - the exact failure that cost a session when a Wyckoff sentinel read -580. THE REPLAY IS THE EXISTING ONE. HistoricalNetVote() already brackets a filter evaluation in a six-field SaveVoteState/EvalShift/Direction/RestoreVoteState sandwich, because Direction() writes the live journaling state and replaying a three-week-old bar without it would overwrite what the live tick just decided and journal history as if it had happened now. Same bracket, new consumer. CACHED ON THE PARENT, not per member. The four AI models on a chart share the same filter objects, so a naive implementation replays fifteen ladders over 179k bars four times - and worse, could let members see different vectors. One computation per bar, served to all four: an ensemble voting on different evidence is not an ensemble. RSI, MACD AND TIME FEATURES ARE BACK ON. They measured 0/1, 0/3 and 0/6 unanimous under the leg-ride label and I cut them last night on that evidence. Under this architecture they are CONTEXT rather than predictors - an RSI-oversold MACD cross is a different trade from a bare one, and day-of-week conditioning is exactly the conjunction a meta-labeler is meant to find. Standalone predictiveness and conditioning value are different questions and the width cut answered only the first. Width is now retrain-forcing through m_neuronsCount as always, so the models re-key themselves. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
095bd27a67 |
feat(signals): restore CCI, Stochastic, WPR, RVI and Parabolic SAR
Five more modules, 16 patterns, recovered from
|
||
|
|
c99020abb7 |
feat(signals): complete the Bill Williams suite - Alligator, Fractals, Gator, BWMFI
Four new modules, 11 patterns, built on the stdlib CiAlligator/CiFractals/CiGator/CiBWMFI
indicator classes and ported to CExpertSignalCustom like the rest.
ALLIGATOR (4). Three smoothed averages of median price, displaced forward - 13/8, 8/5, 5/3, which
is Williams specification rather than a choice. Patterns encode his own reading: ordered lines are
context (10), a lips/teeth cross is the trigger (40), price clear of the lips with the mouth still
opening is the trend running (45), and AWAKENING - previous bar intertwined, this one ordered -
carries the most weight (55) because it times the transition instead of reporting a state that may
have been true for twenty bars. A SLEEPING alligator produces no vote at all rather than a weak one.
FRACTALS (3). A five-bar pattern, and the two-bar confirmation delay is the whole difficulty: a
fractal at bar i is only knowable at i+2, so reading it at i is exactly the lookahead that faked a
+4 sigma reading in
|
||
|
|
3ed053e3a8 |
feat(signals): restore the classic votes as a meta-labelling PRIMARY, with fixed weights
WHAT AND WHY. |
||
|
|
abda14856f |
refactor(features): delete the cross-asset panel - 564 lines and 117 touch points of dead code
IT HAS BEEN INERT SINCE BEFORE THIS SESSION. `EnableCrossAsset` is `const bool ... = false`, the only
thing that ever calls `UseCrossAsset(...)` is Warrior_EA.mq5:819 passing exactly that constant, and
`BuildCrossAssetPanel()` returned at its first line. Its 6 features were never added to the 92, and
its fingerprint tag `|XA:` was never appended - so REMOVING IT CHANGES NO FINGERPRINT AND NO FEATURE
VECTOR. Nothing retrains because of this commit.
WHAT WENT:
System\CrossAsset.mqh 564 lines, deleted outright
the panel build, the per-bar emit block, the feature-table entry (FeatureBuilder)
the member, the accessors, the persistence fields (ExpertSignalAIBase)
three view interfaces + their two implementations
the topology feature-count contribution and the fingerprint tag
the blocking warm-up in OnInit, the sweep-guard rule, the input constant
THE PERSISTENCE FORMAT LOSES A FIELD, and that is safe ONLY because every model was wiped an hour
ago. The .cfg carried a length-prefixed cross-asset pin between the direction-confidence threshold
and the alt-data name list. Both the writer and the two readers (the loader and the cfg walker) drop
it together, so the remaining fields stay aligned - but a .cfg written by the PREVIOUS build would
now have its alt-data pin read from the cross-asset slot. There are no such files.
WHY DELETE RATHER THAN LEAVE IT SWITCHED OFF: it was 117 references across 22 files that every future
reader had to understand before concluding it did nothing - I spent part of tonight doing exactly
that, twice, and got the reasoning wrong the first time (I claimed UseCrossAsset(true) was never
called anywhere; it is called, with a const false). Dead code that takes two passes to prove dead is
not free.
NOT A MEMORY FIX. It was proposed as one and it is not: the panel allocated nothing at runtime. The
27 GB was the indicator tuner (
|
||
|
|
0405f1e8a4 |
feat(training): optimizer is per ARCHITECTURE, and the era cap goes 100 -> 10000
TWO OPERATOR REQUESTS.
1) THE OPTIMIZER IS NO LONGER AN INPUT. One input applied one optimizer to all four ensemble
members, which is the opposite of what a voting ensemble wants - members that fail in correlated
ways average to nothing. Each model class now names its own via the new virtual
PreferredOptimizer(), so the choice lives with the architecture instead of in a global switch.
It also had to stop being an input on this project's own rule: it is retrain-forcing (a field of
the model filename), and retrain-forcing values are not inputs, because MT5 stores inputs PER
CHART and an already-attached EA ignores a changed default - the exact trap that cost a deploy
cycle in
|
||
|
|
d2574f4c3a |
feat(label): measure the horizon frontier - and it says the horizon is NOT the lever
The label horizon has been called "the only lever that raises evidence" for weeks and was never measured. This measures it, from PRICE and the leg replica alone - no network, no training, no era. What it computes per ZigZag depth: legs (which IS the effective sample size, because every bar inside a leg shares that leg outcome), mean leg life (= the label overlap), the ORACLE ride a perfect caller takes, the round-turn cost over the same window, the break-even capture, and the total at 2/5/10% capture. EURUSD, 60k bars: depth 4: legs 8994 life 6.7 oracle 2.055 BE-capture 1.62% @2%=+71 @5%=+625 @10%=+1549 depth 6: legs 6780 life 8.8 oracle 2.556 BE-capture 1.33% @2%=+116 @5%=+635 @10%=+1502 depth 8: legs 5330 life 11.3 oracle 3.046 BE-capture 1.13% @2%=+141 @5%=+628 @10%=+1440 depth 12: legs 3763 life 15.9 oracle 3.865 BE-capture 0.90% @2%=+160 @5%=+597 @10%=+1324 depth 16: legs 2887 life 20.8 oracle 4.581 BE-capture 0.75% @2%=+165 @5%=+562 @10%=+1223 depth 24: legs 1662 life 36.1 oracle 6.347 BE-capture 0.52% @2%=+156 @5%=+473 @10%=+1000 THREE RESULTS, TWO OF WHICH KILL MY OWN FRAMING: 1. Cost never binds. Break-even capture is 0.39-1.62% at every horizon, against rides of 2-8 ATR. The "shorter horizon sells payoff to buy evidence" trade-off this report was designed around barely exists. 2. At the ~2% capture this fleet has demonstrated, the horizon is nearly FLAT: +141 to +165 across depths 8-36, with the SHIPPED depth 12 already within 3% of the peak. Shortening to depth 6 makes it WORSE (+116). The horizon is not the lever. 3. Capture rate is first-order - 2% -> 5% roughly quadruples the total at any depth - and shorter horizons only win once capture is high (at 10%, depth 4 is worth 2x depth 16). The first cut of this report ranked by the ORACLE and therefore picked depth 4 on every chart, which is simply wrong: it credits a horizon with money no model here has ever taken. Ranking now uses the demonstrated capture, and the capture columns are the ones to read. Also fixed while here: LEG_STATE_DEPTH was NOT in the model fingerprint, though it sets where every pivot falls and therefore which leg each bar belongs to, its direction, its ride, and whether it is Buy/Sell/Neutral at all. Two models at depth 12 and depth 6 train on completely different targets and were sharing a key. Now TGT:LEG1:<minride>:D<depth>. The ZigZag replica takes depth/deviation/backstep as arguments defaulting to the shipped #defines, so every existing caller - including the in-situ verification against the stock indicator - is byte-identical. WHERE THIS POINTS: capacity is dominated by WIDTH, not horizon. Depth 12 gives 3763 independent legs; the mask-off change took inputs 180 -> 552, so the first dense layer is ~17,600 weights against 3,763 observations. Halving the horizon buys 1.8x observations; tripling the width already spent 3x. The fix is the redundancy filter - correlation between INPUTS, no label involved - which is designed in project_pca_reduction_plan and never built: measured worst pair |r|=0.999, PCA takes obs/param 0.49 -> 11.9. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
9550f94997 |
fix(ichimoku): the throwaway probe turned out to be the fix - it is now a deliberate primer
The operator was right that Ichimoku is a stdlib indicator like every other, and it is. Three hypotheses of mine died to measurement today: 1. "MT5 will not serve the Senkou buffers" - false; the wrapper reads them fine 2. "the full-history count at a negative start_pos fails" - false; returns 179048, err=0 3. "a second Create corrupts the buffers" - false; after 2 Creates, Total=10, spanA still ok What actually correlates with the fix is the one thing the probe started DOING rather than reporting. While it copied 64 values the feature path still failed every bar with spanA/spanB EMPTY and no chart completed an era. The moment it began issuing FULL-HISTORY CopyBuffer calls on buffers 2 and 3 at init, the feature path started working on all three charts - 0 rejections, 0 stalls, eras in 21-34s, combined vote scoring - with no other functional change between those two builds. The reading: the Senkou plots are shifted kijun bars FORWARD, and that shifted region needs one full-range materialisation before CIndicatorBuffer::Refresh - which asks at start_pos = -offset for m_size values - will serve it. Priming costs two CopyBuffer calls per init. Not priming cost the fleet hours across two days. CAUSALITY IS INFERRED FROM SEQUENCE, NOT PROVEN BY ISOLATION, and the header says so. It is renamed from WarriorProbeIchimokuBuffers to WarriorPrimeIchimokuBuffers and marked DO NOT DELETE AS SCAFFOLDING, because I was one turn away from removing it as spent diagnostics - which would have re-broken the fleet and left no trace of why. Buffer 3 is primed alongside buffer 2 even though nothing reports on it: the feature block reads SenkouSpanB every bar exactly as it reads SenkouSpanA. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
b724fc58a2 |
fix(ichimoku): back ON - the stdlib was never the problem, the double-shift was
I switched Ichimoku off earlier today on the reasoning that "MT5 will not reliably serve the Senkou
buffers", from the symptom that spanA/spanB read EMPTY at every index while tenkan/kijun read fine
on the same bar. That conclusion was wrong. The operator pushed back - it is a stdlib indicator like
every other - and an init-time probe settled it by measuring the thing I had only inferred:
wrapper: BufferResize(39502)=ok BarsCalculated=39502
at idx 5 -> tenkan=ok kijun=ok spanA=ok spanB=ok
raw: spanA CopyBuffer(start=0)=64 err=0
spanA CopyBuffer(start=-26)=64 err=0 spanA[0]=29432.895
On all three charts. CIndicatorBuffer::Refresh reads shifted buffers with
`CopyBuffer(handle, num, -m_offset, m_size, m_data)`, and that NEGATIVE start_pos is deliberate and
works - it is how the forward-plotted cloud region is addressed. The wrapper is coherent and the
indicator serves data.
What was actually broken is the double-shift fixed in
|
||
|
|
08591e9c4e |
fix(fleet): three charts, because the terminal was committing 29.5 GB of address space
WHY VS CODE KEPT CRASHING, and why the terminal died at 03:51 - one cause: terminal64 working set 2,222 MB COMMIT 29,509 MB Against a commit limit of ~47 GB (32 GB RAM + a 15.4 GB pagefile) on a box also running six 1 GB tester agents and fourteen VS Code processes. At roughly 76% committed, every new allocation starts failing. That is the bad_alloc that killed the terminal inside WarriorCPU.dll ( |
||
|
|
70a3c9a719 |
fix(dll+mask): the overnight crash was an unhandled bad_alloc, and the screen verdict is in
THE CRASH. terminal64.exe died at 03:51 and took the night's training with it:
Faulting module name: WarriorCPU.dll
Exception code: 0xc0000409 (__fastfail)
Fault offset: 0x00000000000182ed
NOT a system OOM - 32 GB total, 25.4 GB free afterwards - and not a stack overrun in the
kernels, which are all REQUIRE-checked against real buffer sizes. 0xc0000409 is what the
CRT raises via __fastfail when C++ calls std::terminate, and the cause here is an
allocation that threw std::bad_alloc straight across a __stdcall DLL boundary into MT5's C
code, which cannot unwind it.
CPU_BufferCreate had no exception handling at all: AllocBufferSlot may push_back onto the
slot vector and assign() reserves the whole buffer, and either can throw. One try/catch in
the entire DLL against five throwing allocation sites, no nothrow, no new-handler. So an
allocation failure - a NORMAL outcome when forty members each hold a multi-hundred-MB
adopted pool - killed the terminal instead of returning an error.
It now returns -1 with CPU_ERR_ALLOC, which is the value the function already returns for
bad arguments, so every existing caller handles it. First such crash in 14 days of event
log, and it landed the night the feature width went 288 -> 360; the fix makes that class of
failure survivable rather than fatal.
THE SCREEN VERDICT (mask v8). The measurement pass DID complete before the crash - ten
keep-screen votes banked. Width 360 -> 180.
KEPT rsi 1/1 on both charts that voted at 60 columns, macd 3/3 and 2/3, ichimoku 7/8
and 3/8. Their August verdict was taken under the OLD label and does not survive
re-measurement. Two charts only - the weakest evidence in the mask, and the next
screen supersedes it.
KEPT cumdelta - 4 of 6 columns on ALL FIVE charts at 48 columns and 4/6 again on
NAS100 at 60. Six of seven chart-observations, as consistent as the ma block.
DROPPED sot (1 slot of 28), wyckoffEvent (0/16 on five charts, 0/16 and 1/16 on the two
at 60 - the largest block in the set and where every CONSTANT column appeared),
wyckoffFail and wyckoffBarInv (inconsistent across charts, which for a fleet-wide
mask is the same as absent).
The August verdict was right about the FAMILY and wrong about one member. Screen per
COLUMN, kill per column.
Halving the width also halves the pool footprint that provoked the allocation failure.
RETRAIN-FORCING: mask version and input width both ride the fingerprint.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
2679b02970 |
fix(features): the eight restored toggles are const - chart-stored inputs were pinning them false
The fleet ran a whole cycle at 48 columns while the source said 60, and every log line agreed with the source. MT5 STORES EA INPUTS PER CHART, and an already-attached EA ignores a changed DEFAULT entirely. The five AD toggles took their true default because they were genuinely NEW to the attached build. The three oscillators were introduced in the SAME commit ( |
||
|
|
c6d001d94f |
fix(fleet): a chart whose EA was killed is invisible to fleet expansion - it now revives it
Five of ten charts sat idle for over an hour tonight with no missing chart and no error to
see. Sequence: an unchecked ArrayResize overran (fixed in
|
||
|
|
a225dca7f6 |
feat(features): enable RSI, MACD and Ichimoku for the measurement pass
Completes
|
||
|
|
151f2bc1eb |
fix(pool+vote): unchecked ArrayResize wrote out of range once the feature width tripled
Nine live 'array out of range' errors within minutes of the width going 72 -> 288: TrainingPool.mqh (298,32) on five charts, ExpertSignalAIBase.mqh (3254,24) on a sixth. BOTH ARE THE SAME DEFECT AND NEITHER IS NEW - the wider vector only made them reachable. Each site calls ArrayResize and then writes at the index it asked for, without checking that the resize succeeded. ArrayResize returns -1 on failure; the write then lands past the end of an array that never grew. TrainingPool also asked for an absurd reserve. The hint was 16384 * m_width, i.e. sixteen thousand ROWS - 196k doubles at a width of 12, and 4.7 million (38 MB) at 288, per pool, on a 2013 Xeon running forty members. A reserve proportional to row width turns a wider feature vector into a quadratically larger allocation for nothing. It is now a flat 65536 ELEMENTS, and the result is checked: a pool that cannot grow refuses the row. The ensemble vote buffer checks the two arrays that actually bound the write - they are resized in one block, so any failure in it surfaces there - and drops the bar with one throttled line rather than corrupting memory. Losing a bar reads as low coverage; the alternative reads as anything at all. Found because the restored Wyckoff groups took the width to 288, which is the point of a measurement pass: it exercises the code at a size nothing had run at before. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
8866ff7c3b |
feat(features): restore RSI, MACD, Ichimoku and the five Wyckoff groups for re-measurement
Reverts
|
||
|
|
1016e52621 |
feat(features): mask v6 - drop crossasset and volume, and state what the capacity rule really says
Tag lean-2. Width 96 -> 72. crossasset scored 12 and 17 votes of 24, volume 12 and 13 - the four weakest survivors in the set, and this file already named volume "the block most likely to fall out of the set on the next screen". They were affordable when nobody was watching the budget. They are not now. THE RULE OF THUMB, APPLIED HONESTLY RATHER THAN QUOTED. Classical practice is 10-30 independent samples per parameter; 1 is the absolute floor. Measured live after this cut: chart member params indep obs obs/param SP500 Perceptron 2336 1139 0.49 SP500 Convolutional 2080 1139 0.55 SP500 LSTM/ConvLSTM 272 1139 4.19 BTCUSD Perceptron 2336 1231 0.53 NAS100 Convolutional 2080 867 0.42 Cutting 186 -> 96 -> 72 inputs moved the ratio from ~0.4 to ~0.6. It BARELY MOVED, and that is the finding: the binding term is not the input count, it is the 16-unit minimum first layer. Run it to the end - SP500 has 1139 independent observations, so 10 obs/param allows ~114 parameters, and at a 16-unit floor that is SIX INPUTS. Not 72. SO THE RULE DOES NOT SAY "make the network smaller". It says a network of any usable size is the wrong model at this sample count, and the only statistically supportable capacity here is roughly linear. The linear baseline has already been run on exactly this data: -0.1pp edge at 100% coverage over 116 independent calls. No edge. THREE MEASUREMENTS, ONE CAUSE. The in-sample book is 3-6x the out-of-sample one; capacity is 0.5 observations per parameter; a linear model finds nothing. All three are the same quantity: we have 1,100-3,600 independent observations, and the LABEL OVERLAP is what makes it so few. EURUSD turns 125,322 rows into 3,631 independent ones by dividing by the 34.5-bar leg-ride lifespan. The previous label's lifespan was 5 bars, so that one change cut the effective sample ~7x. The label horizon is the only lever that raises the evidence instead of shrinking the model to fit the lack of it. Nothing here fixes that; this commit just stops spending capacity on columns that were never earning it. RETRAIN-FORCING: FEATURE_MASK_VERSION 5 -> 6 rides the fingerprint, so every chart starts fresh. Verified live - 73 inputs (72 + bias) on all ten charts. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
03ecf11efd |
feat(features): mask v5 - drop alt and fracdiff, width 186 to 96
The mask exists to hold the first-layer budget effN/(width+1) near 60, which is what its
294 to 78 cut achieved. The LIVE keep-screen now reports that budget at 6.1 against a
16-wide floor. An order of magnitude below design, and inside the regime CTopology
already prints a warning for on nine charts.
THE WIDTH TRIPLED WHILE NOBODY WATCHED THE BUDGET. v3 re-added the alt block and v4 added
fracdiff, both KEPT BY CONSTRUCTION so the keep-screen could measure them - a fair deal,
because the screen only reports on columns that are actually emitted. Both measurements
have now come back, and both are empty:
alt -0.15pp ALT-only against price-only +5.08pp, and 60 of its 72 inputs are
mostly-zero repeats of the anchor bar
fracdiff CPCV ablation -0.47pp to +0.75pp against split standard deviations of 1.05
to 3.53pp, sign going both ways across six charts
The deal for a column kept by construction is that it goes when the measurement arrives.
They go. 90 of 186 inputs - 48 percent of the vector - were columns already measured to
carry nothing.
AND v3 GOT THE COST WRONG, which is how it got away from us. Its note claimed the alt
block cost twelve inputs rather than twelve per bar because the window layout zeroes
repeated readings. Zeroing does not remove an input: the first layer still carries and
fits a weight for every one of the 72. The live width proves it - 31 columns x 6 bars is
186, which only adds up if alt is 12 per bar.
Width 186 to 96, budget 6.1 to about 11.8. Still under the 16 floor, so this is a step
and not a cure - but every input removed was measured worthless first, so it costs
nothing to find out.
The alt fetch pipeline stays ENABLED and both blocks keep being BUILT, so re-enabling is
a one-line change and the finding stays falsifiable.
RETRAIN-FORCING BY DESIGN. FEATURE_MASK_VERSION rides in the model fingerprint, so 4 to 5
means no chart can load its old weights and every one starts fresh at era 0 - which is
exactly what the era cap in
|
||
|
|
ba8bc8e1c1 |
fix(training): stop at 100 eras - the run was picking the best of 1337 re-scorings of one window
Tag era-cap-1. MaxErasPerRun 10000 -> 100, and the era cap stops prompting. THE ERA CAP WAS LABELLED A "runaway backstop, not a training control" and left at 10000 because "the plateau ladder decides when a run ends". The ladder does not end runs. PLATEAU_PATIENCE_ERAS is 8, but ANY new best resets the counter, and on a noisy score a new best arrives by luck often enough that the ladder wanders indefinitely. Measured live today: SP500 era 1337, XTIUSD 1126, USDJPY 721, EURUSD 638. WHY THAT IS NOT FREE, and it is the mechanism behind the overfitting the operator has been pointing at all day. Successive eras RE-SCORE THE SAME OOS WINDOW - they add no independent observations at all (project_window_cut_verdict: "effective n at a rung is coverage x OOS_bars / lifespan, computed ONCE, not once per era"). So the checkpoint is the MAXIMUM of N draws of a noisy statistic, and the winner's OOS score is inflated purely by construction. It gets worse with every era the run survives: chart current era BEST era selection family SP500 1337 1323 best-of-1337, book now NEGATIVE XTIUSD 1126 1057 best-of-1126 USDJPY 721 666 best-of-721 EURUSD 638 624 best-of-638, book now NEGATIVE AND NOTHING IS GAINED PAST ~ERA 20. Measured 2026-08-27 across six charts and two runs: every precision-on-era slope under 2.5pp per 100 eras, signs disagreeing across charts AND across runs, and all six still holding their ERA-20 checkpoint at era 66-71. "In-sample error keeps falling; OOS does not follow." THE LIVE NATURAL EXPERIMENT ran itself today. The two charts that restarted sit at era 69 with their best checkpoints at era 20 (USDCAD) and era 24 (XAUUSD) - exactly where that measurement said they would be. XAUUSD books +0.66, second best on the fleet. 100 keeps the burn-in (ENSEMBLE_CHECKPOINT_MIN_ERA 20) plus ~80 eras of candidates, ample for a ladder needing ~24 eras of patience to reach PLATEAU_STAGE_DEPLOY, and cuts the selection family by 13x. It also makes a full retrain roughly an order of magnitude faster, which is what makes experimenting on topology or labels affordable at all. THE PROMPT HAD TO GO WITH IT, and this was caught before deploying rather than after. PromptContinuePastEraCap raises a MODAL MessageBox on any non-tester chart. Reaching a 10000-era cap was rare enough to be worth interrupting for; at 100 the cap is the NORMAL path, so every chart would raise one. Ten charts, ten modal dialogs, terminal frozen until each is dismissed - on a fleet that is not a prompt, it is an outage. It now prints the same text and deploys the best checkpoint, which is what the plateau ladder does anyway. NOT YET DEPLOYED. Every live chart is already past era 100, so a restart on this build would have them all hit the cap at once and ship their existing late-era checkpoints - the very ones this commit argues are noise-selected. The change only pays on FRESH runs, so it wants a fleet reset to go with it. Not retrain-forcing by itself; worthless without a retrain. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
1c329004c7 |
feat(gate): positive expectancy is the whole bar - stop charging spread, stop requiring alpha
Tag expectancy-1. Two operator decisions, both recorded with the reasoning so neither gets
quietly "fixed" back.
SPREAD IS NO LONGER CHARGED. The book being judged is a ZIGZAG LEG RIDE, and a leg is
dozens of times the bid-ask difference. Measured on this fleet, the round-turn spread is
0.008 ATR on BTCUSD against books of +0.17 to +0.52 - about 4%. The operator reports the
same result in SQX, where setting 0, 6 or 60 pips does not move the outcome, and that
their broker does not bill it as a separate line. Spread is still SAMPLED once per bar and
still PRINTED, so the number stays visible; it simply stops gating.
THE CONSEQUENCE, STATED UP FRONT RATHER THAN DISCOVERED LATER: indices pay no commission
either, so SP500, DAX40 and NAS100 now have a cost of EXACTLY ZERO and their test reduces
to "book > 0". That is literally positive expectancy, which is the stated objective.
Per-chart, what stops being charged (from the live init lines):
SP500 0.68 | DAX40 1.73 | NAS100 2.20 -> all become zero cost
XAUUSD 0.55 of 0.60 | XTIUSD 0.09 of 0.11 | BTCUSD 3.58 of 28.00
USDJPY 0.003 of 0.009 | EUR/GBP/CAD ~0.00003 of ~0.00007
ALPHA IS REPORTED, NOT GATED. It gated for roughly an hour. In the operator's words: "We
do not require beating buy and hold either. I already explained the goal : positive
expectancy, simple as that."
THE OBJECTION WAS RAISED AND OVERRULED, which is the right order and is recorded rather
than re-argued: a book paying less than the drift carries the DRAWDOWN PROFILE of the
underlying trend, and a prop account fails on drawdown. Against that - buy-and-hold is
not a strategy a prop account can run, nobody pays a prop trader for alpha, and DAX40
booking +0.17 against an always-long +0.234 still MAKES +0.17. Operator's account,
operator's call.
The refusal branch became unreachable (tradeableOK is now ratesOK AND bookPays, so a
paying book implies deployable) and was converted into a DRIFT WARNING that prints on the
deployable era instead of withholding it. Same information, no longer a block.
THE GATE IS NOW: measurable, covers the base rate, fires both ways, and the book beats
commission. Four conditions, down from six this morning, and every one of them is a thing
an operator asked for.
Not retrain-forcing.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
50f71f42a3 |
feat(gate): measure the IN-SAMPLE book and compare it to the out-of-sample one
Tag is-oos-2. The operator's ask, in their words: "the goal is to train neural networks to recognize the patterns as best they can, deploying once we are not making progress for x number of eras. performance of is and oos should be similar. just like SQX does." The plateau ladder already IS "deploy once no progress for X eras". The IS-vs-OOS comparison did NOT exist: this project scored the out-of-sample slice only. There is an in-sample MSE as a training diagnostic, but no in-sample BOOK to compare the OOS book against, so nothing in the gate could see a model that had memorised its training span. This adds it. SAME RUNG, SAME LABEL, SAME DECISION RULE - the only difference is which bars. The derivation is copied from the OOS branch deliberately (softmax with the as-of leg side, then the prior-corrected adjusted signal) so the two books differ by the DATA and not by the decision rule, which is the whole point of comparing them. ITS OWN ARRAYS, NOT A FLAG ON THE OOS ONES. The vote rows feed coverage, precision, the chance rate, the cost mean and the exact binomial bar - they ARE the deploy gate. Letting in-sample rows into them would corrupt every one of those silently. Seven arrays instead of twenty-two, because only the book is wanted. THE FIRST CUT OF THIS WAS DEAD CODE, and it shipped as is-oos-1 before the fault was found. The in-sample scoring was placed inside the pass-3 loop, which starts at oosCutoff-1 and counts DOWN to 2 - so it walks the OUT-OF-SAMPLE bars only, and the in-sample ones are the higher indices it never reaches. The guard could never be true and the measurement never printed. It now rides its own descending cursor, stepped once per OOS bar, which also means it INHERITS this loop's yielding and can never become the blocking pass that would livelock an era (project_era_slower_than_bar). SAFETY CHECKED BEFORE WIRING, not after: ApplyClassificationSoftmax and AdjustedSignalFromSoftmax were both read for member side effects. Both mutate TempData and nothing else, so scoring an in-sample bar cannot reach the operating-point fit or any OOS tally. The in-sample step runs BEFORE the OOS bar rebuilds its own window, since TempData is shared. No backprop anywhere in it - this is a scorer, and training on these bars again inside it would be a second unshuffled epoch. Also caught before it could mislead: the OOS path adds signedVote RAW, because LiveVoteContribution already carries m_weight x the tier's pattern weight and voteWeight only builds the divisor. The first draft multiplied them again, which would have squared the weight and made the two books measure different things - a false overfitting signal from the very comparison built to detect one. MEASURED AND NOT GATED, same discipline that let the split-half test be rejected on evidence within the hour rather than becoming another closed door. Bounded cost: at most one extra feedForward per OOS bar, strided to ~ENS_IS_TARGET_BARS samples. WHAT IT CANNOT SEE, printed in the line itself so the number is not overread: our other overfitting route is SELECTING the best era and rung ON the OOS slice, and this comparison is blind to it. Only a third slice that neither trains nor selects would speak to that. Not retrain-forcing. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
f1268a16d9 |
feat(cost): the broker charges by ASSET CLASS - commission schedule, detected not guessed
Tag commission-1. Cost_CommissionPerLotPerSide (one number, shipped at 0.0, never set)
is replaced by the operator's actual contract, applied per class:
indices none forex 4 USD per lot
crypto 0.03% notional metals 0.001% notional energy 0.03% notional
CLASS IS DETECTED FROM SYMBOL_PATH, which is what the broker itself organises its tree
by, with symbol-name and SYMBOL_TRADE_CALC_MODE as fallbacks for a flat Market Watch.
Verified in situ on all ten live charts rather than asserted - every one resolved
correctly from its own path (Indices\, Forex\, Crypto\, Energy\, Precious_Metals\).
A PERCENTAGE OF NOTIONAL NEEDS NO CONTRACT SIZE AND NO FX RATE. Commission in money is
pct x contractSize x price x (quote->account rate); the price-equivalent divides by
money-per-price-unit, which is tickValue/tickSize - and tickValue already carries the
same contractSize and the same rate. They cancel exactly, leaving
price_equiv = pct x price for ANY quote currency
so nothing stale or missing can be read. Worth stating because it looks too easy.
THE FACTOR-OF-TWO, and the two branches need it OPPOSITE ways round. The percentage
branch builds the round turn itself, so a per-side quote is MULTIPLIED by 2. The flat
branch hands a per-side figure to WarriorCommissionRoundTurnPrice, which does its own
doubling, so a round-turn quote is HALVED on the way in. Writing them the same way round
would have been a factor-of-four error between asset classes. The operator's figures are
read as the FULL ROUND TURN (Cost_CommissionIsRoundTurn, default true), which is how a
prop contract quotes it; false doubles every figure without editing any of them.
MEASURED EFFECT - COMMISSION DOMINATES SPREAD ON HALF THE FLEET, and the gate has been
charging spread alone until now:
BTCUSD comm 24.42 + spread 3.58 = 28.00 cost was UNDERSTATED 7.8x
USDJPY comm 0.006 + spread 0.003 = 0.009 3.0x
EURUSD comm 0.00004 + 0.00002 = 0.00006 3.0x
GBPUSD comm 0.00004 + 0.00003 = 0.00007 2.3x
XAUUSD comm 0.04 + spread 0.55 = 0.60 1.1x
indices comm 0.00 unchanged
In ATR terms BTCUSD goes 0.008 -> ~0.063 and still clears on a +0.52 book, but GBPUSD
goes ~0.014 -> ~0.033 against a +0.03 book and should now FAIL. That is the correct
outcome: it was the thinnest book on the fleet and it was being charged a third of its
true cost.
The resolved class, both cost halves and the spread sample count are PRINTED at init, so
a symbol filed in an unexpected folder shows up as a wrong class rather than as a
silently wrong number.
Not retrain-forcing.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
728bc647ba |
feat(gate): the book must beat the DRIFT, not just its cost - and the split-half test is rejected on measurement
Tag alpha-gate-1. Second condition added to the ensemble deploy gate:
tradeableOK = voteGate.ratesOK && bookPays && alphaPays
WHY book > cost WAS NOT ENOUGH. book-gate-1 admitted DAX40, which books +0.20 ATR per
call against a 0.042 cost - a healthy 4.8x - while the always-long book over the same
window pays +0.234. The vote earns LESS than passively holding: its alpha is NEGATIVE,
measured -0.057 +- 0.073 over 18 consecutive eras. That is beta sold as signal, and it
stops paying the day the trend turns, which is exactly when a prop account's drawdown
limit is being tested. Neither more training nor a higher rung fixes it - the model has
found the trend rather than the turns - so the refusal now says that in those words
rather than leaving an operator to conclude the model is merely undertrained.
AND THE SPLIT-HALF CONSISTENCY TEST IS REJECTED, ON ITS OWN EVIDENCE. It shipped in
|
||
|
|
7606517dcd |
feat(gate): deploy on the book, not on the significance of the win rate
Tag book-gate-1. The ensemble deploy condition was
tradeableOK = voteGate.tradeable && bookPays
and voteGate.tradeable ANDs in `precision > chance + EDGE_MIN_SIGMAS x SE`. That test is
now REPORTED instead of GATING. Measured on the live fleet, one era per chart:
chart verdict book cost zero-skill alpha precision vs its bar
NAS100 DEPLOY +0.870 n/a +0.163 +0.707 37.8% vs 33.6%
XTIUSD DEPLOY +0.590 n/a +0.087 +0.503 31.7% vs 30.4%
XAUUSD refused +0.660 n/a +0.256 +0.404 34.9% vs 36.4%
BTCUSD refused +0.370 0.008 +0.032 +0.338 24.6% vs 30.7%
USDJPY refused +0.290 0.014 +0.124 +0.166 29.9% vs 30.7%
GBPUSD refused -0.010 0.012 -0.030 +0.020 27.2% vs 27.1%
DAX40 refused +0.180 0.042 +0.234 -0.054 29.3% vs 37.5%
EURUSD refused -0.130 0.018 -0.053 -0.077 23.1% vs 24.1%
USDCAD refused -0.040 n/a +0.070 -0.110 24.4% vs 27.6%
SP500 refused -0.020 0.074 +0.248 -0.268 25.1% vs 32.2%
THREE BOOKS PAYING 20-50x THEIR OWN COST WERE BEING REFUSED, and not for any reason to do
with the model. effN deflates by the 34.5-bar label overlap, so the bar is chance +4.3pp at
USDJPY's 14,584 calls and chance +15.0pp at DAX40's 1,385. BTCUSD was asked for a 14pp
precision edge over chance on roughly 46 independent observations. Nothing produces that.
This was not a strict gate, it was a closed one, and it was closed by call COUNT rather than
by edge.
IT ALSO CONTRADICTED THE GATE NEXT TO IT. WarriorRungBookProfitable's own header argues that
demanding significance "would deploy NOTHING, ever, which is the prove-your-edge trap this
codebase has backed off twice", and gates the book on a POINT ESTIMATE for that reason.
Then tradeableOK demanded significance anyway and overrode it. Two gates, opposite
philosophies, and the closed one won every time.
AND PRECISION IS THE WRONG QUESTION, which this same file already said twelve lines away:
precision "says a turn was CALLED, not that the leg after it paid". BTCUSD calls 24.6% of
turns right and earns +0.37 ATR per call, because its winners are much larger than its
losers. A gate on the win rate cannot see that; a gate on the book can.
NEITHER SUPPLIED BOOK RECOMMENDS A SIGNIFICANCE GATE. mql5book.pdf pp. 1482-9 optimises and
then FORWARD-TESTS, ranking on a criterion built to reward a smooth equity curve - signed by
slope, weighted by sample size. It never asks a win rate to clear a sigma bar, and it is
honest about the yield (of its top 1000 in-sample passes, 323 were profitable forward).
neuronetworksbook.pdf has no cost model in 690 pages. We invented this bar ourselves.
WHAT CHANGED
* SDeployVerdict splits its verdict into ratesOK ("is this a strategy at all" - measurable,
covers the base rate, fires both ways; none of which needs statistical power to read) and
edgeOK (the significance test). tradeable = ratesOK && edgeOK, UNCHANGED, so the member
gate and its ranking are untouched. Only the ensemble reads ratesOK.
* The rung selection reads ratesOK too. Selecting the certified operating point through a
test the verdict no longer applies would have picked the rung by the wrong criterion -
and on a low-coverage chart would have picked none.
* The Sidak family-wise test no longer gates. It corrects the PRECISION z, and precision is
no longer the decision; leaving it in the conjunction would have kept the closed gate
closed through a side door. It is still computed and printed, because a best-of-N maximum
is still a maximum and an operator should see how selected the number is.
WHAT STILL GATES: ratesOK, and the book beating its own measured per-rung cost.
WHAT IS NOW MEASURED AND DELIBERATELY NOT GATED: the SPLIT-HALF BOOK. The window is split at
its median timestamp and the ride book reported for each half at the certified rung. This is
the CONSISTENCY question - did the book hold up across the window, or is the total one good
stretch carrying a bad one - and it is what the coding book ranks on instead of significance.
It prints on every era. When there is fleet evidence that it discriminates a real book from a
selected one, it becomes the deploy condition. Replacing an unreachable bar with an unmeasured
one is precisely how the precision gate closed, so it is not being done blind.
THE HONEST CAVEAT, stated rather than buried: the book is still a maximum taken over
eras x rungs, so it is still a selected number, and at effN ~46 with ~3.1 ATR per-call SD its
own standard error is near 0.46 ATR - BTCUSD's +0.338 is under one sigma. This is a
point-estimate gate, knowingly. The controls that remain are the plateau ladder (this branch
is not reached until the best vote survives PLATEAU_STAGE_DEPLOY-1 warm restarts with no
improvement), the coverage and both-sides checks, and the live expectancy stop - which is the
only one of the three that judges the book the account actually experiences.
Three stale comments corrected to match (the sweep's [DERIVED] explanation, the stage-3
refusal attribution note, and the refusal branch's own condition).
Not retrain-forcing.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|