Commit graph Warrior_EA/research
Author SHA1 Message Date
AnimateDread
5a4b36bb61 Add AFML parts A, B, C and HCC history decoder
- Implemented AFML part A for testing the dip-z book against search artifacts, including PBO, DSR, and CPCV metrics.
- Developed AFML part B to generate time and tick bars from M1 broker data, including return distribution statistics.
- Created AFML part C to build a pipeline for dip-z primary analysis, incorporating features and a random forest model for classification.
- Added HCC history decoder to read and process broker M1 `.hcc` files, ensuring proper handling of data structure and integrity.
2026-09-29 20:33:42 -04:00
AnimateDread
75d7362161 feat(warrior): the vol-gated dip-buy book, ported into Warrior_EA
Warrior's defaults are now the validated book: DIP_ZSCORE alone, long only, H4,
risk 0.25%, one chart per index with a shared Magic.

- System/BarCache.mqh: whole-history closed bars, Wilder ATR, GK sigma and the
  expanding vol percentile (no 1024-bar stdlib ceiling)
- System/AccountGuard.mqh: open-risk cap, kill switch, cross-chart lock and
  Friday flat, shared through terminal globals by Magic
- CWarriorExpert: guard on every tick; a transient open failure retries the bar
- CWarriorSignal::SetupStop: the dip owns its 3 x Wilder ATR stop from the bid
- SignalDipBuy: no entry vote while holding (a still-dipping time exit never
  closed, and Processing re-entered on the exit bar); no entry on a stop bar
- WarriorMoney sizes on equity; WARRIOR_RISK allows fractional risk
- TradeLog + research/compare_ea.py: trade-for-trade check vs WarriorDipZ -
  SP500/US30/DAX40 identical to the cent, NAS100 96.9% (stale-quote timer fills)
- research/nn_cross_index.py: pre-registered cross-index NN meta-label - FAIL
  (AUC 0.564, CI [0.498, 0.630]); DipMetaCut stays off

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 21:14:53 -04:00
AnimateDread
6192393511 research(fx): forex and metals - thirteen registered families, nothing passed
Hypotheses were registered in FX_PLAN.md before each round. Trend,
breakout, cross reversion, hour seasonality, month-end USD, carry-cross
dip-buy, metals dip-buy and flight-to-safety all fail the bar.

The weekend-gap fade looked like the best result of the project on bar
data (OOS t 20, 28/28 pairs) and loses on real ticks (EURCHF PF 0.52,
AUDNZD PF 0.53): the Sunday-open spread is as wide as the gap.
WarriorGapFade is kept as the research artifact that proved it and is
flagged DO NOT TRADE.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 13:24:57 -04:00
AnimateDread
ff46cfd57e research(dipz): the vol-gated dip-buy on four indices - screens, bear test, reconciliation
Recovered live config (z20 <= -1.5, exit SMA20 / 10 bars, 3xATR) plus a
Garman-Klass vol-regime gate. Expectancy is monotone in the vol regime in
IS, OOS and full sample, 4/4 indices; the gate reverses on USDJPY/XAUUSD.
D1 2008-2026 survives 2008/2020/2022 (maxDD 2.8%, ret/DD 7.74); the gate
halves trades, so it belongs on H4, never D1. reconcile.py matches the EA
to the backtest trade by trade; combine_charts.py rebuilds the account
curve from per-chart tester runs.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 13:24:57 -04:00
AnimateDread
9d60289f9f research(vol): meta-label blueprint, and the premise test that sank the 70% forecast
BLUEPRINT.md reviews the feature/label layers and designs fractional
differencing, Garman-Klass / Yang-Zhang targets and a 48-72h expansion
label; mql5_patches/ holds the MQL5 side (FFD safe past the 1024-bar
series ceiling, vol estimators, the label + veto gate, NY-time swap window).

premise_test.py measured the premise on real broker bars: the headline
AUC 0.75 was a day-of-week / path-length artifact (a Friday-clipped path
is shorter, so it touches K*ATR less). Honest residual 0.56-0.62, mostly
within noise once overlap is deflated.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 13:24:57 -04:00
AnimateDread
ee0dc78382 chore(repo): move research/ and references/ out to ..\Warrior_Research
This repo now holds only EA (MQL5) sources. The Python research scripts and the
third-party MQL5/PDF reference material live in a sibling workspace folder,
..\Warrior_Research\, with their own git repo (initial commit ec2214a there).

Nothing in the EA depends on either folder at build or run time, and the research
scripts address ..\Market Data\ and the MetaTrader Common\Files directory by
absolute path, so the relocation breaks no path. EA comments that cite scripts by
name (research/edge.py, research/altdata/export.py, research/test_spread.py, ...)
stay accurate - only the parent folder moved.

.gitignore drops the two rules that only existed for the moved trees
(references/*.pdf, research/edge_rows.npy); they were carried over to the new repo.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-25 11:00:24 -04:00
AnimateDread
4c7c063e77 chore(research): drop the scratch placebo driver, now a sqxrepl subcommand
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-24 13:31:07 -04:00
AnimateDread
5ac948595b research(metafilter): shuffled-outcome control, and correct the "data adds nothing" verdict
The verdict recorded in this docstring - that volume, time and alt all land at +0.019-0.020,
identical to price alone, so none was being used - came from a run whose weekend clock was a
day early and keyed on a bar most feeds never trade. That left two feeds with 74-78% unresolved
trades and sd(R) of 0.21 against everyone else's 0.62, which handed them overwhelming weight in
the inverse-variance pooling.

With the clock fixed the ordering inverts. `geom` becomes the WORST row rather than the
equal-best one, and price+time nearly doubles it:

    price+time +0.047 | price+vol +0.038 | price +0.037 | ALL +0.028 | price+alt +0.027 |
    geom +0.026

So price and time DO add ranking power over the strategy's own entry arithmetic. What survives
both versions is the alt result: every set containing alt columns scores below the same set
without them, agreeing with the direction screens at H4 and D1.

`shuffle` permutes the training outcomes while leaving fold boundaries, purge, threshold rule,
kept fraction and scoring identical. A lift that survives that comes from the machinery, not
the data - and this session has already produced two results that did exactly that, so the
+0.047 does not get believed until this run comes back near zero.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-24 13:16:40 -04:00
AnimateDread
a86019c46a fix(edge): save_rows walked only the top level, and my first fix was the reason
The encoder was rewritten once to convert every value rather than the one field known to hold
an array. It still only looked one level deep, and `evaluate` attaches the ENTIRE built matrix
under a 'd' key - so it stepped past d['X'] and killed the last line of a second 40-minute run
with the same TypeError the first fix was meant to end.

plain() now recurses into dicts and lists. The matrix and its column index are dropped rather
than converted: they are working state, tens of MB per instrument, reconstructible from
features.build, and nothing re-analysing these rows needs them.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-24 13:12:21 -04:00
AnimateDread
1c2030bcb8 research(sqxrepl): measure the hidden long-beta, as a beta subcommand
The portfolio case for this strategy is that more uncorrelated strategies alongside it beat
an index. That only holds if the members do not share a hidden common factor, and a backtest
correlation matrix cannot see one: it is dominated by the calm months that make up most of a
sample, while the shared exposure surfaces in the month that breaches a drawdown limit.

So measure it directly. Aggregate each strategy's R by month, correlate against the
underlying's own monthly return, and split into up-months and down-months where a long bias
actually shows.

The answer is not marginal: mean correlation +0.69, positive on 14/14 feeds across crypto,
indices, energy, FX and metals, R2 up to 0.62 on USDJPY. Mean monthly R is +1.94 in up months
against -1.61 in down months. Every vendor pair agrees to within 0.03. These are not seven
independent bets, they are one bet placed seven times.

The placebo decomposition says why: across eight markets the barrier term is close to the
negative of the drift term (BTCUSD +0.055/-0.041, SP500_d +0.083/-0.030, USDCAD_d
-0.056/+0.059). The stop and target are a trend-capping device - they clip the gain where the
asset rises and limit the loss where it falls - so the residual cannot be an edge. It is the
same exposure with both tails trimmed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-24 13:06:19 -04:00
AnimateDread
5b33c8396b research(sqxrepl): three-arm placebo, and a placebo subcommand to run it
The two-arm version answered only half the question. It showed the MACD entry beaten by random
timing on every feed, but left the residual +0.01 to +0.05 R unexplained, and the first story
built on it - that the strategy is a drift harvester - died on the second market.

The drift arm settles it by running the same random entries with the stop and target moved out
of reach, so every trade holds to the weekly close. Its R is then rebuilt by hand on the REAL
stop distance, because simulate divided by the widened one; leaving that alone would report
every drift trade as ~0 R and make the comparison vacuous. real-timing is what the signal is
worth, timing-drift is what the barriers are worth over just holding.

The driver moves into this module as a `placebo` subcommand instead of living as a loose script
beside it, and reports per-market statistics for both terms.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-24 12:54:44 -04:00
AnimateDread
cc68a768b1 fix(sqxrepl): the "buy and hold" benchmark shared the strategy's own exit bar
It was reported in every table as the drift this strategy inherits, and as the column that
isolates what the RULES contribute. It is neither. buy_and_hold_R marks each trade at
exit_idx - the bar its own barrier fired on - so a trade that exits at its target is compared
against the close of the bar that touched the target. The two are nearly the same number by
construction.

Measured, the per-trade difference has sd 0.10 against R's own 0.62. That is where a
per-market t of 8.23 and p=0.00017 came from: a quantity that mostly cannot vary will always
look significant. Every "skill over buy-and-hold" figure quoted from this module is withdrawn.
What the column legitimately shows is exit slippage, and it now says so.

placebo() replaces it. Same number of entries, same previous-day-low level, same ATR-scaled
stop and target read at the entry bar, same weekly close, same non-overlap, same fill engine -
only WHEN the orders are placed moves, drawn from the bars the rule could have fired on so the
null inherits the same calendar exposure. Geometry then appears in both arms and cancels, and
only the MACD timing is on trial, which is the question that was being asked all along.

The run also now prints whether any vendor PAIR disagrees in sign. Before the weekend-clock
fix SP500_d and SP500_5 disagreed (-0.003 against +0.064) and so did the two FTSE feeds; they
are the same market at correlation >= 0.999986, so that disagreement was evidence of a
machinery fault and nothing noticed it. Now it cannot pass unremarked.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-24 12:42:41 -04:00
AnimateDread
37764981ea fix(sqxrepl): the weekend clock was a day early and keyed on a bar most feeds never trade
Two bugs in one function, and together they invalidate every number this module has printed.

1. `(days + 4) % 7` makes Friday 5 and THURSDAY 4, so every `dow == 4` test matched Thursday.
   Trades were force-closed a day early and the no-entry window blocked Thursday night to
   Saturday night. Epoch day 0 is a Thursday, so Monday=0 needs +3. Verified against a known
   calendar week instead of re-derived by argument.

2. The close was keyed on a literal 23:45 stamp. That assumes every feed trades up to it and
   they do not - FTSE_d has 175 such bars in its entire history against XAUUSD_d's 13,657,
   because an index CFD session closes hours earlier. Most FTSE trades found no close ahead of
   them, and a trade past the last close bar got a negative horizon that maximum(_, 1) turned
   into a ONE-BAR hold: a silent instant exit indistinguishable from an ordinary unresolved
   trade. The close is now the last bar of the trading week, which is feed-agnostic and is
   what 'flat for the weekend' means.

The tell was in the diagnostics, not the result: SP500_d and FTSE_d showed 74-78% unresolved,
sd(R) of 0.19-0.21 and 1.1-hour holds while every other feed sat near 0.62 and 20 hours. Those
two carried a third of all trades and, having almost no variance, dominated the inverse-
variance pooling - which is where metafilter's implausible t_mkt of 9 came from.

Corrected, the five feeds agree: 22-34% unresolved, sd(R) 0.61-0.64, RR 0.25-0.36, win 67-74%,
and expR still positive on all of them at 1bp/side. The finding survives; its statistics do not
and are being re-run.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-24 12:35:02 -04:00
AnimateDread
061128a88c research(metafilter): correct the claim that the geometry control carries no information
It was introduced as 'no market information whatsoever'. That was wrong. With g = fill minus
the previous day's low, rr = (6.51*ATR25 - g) / (g + 3.23*ATR15), a monotone decreasing
function of g/ATR - so rr is a NORMALISED DISTANCE ABOVE YESTERDAY'S LOW, a price feature in
the same family as donch and smadist, reparameterised until it looked like bookkeeping.

What survives the correction is the part that matters: volume, time and alt add nothing, every
combination lands at +0.019 to +0.020, and price alone already reaches +0.020. What changes is
the explanation - the lift is one price relationship, not an absence of one.

And the relationship is not monotone, so 'prefer a better payoff ratio' is the wrong summary.
By decile on SP500, XAUUSD and USDJPY alike it is an inverted U: filling far above the low
pays ~0, the middle band (rr 0.13-0.40) pays +0.05 to +0.14, and filling AT the low is
negative on all three. Buying the level the strategy aims at is the losing case.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-24 12:31:04 -04:00
AnimateDread
22e75741b0 research(metafilter): the geometry control, which turns the finding inside out
First run returned +0.020 R lift at t_mkt 6.5 - an order of magnitude beyond anything else in
this project - and the tell was in the same table: price, price+alt and ALL returned the SAME
lift to three decimals. A model given more information that does exactly as well as one given
less is not using the extra information, so whatever it found was in something all five sets
shared.

`geom` is that something: two columns of the strategy's own entry arithmetic, the realised
reward-to-risk ratio and risk as a fraction of price, both known at entry and carrying no
market information at all. It scores +0.022 - the LARGEST lift in the table - and ALL+geom at
+0.020 is no better. Every data family contributed nothing, which is exactly why they all
agreed.

The cause is the entry. A buy stop at the previous day's low sits below the market, so it
fills a median 4.75 ATR from the level its stop was sized against and the reward:risk of each
trade is close to arbitrary. The model was ranking that, not the market.

Also adds cut_from='train'. The keep-threshold was taken from the TEST fold's own prediction
quantile, which keeps exactly q by construction but cannot be known in advance - so the filter
as first measured was not implementable. The training quantile is fixed before the fold is
seen and lets the kept fraction float, which is the version that could be traded.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-24 12:27:04 -04:00
AnimateDread
5c3bd493ed research(metafilter): ask whether the data can RANK trades, not call direction
edge.py asks a question that is mostly closed in this project: can a model call direction on
a symmetric barrier. Filtering is a different and easier question - the rules have already
chosen the side, and the model only has to rank trades that were going to be taken. A series
that cannot say 'up or down' can still say 'not today', and nothing built here so far could
have detected that.

For each trade the strategy takes, the feature vector is read ONE BAR BEFORE the decision bar
- every column in features.py is a function of bars <= i including i's own close, and the
order goes in at bar i's open, so reading row i would hand the filter the outcome of the bar
it is deciding on. A gradient-boosted regressor predicts R under a purged walk-forward, the
top q of each test fold is kept, and the lift is measured against the mean R of ALL trades in
those same folds.

That baseline is the one that cannot be gamed by the strategy being good: a random subset of
the same size has expected mean equal to the fold mean, so the difference is exactly what the
ranking contributed, and a profitable strategy raises both columns together rather than the
gap. Families are ablated as in edge.py and the verdict is read from the per-market t.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-24 12:15:28 -04:00
AnimateDread
aa53d5d2ca research(sqxrepl): replay the generated SQX strategy through the validated fill engine
A generated strategy is not the shape edge.py measures. That screen asks whether direction is
callable on a SYMMETRIC barrier; this one is long-only with a target near twice its stop, so
it can pay at a win rate well under 50%. 'Direction is at chance' and 'this makes money' are
not in conflict - they are different measurements, and the way to settle which applies is to
replay the rules rather than argue from the screen.

Both readings of the entry are implemented because they are not the same strategy. Taken
literally, a BUY STOP at the previous day's LOW sits below the market, triggers at once and
fills a median 4.2 ATR above its own level - 91% of the time - so the stop loss is measured
from a level the trade never touched and realised reward:risk lands at 0.28 rather than the
~2.0 the coefficients imply. Read as a pullback (LIMIT), the geometry comes out at 2.1 as
designed and only 40-60% of orders ever fill. A trade export decides which one the generator
ran; nothing else can.

Also carries a correction. The docstring first claimed the omitted trailing stop could not
bind because activation sat far out. Measured, activation is at 0.54R - it arms before the
trade is one unit of risk in profit. The claim was wrong, the number is now printed every
run, and the omission is recorded as the largest deviation rather than a small one.

edge.save_rows writes to a temp file and renames. The first D1 run crashed mid-dump and left
a truncated JSON at the canonical path, which is worse than no file.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-24 12:09:52 -04:00
AnimateDread
2c68314526 research: count markets, not feeds - and stop throwing away a 40-minute run's rows
Every pooled statistic here was reported as a bracket: SE_INDEP, which claims 28 cells are 28
independent tests, and SE_CORR, which claims they are one. Neither is the number. The 15 feeds
are 7 markets - six PAIRS entries are one instrument quoted by two vendors at correlation
>= 0.999986, and ES_fut, SPY_d1, SP500_d and SP500_5 are four claims on the same index - so
market_t() collapses each market's cells by inverse variance and takes a plain t across the
market means. Its degrees of freedom are separate price series, which is the only n this
catalog can defend, and it is now the column the verdict is read from.

catalog.MARKET is where that collapse lives, next to PAIRS, because it is the same fact.

save_rows() writes both arms' scored cells, per-trade diff vectors included, to a JSON beside
the bar cache. The screens cost ~40 minutes and produced nothing but a printed table, so
re-pooling, collapsing feeds, or reweighting a threshold meant paying for every fit again -
which is why the H4 run's rows are gone and it has to be re-run to get them back.

Also folds two spellings of the Market Data root into catalog.ROOT; sqxbars had its own copy.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-24 11:30:54 -04:00
AnimateDread
76dfa223d3 research(edge): one timeframe-scaled table of sample-size floors, and k as an argument
Six bare literals across four functions all encoded the same judgement - below this many
rows, a fit or a score is noise - and all six were absolute counts tuned on H4. D1 has six
times fewer bars per year, so a D1 run would have dropped almost every instrument from the
sample without saying it had: build wanted 3,000 labelled rows and no D1 series but SPY has
that many, and evaluate wanted 5,000, which nothing has. The run would still have printed a
table, just a much emptier one, and the emptiness is exactly the kind of thing that reads as
a null result.

floors_for(tf) scales them by bars-per-year so the judgement stays 'this many YEARS', with
clamps so the coarse end cannot scale down into a sample no statistic survives. FLOORS is
set once at the entry point, after the timeframe is known, and printed with the run.

k and the barrier window join it as arguments. k was pinned at 2.0 in three places while the
banner claimed 'k=2' unconditionally; at D1 that same k resolves in ~11 days and leaves ~250
independent trades per instrument, which is a different experiment from the H4 one and needs
to be requested rather than assumed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-24 11:27:27 -04:00
AnimateDread
7106504061 research(edge): matched-fold A/B on training breadth, and three fixes that changed the answer
The screen now runs two arms that differ in exactly one thing - what the model was
trained on - and compares them cell by cell.

THREE CORRECTIONS, each of which biased toward finding an edge:

  The pooled arm's STATIC BENCHMARK was fitted globally on all instruments' training
  rows. Any instrument drifting against the pool got an "always long" benchmark while
  it was actually falling, handing the model a win it had not earned. Now fitted per
  instrument on that instrument's own purged rows.

  The two arms used DIFFERENT FOLD BOUNDARIES - the pooled arm cut on calendar time
  (it must: the instruments have different bar counts and start dates), the solo arm
  on row index. So any difference between them mixed "pooled training helps" with
  "the arms saw different years". Both now share calendar folds, and the solo purge
  moved from bar index to exit TIME to match.

  The pooled model was REFIT PER THRESHOLD. The fit does not depend on the threshold,
  only the walk does, so this doubled the cost of the most expensive arm for an
  identical model. One fit now serves all thresholds.

Dropped the EMBARGO constant: purging on exit time already keeps a training row only
if its trade had closed before the test window opened, and a bar-count embargo on top
is a second, weaker statement of the same rule.

RESULT, on identical columns, folds, purge, benchmark and scoring:

  price       pooled +1.18pp   per-instrument -0.62pp
  price+vol   pooled +1.33pp   per-instrument -0.83pp

Per-instrument training is NEGATIVE on every feature set. Nothing clears a defensible
bar in either arm - the best of 140 cells reads t=2.79 against a Sidak bar of 3.56,
which is what the maximum of a null grid looks like - but the GAP between the arms is
the largest effect in the run, and the EA trains a net per chart.

Also: alt data is negative in every alt-containing cell of both arms. Six macro
columns against ~1,900 independent trades buys overfitting, not information.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-24 11:10:02 -04:00
AnimateDread
4cd6841c07 research(edge): combined price+volume+time+alt screen, built so drift cannot pass as skill
features.py assembles one causal matrix per instrument - 18 price columns in ATR
units, 4 volume ratios, 5 time encodings and the alt block (COT positioning, VIX,
curve, dollar, breakevens, Fed funds, and the instrument's own implied vol where one
exists) - plus a symmetric k*ATR first-passage label. edge.py evaluates it.

Four things it does that the naive version does not, each of which has already cost
this project a retracted result:

  BENCHMARK    Gold's up-first rate is 53.49% over 23 years against a 51.00%
               break-even, so "go long" alone looks profitable. Skill is the excess
               over the best STATIC side fitted on the training fold, and its t is
               PAIRED - the static walk takes the same trades, so drift, regime and
               sample composition cancel. An unpaired t against break-even is the
               drift's t, not the model's.
  PURGING      Barrier labels stay open for many bars, so training rows whose trade
               exits after the test window opens are dropped, plus an embargo. In the
               pooled arm the purge is on TIME, not bar index - the instruments have
               different calendars and index-purging would align 2015 with 2021.
  INDEPENDENCE The scorer walks each fold sequentially - take a signal, jump to that
               trade's exit, look for the next - so it counts what an account could
               have taken instead of counting the same swing once per bar.
  NULL         A rotation null is provided for the best-of-N problem: rotating the
               fitted predictions against the labels keeps both series'
               autocorrelation and destroys only their alignment.

Two design errors found and fixed while building it, both worth keeping visible:

  The alt `_na` missing-flags were a DATE PROXY - each flips once at its series'
  first publication, so a tree reads "before 2010" and fits that era separately.
  `alt_mode='restrict'` (now the default) keeps only the published era and carries no
  flags. It matters: gold's price-only cell went from t=3.55 to t=0.35 under it, so
  that apparent edge lived entirely in the pre-2010 sample.

  Coverage was reported as trades/bars, which reads 8% for a walk that is actually
  taking ~90% of every slot available. Non-overlapping trades make the ceiling
  bars/duration, and the honest number says there is nothing left to take.

Also fixes a regression this change introduced: `book._resample` divided the summed
spread by `v` to get the mean, which was correct only while `v` was a bar COUNT.
Volume is now real (carried from the .dat's sixth field rather than discarded), so
that divisor is now the bar count explicitly - USDJPY H4 reads 0.640 bp against the
catalogued 0.64. Frame.has_volume distinguishes real volume from a bar-count proxy,
because a "volume feature" built on the latter is measuring session length.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-24 10:35:31 -04:00
AnimateDread
c5890ef17f research(detect): the ratio was read backwards - rebuild around a decision number
The first version reported `required / floor` (the spread's win-rate hurdle over the
smallest uplift the sample can resolve) and called a LARGE ratio cost-bound,
concluding H1 was untradeable on all 14 instruments. That inverts the meaning. A
large ratio means the hurdle sits many standard errors away, so a break-even-sized
edge would be seen at overwhelming significance - USDJPY H1 read 5.89, which is a
break-even edge showing at ~12 sigma. That is a well-POWERED cell. The bad case is a
SMALL ratio: cost cheap, nothing measurable.

The tell was in the same row and went unchecked: it also said the spread was 4.5% of
a 1-ATR stop and break-even was 52.25%. Neither supports "cannot pay its spread".
When a derived ratio disagrees with the raw quantity it came from, the raw quantity
wins.

Rebuilt around the number that actually decides whether to act:

    CONFIRM = 50 + 50*costR + 2*SE

the win rate a symmetric kR:kR setup must hit to be PROVABLY profitable. The two
terms pull opposite ways in trade size - widen the stop and the spread shrinks as a
share of the move, but each trade eats more history so SE rises - so CONFIRM is
U-shaped and its minimum is the cell worth testing first. This also makes explicit
that the horizon axis and the stop-multiple axis are the same axis: a 20-bar hold on
H1 is an H4 trade, and the table now prices both.

Durations are MEASURED by walking each trade to its first barrier touch rather than
assumed to follow the diffusive k^2 scaling, which is off by an instrument-dependent
factor. Trades open at the window cap are reported, since they make n_eff optimistic.

Result on the four longest histories: the best cell needs 51.9-52.9% and the surface
is FLAT from H1 k=2 to H4 k=2. There is no magic horizon. Add commission and the
working target is ~53% - against a deploy gate that asks ~66% at 18% coverage purely
because its OOS window holds ~63 independent observations. The binding constraint is
the gate's window, not the market.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-24 09:47:55 -04:00
AnimateDread
4d726c743e research(data): one catalog for 15 instruments, and the detectability map
The tick .dat files and the bidask/bars/sqxbars caches are gone, so every screen
that reached for fills.Book was dead and breadth.py's five symbol keys no longer
matched the SQX export (it has been rewritten with ONE underscore, turning every
lookup into FileNotFoundError). Rebuilt the data path from the only source left -
the SQX bar files - and pointed it at a local copy so research never reads the
live SQX install.

catalog.py is now the single place that says what an instrument is: path, asset
class, synthesised spread, and the price range that PINS the decimal scale. The
scale used to be fitted against an MT5 reference series that no longer exists, so
it is now asserted per file instead of inferred, on three checks that agree on 1e6
for every file - price level, round tick GCD, and the medians recorded when the
decoder was last validated at corr 1.000000 (FTSE 7246, WTI 65.4, USDCAD 1.26 all
reproduce exactly). breadth.py's two duplicated dicts are gone; catalog owns it.

Validation: five of the six dual-feed pairs agree at corr >= 0.999986 with a
sub-basis-point median difference. WTI is the exception at corr 0.9988 / -14.7 bp,
because the vendors roll the continuous contract on different days - so the WTI
pair is NOT a clean replication arm and must not be quoted as one.

detect.py answers the question the closed verdicts never did. "No edge" has two
opposite causes that look identical in a results table - the effect was smaller
than the spread, or the window could never have resolved it either way - and only
the second one is fixed by more data. So it computes both bounds per cell: the
win-rate uplift needed to pay the spread, and the uplift that is distinguishable
from chance on the trades the history actually holds.

Also fixed, all found by running the above:
  book.py       W1 buckets, with the 4-day offset the epoch's Thursday needs, or
                every weekly bar would straddle a weekend
  breadth.get   serves each symbol's finest AVAILABLE base series and refuses to
                resample upward rather than inventing intrabar highs and lows
  catalog.keys  filtered on M1, which silently dropped the futures tree (M60) and
                the 33-year SPY series (D1) from every screen that iterated it
  breadth.cells one generator, shared, that skips the unbuildable tick-derived arm
                instead of dying on it - was duplicated in two test scripts

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-24 08:59:58 -04:00
AnimateDread
120afde2a3 research(altdata): H4 timeframe screen - information survives, diluted; costs 2.6x worse
Same harness as the D1 screens, forward 30 H4 bars (~5 days), 199 perms, on
the surviving htf mid bars (19.8k-36.4k bars per symbol). Question: does the
wired alt/volume information carry to H4, for the chart-timeframe decision.

Answer: the signal survives but is roughly halved, and the cost side worsens
2.6x per step down. SP500 vix_chg5 clears the family bar with MI|vol 0.0112
(vs 0.031 at D1); volLevel50 (the EA activity feature) is incremental on all
four symbols at H4 - USDJPY family-clean, and it is EURUSD strongest
non-control signal there too. Gold gvz_chg5 stays incremental (0.0038 vs
0.0197 at D1 - a fifth of the strength). COT is null at H4 on FX (weekly
cadence pasted across 30 bars/week dilutes it below detection on EURUSD/JPY;
survives conditionally on SP500).

Cost table (median spread/ATR; the 1.74xATR geometry in spread units):
SP500 138->52->25, EURUSD 282->113->57, USDJPY 230->87->44, XAUUSD 92->34->16
for D1->H4->H1. Every step down multiplies the cost share ~2.6x.

Also fixes the disaggregated-COT column name for gold/WTI in the screen
(M_Money vs Lev_Money - the same crash the EA-side catalog documents).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-16 17:49:37 -04:00
AnimateDread
40ddf06ce6 docs(altdata): adjudicate the NASA API trio - POWER queued behind NATGAS, GIBS unconsumable
POWER is the real find of the three: daily temperature -> degree days ->
natural-gas demand is the textbook gas fundamental, numeric and daily. But it
is point data needing construction into a national series (NOAA CPC ships that
ready-made), and its target symbol is not traded yet - fetch code written for
a chart nobody attaches first runs months later, unobserved, which is the
silent-FRED failure shape. Queued for the AvaTrade expansion, not refused.

FIRMS: re-raised, nothing changed since it was parked - point fire detections
behind the same unproven proxy chain. GIBS: imagery tiles, not numbers; our
CONV is 1D and NASA already sells the extracted products.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-16 17:35:58 -04:00
AnimateDread
626b591ca6 feat(altdata): wire everything the sources serve - screens become priors, not gates
Owner decision (stated twice): available data gets wired; the networks judge
usefulness; the deploy gate remains the arbiter of what trades. Implemented:

  MACRO block (6) on every symbol: 10y yield 20d change, curve slope, 5y
    breakeven 20d change, Fed-ECB policy gap, CPI yoy, unemployment 12m change.
    Screened null vs forward range on all four research symbols - recorded as
    the honest prior in the catalog comment, wired regardless.
  RISK block (3) extended to every symbol (FX majors, metals, energy, BTC all
    now carry vix/vix_chg5/usd_chg5).
  IVOL pair extended with the level alongside the change.

Vintage integrity kept where it is free: CPI is fetched as CPIAUCNS (NSA,
essentially never revised) so the plain-FRED backfill stays first-print-clean;
yields/curve/breakevens/policy rates are unrevised by nature. UNRATE is the
one exception (seasonal refits, ~0.1-0.2pp) - the EA cannot run the ALFRED
protocol, accepted and documented at the declaration site.

UpdateFred gains a staleDays parameter so the monthly series do not fire a
pointless fetch attempt every hour for three weeks after each print.
FeatureValue now takes the day and does its own as-of lookups - adding a
source no longer widens a parameter list. Feature counts: 12-15 per symbol;
symbol feature-order changed, safe only because no models exist yet.

export.py mirrors the new catalog for the five research symbols (13-15
features), smoke-tested: all five CSVs written, 6,072 daily rows each.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-16 17:29:22 -04:00
AnimateDread
3b77500726 research(altdata): macro/rates/country data is NULL vs forward range on all four symbols
Screened yields (DGS2/DGS10), curve slope, inflation breakevens, Fed policy,
the Fed-ECB policy differential, and monthly US unemployment and CPI - all on
ALFRED first prints, 499 permutations, against forward 5-day range.

NOT ONE macro feature clears the family-wise bar on any symbol. The only thing
that clears anywhere is the trailing-range positive control, which is what it
is there to do. Best a-priori candidate, the Fed-ECB differential on EURUSD,
came in at MI 0.00170 p=0.088 - nothing. The two features flagged INCREMENTAL
(dgs2_chg5 on SP500) have null marginal MI and are isolated conditional cells
at the expected false-positive rate, not findings.

The `distinct` column quantifies the power argument instead of asserting it:
unemployment takes 51-66 distinct values across 3,745-6,159 bars, CPI 174-277,
against 6,159 for a continuous feature. A monthly series pasted onto daily bars
carries about 1% of the resolution, and it showed - the monthly features were
among the weakest in every table.

The contrast with the implied-vol screen is the useful part: the options
market FORWARD-LOOKING view of an instrument (gvz_chg5 on gold, MI|vol 0.0197)
carries real information about its range, while the economy BACKWARD-LOOKING
state carries none. Mismatched timescales - rate levels move over months,
5-day range moves daily.

Also makes load_bars fall back to htf/{SYM}_D1_mid.npz when the tick-derived
build is absent (the 2026-08-16 disk cleanup removed bars/ but htf/ survived),
with need_ticks=True turning that fallback into a loud failure for the
order-flow screen rather than silently testing flow features on OHLC data.

No EA change: nothing survived to wire.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-16 17:21:50 -04:00
AnimateDread
d83cecc011 feat(altdata): wire instrument-specific implied vol; fix FRED vintage path
Wires the screen_ivol survivors (41d726c). New per-symbol `ivolSeries` in the
catalog feeds a generic `ivol_chg5` feature from whichever CBOE vol index the
instrument owns, so one code path serves every symbol:

  XAUUSD  + ivol_chg5 (GVZ)  - MI|vol 0.01971 p=0.002, 3.6x the positive
                               control and 4.6x the vix_chg5 gold had alone.
                               vix_chg5 KEPT: this appends, it does not replace.
  EURUSD  + vix_chg5         - screened, incremental p<=0.006, and its first
                               real feature ever (it had only exploratory EIA).
  USDJPY  + vix_chg5         - screened, incremental.
  NAS100 / US30 / US2000 + ivol_chg5 (VXN / VXD / RVX) - exploratory by analogy.
  XTIUSD / XBRUSD + ivol_chg5 (OVX) - exploratory, no oil bars to screen yet.
  SP500 unchanged - its features already screened clean and VXN/VIX3M edging
  out VIX is a correlated within-family best-of-N, not a real ranking.

On EURUSD/USDJPY the screen put VXD marginally above VIX, but they are
near-duplicates and the gap sits inside the noise, so the tie is broken by a
rule rather than by the number: take the series already in the fetch path.

Also fixes a real collector bug: fetch_vintaged built ALFRED realtime windows
out to 2028, and FRED rejects realtime_end after today - so every REVISED
series (unemployment, CPI, GDP: exactly the ones needing the vintage path) was
unreachable, while unrevised series never noticed because they bail earlier.
UNRATE and CPIAUCSL now return first prints correctly.

Adds screen_macro.py (rates, curve, breakevens, Fed/ECB policy differential,
plus monthly country stats) with a `distinct` column that reports the honest
effective sample size - a monthly series pasted onto D1 bars is a step
function, and that column is what decides whether it can clear a gate at all.
Not yet run: the Market Data bars directory is being regenerated right now.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-16 17:14:03 -04:00
AnimateDread
41d726c746 research(altdata): instrument-specific implied-vol screen - GVZ is a major find on gold
No free historical GEX exists: probed the CBOE chain endpoint with date/dt query
params (both silently ignored, returned today) and dated/historical paths (403),
and the CBOE index-history CSVs are 403 too. The forward recorder stays the only
path to GEX history.

But the options market publishes its per-instrument view of future range as the
CBOE vol indices, and FRED carries the whole family free with 15-25 years of
history - screenable today with the existing collector and harness. Fetched
GVZ (gold), OVX (oil), VXN, VXD, RVX, VIX3M.

HEADLINE - XAUUSD: gvz_chg5 (gold IV 5-day change) MI 0.02103, MI|vol 0.01971,
p=0.002. That is 3.6x the trailing-range positive control and 4.6x the vix_chg5
this project currently ships on gold - the second-largest incremental MI of the
whole campaign, on a symbol that carries exactly one screened feature today.

Vol-change is incremental on all four symbols: SP500 (known), USDJPY vxd_chg5
0.00492, and EURUSD vxd_chg5 0.00412 / vix_chg5 0.00379 - notable because
EURUSD has no screened features at all and its own trailing range is a weak
control there, so external vol carries information its own history does not.

Caveats recorded in the script and memory: SP500 within-family ordering
(VXN > VIX3M > VIX, all ~0.031-0.038 conditional) is a best-of-N artifact and
must not be cherry-picked; XAUUSD noise control misbehaved this run (MI|vol
0.00271 p=0.002), so anything under ~0.003 conditional on gold is unresolved -
gvz_chg5 at 7x that floor is unaffected; EVZ (euro IV) is DISCONTINUED since
2025-03 and must never be wired.

Nothing wired - the EA is mid-deploy and this would re-key every model again.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-16 17:06:28 -04:00
AnimateDread
1bee06c945 docs(altdata): FlashAlpha free tier probed - no validation possible, verdict hardens
Spent 3 of 5 daily requests. All three were informative:
  ETF data (SPY/QQQ/IWM) requires Basic - free tier is single stocks only.
  Full-chain GEX (all expirations) requires Growth - free and Basic must
    query one expiration per request, so even with history a full-chain
    backfill would be 24-54 requests per day of history.
  AAPL?expiration=2026-09-18 returned 200 with the right schema but a nearly
    empty payload: 13 of 93 strikes carried any open interest, total call OI
    4,296 against CBOE 373,253 for the same expiry, put OI zero, and every
    near-the-money strike blank.

So the construction could not be validated - not because the math disagreed
but because there was nothing to compare against. From outside it is not
possible to tell free-tier degradation from their flow-signed methodology,
and finding out costs $1,499/month.

Verdict hardens: the free CBOE CDN is strictly better than Basic for this
project - complete chains, every expiry and strike, gamma and open interest
populated, unlimited, $0. Our own AAPL figures were internally coherent
(+0.929 Bn/1% total, Sep-18 expiry +0.154 Bn, near-money gammas 0.013-0.019).
GEX stays externally unvalidated; if that ever matters, use a different vendor.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-16 16:43:23 -04:00
AnimateDread
9e49b2aeae docs(altdata): FlashAlpha backfill is dead - historical API is Alpha-tier only
Pricing checked: Free $0 (5/day), Basic $79 (250/day), Growth $299 (2,500/day),
Alpha $1,499 (unlimited) - and the Historical API is ALPHA-EXCLUSIVE. Basic and
Growth serve live data only.

The archive was the only thing worth buying from this vendor, so nothing in
budget helps: Basic would spend $79/month to make a once-a-day snapshot 15
seconds fresh instead of 15 minutes. Not subscribing.

The free key keeps one genuine use: a single live call to compare their GEX
against our CBOE-computed number, validating the recorder formula against a
commercial implementation (sign and magnitude only - they sign strikes from
classified tape, we use the standard open-interest assumption).

Recorded the EV argument for future sessions: the recorder banks this history
for free in ~12 months, and on this project base rate most alt-data families
die at the incremental gate. Paying four figures to test GEX a year early is a
poor trade. If revisited, price bulk ARCHIVE sellers, not analytics APIs.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-16 16:35:40 -04:00
AnimateDread
3abb2b1a7f feat(altdata): GEX forward recorder (CBOE delayed-quotes CDN, no key)
Option open interest is a snapshot source - no free history exists anywhere -
so the series only accrues from the day recording starts. That is why this
ships BEFORE the redeploy: every day the EA is not running is a day of history
that cannot be recovered later.

Records one row per weekday after 21:00 UTC to gex_{CANONICAL}.csv: net/call/put
dollar GEX per 1% move, call and put OI, the three nearest expiries and the
front expiry code. Feeds NOTHING - wiring a feature that is missing across ~100%
of the training sample would waste input width and hand batch-norm a constant.
It becomes a screening candidate at ~250 rows, gated like every other feature.

Thesis: dealer gamma is a RANGE mechanism (long gamma -> hedging sells rallies
and buys dips, range compresses; short gamma amplifies both ways), and range is
this project's one proven channel.

Verified in situ against the live SPX chain before writing any MQL5: 29,362
contracts, 20,993 with nonzero gamma, 54 expiries, total +90.7 Bn/1% (calls
+305.7, puts -215.0), and 100% of net GEX inside 5% of spot. The CDN publishes
per-contract gamma directly, so no pricing model - and no model risk - enters
the recorded data. Also verified the CDN does NOT gate on User-Agent (the old
"CBOE is UA-gated" note in DESIGN.md was a different CBOE path), so plain
WebRequest reaches it.

Dropped a zero-gamma "flip level" field: the probe returned a crossing above
spot while total GEX was strongly positive, which is incoherent - a static
gamma snapshot cannot give a flip level without repricing. Recording a
plausible-looking wrong number is worse than recording nothing.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-16 16:30:39 -04:00
AnimateDread
0a198fb95f feat(altdata): EIA wired, 24-instrument symbol catalog, mapping dialog for unknown symbols
EIA (user directive: "the NN might find patterns in it for both oil and regular
symbols"). Weekly Petroleum Status Report via the v2 API - crude stocks ex-SPR,
field production, refinery utilization - three features (1y percentile, 4w
change, utilization) on EVERY catalog symbol, not just oil. EIA screened NULL on
WTI's short 7y sample, so these ship as EXPLORATORY inputs: the deploy gate, not
the screen, decides whether a model trained on them trades. Publication stamp
observed+6d mirrors research/altdata/eia.py.

Symbol handling was hardcoded to three if-blocks; it is now a catalog of 24
instruments x alias lists covering The5ers/FTMO/AvaTrade/Dukascopy/OANDA/IC
Markets naming, with prefix matching for the broker suffix zoo (US500.cash,
XAUUSDm, EURUSD.r). Adding an instrument is one AddSpec row. COT caches are
named by CANONICAL so two brokers' names for one contract share a download.

Unrecognised symbol -> a chart dialog (Panel\AltDataMapDialog.mqh, CAppDialog +
dropdown) asks which instrument it is; the answer persists in symbol_map.cfg and
"No alternative data" is a recorded choice, not a nag. Non-blocking by design:
an unmapped symbol contributes 0 features and must never hold up a chart.

Also: UrlEncodePart now escapes '%' - SoQL like-predicates use it as the
wildcard and an unescaped one corrupts the query; docs/ gains the whitelist
URLs, an API-key backup, and the catalog reference.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-16 16:18:29 -04:00
AnimateDread
2f901994fa test(altdata): flow/activity screen - tick ACTIVITY clears family-wise on all 4 symbols for range
VPIN-style toxicity (|imbalance|): null-to-marginal everywhere. But tick
ACTIVITY (count vs 20d mean) clears the family bar on ALL FOUR symbols for
forward range AND survives conditioning on trailing realized range
(SP500 MI|vol 0.025, XAUUSD 0.0088, USDJPY 0.0069, EURUSD 0.0049, all
p<=0.006). On EURUSD it beats the trailing-range positive control itself -
resolving the void-control anomaly: EURUSD D1 range IS predictable, just
not by its own trailing range. USDJPY spread_stress (max/mean) also
family-clean + incremental. Direction: nothing beyond the known SP500
leverage effect. Validates the EA's volume feature block for the RANGE
objective the tuner now optimizes.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-16 15:32:31 -04:00
AnimateDread
51093697e2 test(altdata): EIA x WTI screen - NOTHING clears the family bar, EIA stays research-side
12 features (5 EIA petroleum + 3 COT managed-money + VIX/USD + controls)
vs forward 5-bar range and direction on 1,789 D1 bars resampled from the
decoded XTIUSD M1 file, 499 circular-shift perms. No feature clears the
family-wise bar on either target; every EIA fundamental is null even
marginally (best p=0.13). Sample is short (~7y) so a weak effect is not
excluded - but per the gate, no EIA feature ships. The EIA key stays in
keys.txt for future use (longer history / recorded surprise-vs-consensus).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-16 14:04:37 -04:00
AnimateDread
8ce635ce70 feat(altdata): external feature block wired into the NN feature window
- System\AltData.mqh: CAltDataPanel - publication-stamped CSV panel
  (Common\Files\Warrior_EA\AltData\{SYM}_{TF}.csv), as-of lookup by bar
  open, 0-fill degradation (mirrors cross-asset), hourly live refresh
- Topology: width block AFTER the .cfg name-list pin is pre-read
  (ReadAltDataPinFromCfg) so a grown export can never mismatch a resumed
  model's width or shift its slots
- Persistence: alt pin appended to the .cfg (append-and-length-guard
  convention), adopt-don't-compare on load
- Features: emit block after Wyckoff SBI; EnsureFresh probe in
  BuildFeatureWindow (never fires in tester)
- export.py: fixed a-priori scale constants (never data-fitted)

Widths change SP500 +4 / USDJPY +3 / XAUUSD +1 (fingerprint re-keys ->
fresh models on redeploy); EURUSD exports nothing and resumes unchanged.
Compiles 0 errors / 0 warnings.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-16 13:39:00 -04:00
AnimateDread
6d50e63370 feat(altdata): conditional (incremental) MI + direction-channel screen
MI|vol column = I(X; target | trailing-range tercile), same circular-shift
null. Range target 499 perms: SP500 vix_chg5 survives conditioning at 0.031
(3x trailing range's own within-tercile residual); VIX LEVEL emerges
conditionally (variance-risk-premium structure); USDJPY COT family survives.
Direction target: SP500 vol/VIX-chg clear marginally (equity leverage
effect) but drop to p~0.05-0.06 conditional = redundant with price vol;
USDJPY/XAUUSD/EURUSD direction null across all alt features.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-16 12:57:27 -04:00
AnimateDread
4a56d8da48 feat(altdata): FRED/EIA collectors live + first MI screen - COT clears on USDJPY, VIX-change on SP500
- fred.py: ALFRED output_type=4 first prints, chunked realtime windows
  (2000-vintage cap), unrevised-series fallback (published=observed+1d);
  NFCI excluded from features (revised, no vintage archive)
- eia.py: 4 weekly petroleum series on disk (1982->now)
- screen.py: as-of joined alt features vs forward 5-bar range/ATR on D1,
  3x3 MI, circular-shift null, family-wise max bar, +/- controls

First readings (199 perms): SP500 vix_chg5 MI 0.047 (1.5x the positive
control) + usd_chg5 clear family bar; USDJPY 4 COT positioning features
clear family bar BEATING the positive control; XAUUSD vix_chg5 tops control
but sub-family-bar; EURUSD positive control FAILS -> table void per the
excursion-target rule, needs investigation.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-16 12:45:21 -04:00
AnimateDread
6d1ecb71ce feat(altdata): alternative-data collector package - COT shipped, FRED/EIA ready
Private-use pivot (marketplace dropped): DLL/Python/WebRequest now allowed.
- altdata/cot.py: CFTC COT, no key, 2010->now on disk for all 9 symbols
  (TFF: ES/VIX/BTC/EUR/JPY/CAD/GBP; Disagg: GC/CL); publication-lag stamping
  (Tuesday report -> Saturday 00:00 UTC availability)
- altdata/fred.py: ALFRED first-print vintages (needs free key)
- altdata/eia.py: weekly petroleum status (needs free key)
- DESIGN.md: source adjudication (corrections to the LLM source list),
  vintage + family-wise rules, EA file contract, staged EA-side plan

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-16 12:29:47 -04:00
AnimateDread
f55dbe995c feat(research): drift null for the first-ever family-wise gate pass (SP500 D1 fractal PAI)
2026-08-15, ~6 minutes after attach: PAI-8fea (fractal target, 37
features, D1) converged at era 299 and the plateau deploy CLEARED the
family-wise gate for the first time in project history: dir-precision
73.1% vs 63% break-even, +10.4pp on 350 test calls = 4.04 sigma,
p_family = 0.0081.

This script asks the first two hostile questions offline:
- DRIFT: always-long at the same 2.64/1.66 geometry scores 64.1% on the
  last 15% of D1 history (66.7% on 30%) - drift alone clears BE by
  ~1-3pp, but the model is +9pp above ALWAYS-LONG, so the pass is
  selection, not drift.
- SWAP EXPOSURE: median 7-8 bars to the long target = ~10 nights of
  financing ~ 0.15-0.2% notional vs a ~1.7% target -> a ~1-1.5pp BE
  haircut against a +10.4pp margin. Survives.

Remaining before belief: replication on other D1 symbols, and closing
the live-semantics gap (certified wins assume hold-to-barrier; live
exit paths can cut on vote flips).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-15 16:09:08 -04:00
AnimateDread
4e00955463 feat(research): pooled fractal-pivot direction NN - the last clean look at direction, gate FAILED 4/4
User request: NNs predicting pivots (fractals for label density). Run on
the offline stack that demonstrably CAN learn (+2.6pp XAUUSD meta), free
of every historical in-EA training bug: pooled 4-symbol training,
per-symbol norm, scale-free causal features, target = side of current
price the next confirmed Bill Williams fractal lands on, real M1
ask/bid fills, threshold fitted on calib only, pre-registered 2-sigma
net-expectancy gate.

Result: train CE 0.682 (a whisper below the 0.693 coin), and on test no
symbol passes - EURUSD/USDJPY net zero, XAUUSD gross +0.22 pts vs a
larger spread (net -0.30), SP500 net +0.05 +/- 0.45. The gross-positive
tails are index drift plus sub-spread micro-reversion - the tick-flow
decay shape at swing scale. Seventh independent measurement of the same
fact: entry-time direction information is not in these features.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-14 18:33:11 -04:00
AnimateDread
c85f6a3839 feat(research): stock-ZigZag live replay - the repaint measurement that answers "trade the true swings"
Faithful Python port of ADZigZag (stock MetaQuotes ZigZag 12/5/3,
verbatim rebrand) including the incremental prev_calculated branch, so
the indicator can be replayed bar by bar exactly as it draws live.

SP500 H1, 74,599 bars, fills at real M1 ask/bid:
- FINAL swings: 4,658 legs, mean 50.8 pts = 95 spreads. Perfect
  foresight +50.2 pts/leg. The user premise (swings dwarf spread) is
  fully confirmed.
- LIVE: 81% of drawn newest-pivots later repaint away entirely
  (18,805 of 23,299). Holding the drawn direction at every bar close
  grosses +0.38 pts/trade (t=0.8, zero cost charged) out of the
  50.8-pt average swing - 0.7% of the line the chart ends up showing.
- LONG +1.66 gross / SHORT -0.90 = the index drift, nothing else;
  long net +1.17 pts / 13-bar hold = ~1.5 bp, under one night financing.

The spread subtracts 0.49 pts of a swing that hands over 0.38: cost was
never the obstacle - pivot knowledge is.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-14 18:12:11 -04:00
AnimateDread
9c6c81a706 feat(research): pivot-to-pivot swing expectancy - the spread-vs-swings question measured
User claim: swings dwarf the spread, so cost cannot be what blocks swing
trading at 1:2/1:3 RR. Measured on the validated M1 bid/ask book, SP500
H1, ATR-scaled causal ZigZag at 4 reversal thresholds:

- The premise is CONFIRMED: median swing 36-79 spreads, mean up to 110.
  Perfect-foresight expectancy +28 to +59 pts/leg.
- The conclusion does not follow: trading every confirmed leg (enter on
  the ZigZag confirmation close, real ask/bid fills, exit on the next
  confirmation) grosses -0.2 to -0.5 pts/leg AT ZERO COST, on 3,681 to
  13,871 legs. The confirmation retracement - the event that DEFINES a
  pivot - consumes the entire swing before the spread is even charged.
- Long/short split is symmetric around the index drift (LONG +1.18,
  SHORT -2.17 gross at 3xATR), i.e. no swing structure beyond drift.
- 3.0xATR reversal reproduces the EA ZigZag cadence exactly (median leg
  17 bars, 49 legs/1000 bars vs the EA measured 17 and 44).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-14 16:34:33 -04:00
AnimateDread
f3bded779a feat(research): pooled meta-labeling verdict tooling - memmap loader, per-symbol norm, threshold dose-response
Three additions to meta_pool.py, in the order the campaign needed them:
- memmap + float32-throughout (per-batch float64 cast): the 6.5 GB
  4-symbol corpus OOMed the float64 pipeline on the training box;
- pool2: per-symbol standardization (each symbol by its own train-slice
  mu/sd) + 64/32 capacity + l2 1e-3, after the naive pooled model
  underfit to the prior (train CE pinned at base-rate entropy);
- curve: fixed-ladder precision-vs-threshold on calib and test side by
  side - the dose-response diagnostic that closed the question.

RESULT recorded in memory: pooling transfers real skill (XAUUSD +2.6pp,
SP500 +1.2pp at fitted thresholds, >>2 sigma) but 0/8 fitted operating
points clear break-even, and the high-conviction tail is temporally
unstable - the precision-vs-threshold slope FLIPS SIGN between calib and
test on 3 of 4 symbols, so no ex-ante threshold rule exists.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-14 08:41:24 -04:00
AnimateDread
f878ea4c75 feat(research): optimize memory usage with np.memmap and improve data handling in MLP 2026-08-14 07:35:29 -04:00
AnimateDread
932c94e890 feat(research): offline meta-pool pipeline - loader, EA-mirrored splits, MLP, cov x (p-BE) eval
Loads the EA's MetaExport .f32 datasets (UTF-16 sidecars), applies the EA's
own discipline offline: chronological 55/15/30 split with horizon-length
purges, operating point fitted on the calibration slice only via
coverage x (precision - BE) with the 25% floor, test slice touched once,
deployability at the 2-sigma edge floor. Small leaky-ReLU MLP + Adam in
numpy; `stats` / `eval <tag>` / `pool` commands.

First run on XAUUSD_16388 validated the plumbing and exposed the data:
the 2.5h gold tester run only covered 2004-07..2006-10 (3,214 candidates)
because gold tick volume is huge - and corpus builds do not need ticks at
all (journaling is bar-open-keyed, labels come from bar history later), so
"Open prices only" modeling builds the same corpus ~100x faster.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-13 16:09:14 -04:00
AnimateDread
b06fdd2f0e research: point the seasonal screen at the five breadth instruments
seasonal.py gets a frame-based entry (analyse_frame) and a reusable
report() so the identical statistics - circular-rotation family-wise
null, max-|t| bar, split-half - can run on instruments whose book is
synthesised from M1 bars. breadth_seasonal.py runs it on the five
SQX-decoded instruments (FTSE100, UK100, WTI x2 feeds, USDCAD) that
share no data path with the four originals; the duplicate-market pairs
(FTSE100/UK100, WTI_d/WTI_5) double as replication checks. Caveats
stated in the module docstring: synthesised flat spread (move/spread
is approximate, no intraday spread shape) and file-time clock labels;
drift/t columns are spread-free and unaffected.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-12 14:08:05 -04:00
AnimateDread
371f8aaecd fix: the Adam second moment was never Adam - all four tiers
Root cause of the B=32 regression, and it predates F4 entirely. Every Adam
kernel stored v already square-rooted and then fed that stored value back in
as if it were the variance:

    v_new = sqrt(b2 * v_old + (1 - b2) * g^2)

That recursion has a fixed point at v ~= b2 = 0.999 for ANY gradient below
unit scale, so the denominator stops tracking the gradient and Adam degrades
into plain SGD with lr = lt. Measured against the shipped WarriorCPU.dll
(batch_accum_check.cpp, TestOptimizerScaleInvariance), 4000 steps of a
constant gradient: 3285x less displacement at |g|=1e-5 than at |g|=1, where
a scale-invariant optimizer gives the same distance for both. After the fix
all six magnitudes read 1.199 and v tracks |g| exactly.

It hit conv/LSTM specifically because they sit behind a batch-norm with
running variance ~2.6e+05, so their gradients arrive divided by ~500 - deep
in the degraded regime - while the dense stack near the loss stayed in the
working one. In situ on SP500 H1: lstm1 dW/W 2.62/10.0/7.14% -> 0.024/0.022/
0.003%, conv1 decaying to 0.000% by era 30. NeuronBatchNorm.mqh already
squared v back for gamma/beta and its comment named the kernels as wrong,
which is exactly why gamma/beta kept training while the stages behind froze.

Persisted .nnw needs no migration - v keeps its std-dev meaning.

Also, the two ways F4 exposed it, both mine:

- No LR compensation for B fewer steps per era. sqrt(B) for adaptive methods
  (Krizhevsky 2014; Granziol et al. 2022), applied once in
  InitialEtaForOptimizer(). Linear scaling (Goyal et al. 2017) is for SGD.
- Plateau patience denominated in eras, so raising B made the ladder 32x more
  impatient in its only unit. PAI converged at era 41 on ~49k updates where
  the same config had been finding new bests at era 1028.
  TrainPlateauPatienceEras() stretches it by the same sqrt(B).

TRAIN_BATCH_SIZE 32 -> 8 so the patience stretch stays affordable (8 -> 23
eras per stage, not 8 -> 45). Both helpers are identities at B=1.

Deploy gate: DEPLOY_MIN_SIDE_RECALL_PCT (10%) folded into tradeableOK. The
perceptron reported Sell:0% recall in all 41 eras, cleared the floor on Buy
alone at 36.6% vs 34% chance, deployed, and sprayed buy arrows. Folded into
the ranking key rather than checked at deploy time so a one-sided era cannot
become best-so-far in the first place.

Deinit: the arrow purge now runs BEFORE ExtPanel.Destroy(), an unbounded
CAppDialog teardown that sat ahead of it - the same ordering inversion the
rule there exists to prevent. CONV was force-terminated 4.8 s into OnDeinit
(vs ~1.1 s for the three that finished) having reached none of its cleanup,
so its arrows stayed on the chart. Steps are now timed in the log.

PurgeChart's verification rescan filtered on OBJ_ARROW, the same blind spot
as the bulk delete, so "persisted 10 ... cleared 0" passed silently. It now
walks every object type and reports the object counts when both are zero.

Both build variants compile 0 errors / 0 warnings; both DLLs rebuilt.
FORCES A RETRAIN (already forced by N1) and both DLLs must ship with the .ex5.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 14:02:35 -04:00
AnimateDread
0c01dc279b feat: mini-batch gradient accumulation (F4), front-end-aware capacity budget (F6), split Wyckoff categoricals (N1)
Completes the 2026-08-09 training audit. FORCES A RETRAIN of every
Wyckoff-enabled config (N1 re-keys the fingerprint), and BOTH DLLs must be
redeployed alongside the .ex5 - they carry new exports.

F4 - mini-batch accumulation, TRAIN_BATCH_SIZE=32. Training was pure online
SGD (one weight update per bar), which is the mechanical source of the
era-to-era whipsaw every downstream guard was built to cope with. The O(n^2)
outer product is native - AccumulateWeightGrad / AccumulateWeightGradConv /
AccumulateBufferInto in Network.cl, WarriorCPU and WarriorDML - while the
optimizer step is host-side MQL5 shared by all tiers (ApplyAccumToBlock), so
there is one Adam/SGD implementation instead of four that can drift.
  - the LSTM needs no outer-product kernel (WeightsGradient already holds the
    sample's full dW) but could NOT simply be left un-zeroed between samples:
    CPU_LSTMSeqBackward/DML_LSTMSeqBackward memset it on entry. Hence a
    separate accumulator plus an elementwise add.
  - batch-norm gamma/beta accumulate in host arrays, not new BatchOptions
    slots - BN_OPT_STRIDE is baked into every persisted .nnw.
  - scoped to pass 2; online learning keeps immediate updates. Every save /
    checkpoint / scoring boundary flushes, scaling by the real sample count.
  - degrades to per-sample updates (one log line) on a tier that cannot
    accumulate, so old devices and DLL-free builds are unaffected.
  - verified offline: DirectML/batch_accum_check.cpp drives the real exports
    against an independent reference; at B=1 the accumulator matches the
    shipped unbatched kernel's own gradient to 1.1e-16. Math only - the
    in-situ check remains the per-layer dW/W report on a real era.

F6 - ComputeFirstLayerWidth budgeted against the RAW input width even where a
conv/LSTM front end had already reduced it, so an LSTM's dense stack was
charged for 1,280 inputs when it receives 64. Confirmed from the deployed
.cfg files: CONV, LSTM and HYBRID were all pinned at the 16-unit floor. Now
budgeted against the front-end output and capped at it (never fan out), with
the derivation reordered so both stages settle first.

N1 - EventCode/EventPhase/StructuralPhase are signed categoricals packing
direction and Wyckoff stage into one scalar across a sign discontinuity. Split
into direction + [0,1] magnitude, the same convention the base OHLC block uses.
Information-preserving; 13 readings now occupy 16 inputs.

Compiled clean (0 errors, 0 warnings); both DLLs rebuilt.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 11:48:03 -04:00
AnimateDread
274630f802 fix: training-stability audit fixes F1/F2/F3/F5 - unbiased shuffle, real plateau escapes, fresh optimizer state on restore, pure OOS metric
Four of the six findings from research/training_pipeline_audit_2026-08-09.md
(F4 mini-batching and F6 feature re-encode deliberately deferred - see the
report's implementation-status section for why):

- F1: pass-2 Fisher-Yates (and AutoTune's MI block shuffle) used MathRand()%,
  which is 15-bit - provably non-uniform on every full-history era over 32,768
  queued samples. New 30-bit ShuffleRandomIndex().
- F2: plateau warm restarts were a no-op whenever eta already sat at its
  ceiling (the normal state of a non-regressing plateau) - the ladder was just
  a 24-era countdown. Restarts now overshoot to 5x the ceiling
  (PLATEAU_RESTART_BOOST) and anneal geometrically back over the patience
  window, SGDR-style; ETA_MIN widened 1e-4 -> 1e-5 so the decay schedule has
  real range.
- F3: checkpoint restores put weights back but kept the rejected trajectory's
  Adam moments, so the optimizer immediately pushed back toward the rolled-back
  state (the restore->regress->restore oscillation). CNet::ResetOptimizerState()
  zeroes moments/momentum/step counters (weights, BN statistics, gamma/beta
  untouched) on every mid-run restore, every boosted restart, and the
  deploy-time restore that online learning continues from.
- F5: batch-norm running statistics now freeze for the pass-3 OOS scoring walk,
  so the selection metric the checkpoint ranking and deploy gate read is a pure
  function of the checkpoint instead of partly measuring BN drift. Defensive
  unfreeze in FinalizeTrainRun covers stop-mid-pass; live/online adaptation and
  the OOS continual-learning simulation stay adaptive by design.

Compiled clean (0 errors, 0 warnings) via the staged-tree recipe.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 10:54:09 -04:00