forked from MrBaro75/Warrior_EA
Resurrecting the diagnostic in 4113afd immediately showed why it needed more
than resurrecting. Its first output, on every chart:
strongest of 49 columns x 7999 lags is column 4 at lag +3900, |r| 0.5120
(32.4 SE of the 0.0158 no-information band) -> SURVIVES
XTIUSD and SP500 named the SAME column at the SAME lag with |r| within 0.01 of
each other. Two independent instruments cannot agree to that precision at a
3900-bar lag; that is what identified it as an artifact rather than a finding.
TWO DEFECTS, both of which manufacture significance at long lags.
1. THE NO-INFORMATION BAND WAS GLOBAL, THE CORRELATION IS NOT.
CorrR1D is non-circular: lag k is computed from (n - |k|) overlapping terms
while the normaliser uses all n. Scoring every lag against one 1/sqrt(n)
band understates it by sqrt(n/(n-|k|)) - a factor of 6 at the edge. Lag
+3900 of 3999 rests on ~100 overlapping terms and was being judged as if it
rested on 4000. Each lag now gets its own band, deflated for label
persistence, and the maximum is ranked on z rather than |r| - two equal
correlations are not equally surprising on different sample sizes.
2. THE LAG RANGE ANSWERED NO QUESTION.
The profile exists to say how deep a lookback carries linear structure, and
the input window is HistoryBars. Lags two orders of magnitude past it are
both meaningless and where this estimator is worst. Capped at 4x the window.
The Sidak family is now the lags actually searched, not 2n-1.
Left deliberately: this remains a LINEAR, marginal measure. It can bound the
useful lookback; it cannot prove a shorter window loses nothing.
Build tag -> lagprofile-v2. Not a layout change: no retrain.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
961 lines
48 KiB
MQL5
961 lines
48 KiB
MQL5
//+------------------------------------------------------------------+
|
|
//| Warrior_EA |
|
|
//| AnimateDread |
|
|
//| |
|
|
//| CBaselineComparator - two non-NN learners on the NN's matrix. |
|
|
//| |
|
|
//+------------------------------------------------------------------+
|
|
#ifndef WARRIOR_TRAINING_BASELINECOMPARATOR_MQH
|
|
#define WARRIOR_TRAINING_BASELINECOMPARATOR_MQH
|
|
#include "ITrainingData.mqh"
|
|
|
|
//--- ONCE-PER-CHART, same doctrine as g_ensembleChartMiReportDone (AutoTune.mqh) and for a stronger
|
|
//--- reason. Four un-chunked forest+MLP+OLS fits over a ~1000-column, 4000-row design, to print the
|
|
//--- same numbers four times.
|
|
bool g_ensembleChartBaselinesDone = false;
|
|
|
|
//--- THE BOUND IS THE EA THREAD, NOT THE BAR. This pass used to be described as bounded because it
|
|
//--- "cannot outrun the bar it runs on" - four hours on H4. That is the wrong target and was never
|
|
//--- enforced: the EA is single threaded, so for as long as this runs there is no training, no
|
|
//--- panel, no tick handling, and a removal request just queues.
|
|
#define BASELINE_MAX_TRAIN_ROWS 4000
|
|
#define BASELINE_MAX_SCORE_ROWS 4000
|
|
//--- ROWS x COLUMNS, because every fit here costs O(rows x width) and the row caps above carry no
|
|
//--- width term. 800k = the 4000-row cap at the ~200-column window these numbers were chosen for;
|
|
//--- at this build's 800 columns it buys 1000 rows and an MLP fit of ~13 s instead of 52 s.
|
|
#define BASELINE_MAX_CELLS 800000
|
|
//--- Wall clock for the whole pass, checked between phases. Deliberately far below the ~4,500 ms
|
|
//--- OnDeinit budget's timescale being reached at all: the point is that a stop never has to wait.
|
|
#define BASELINE_BUDGET_MS 45000
|
|
//--- CALIBRATION POINT for the linear phase, which is the one that cannot be bounded by a boundary
|
|
//--- check. LRBuild solves a (width+1)^2 normal-equation system, so its cost grows as width^3 and
|
|
//--- does not care how few rows there are.
|
|
#define BASELINE_LR_MS_AT_800 535000.0
|
|
#define BASELINE_TREES 32
|
|
#define BASELINE_SUBSAMPLE_RATIO 0.66
|
|
//--- Three models are tried, so the family-wise correction is over three - the same rule the deploy
|
|
//--- gate applies to its own best-of-N selection. Reporting each model's uncorrected p as if it were
|
|
//--- the only one tried is precisely the error catalogued four times already.
|
|
#define BASELINE_MODELS_TRIED 3
|
|
//--- MLP baseline. Small and bounded on purpose: this is a CONTROL, not a competitor. One restart,
|
|
//--- a hard iteration cap, and the decay ALGLIB's own docs recommend when you have no reason to
|
|
//--- pick another.
|
|
#define BASELINE_MLP_HIDDEN 8
|
|
#define BASELINE_MLP_DECAY 0.001
|
|
#define BASELINE_MLP_RESTARTS 1
|
|
#define BASELINE_MLP_WSTEP 0.01
|
|
#define BASELINE_MLP_MAXITS 100
|
|
//--- Folds for the MLP's cross-validated error bar. Three, not the usual five or ten: each fold is
|
|
//--- a full retrain, so this multiplies the MLP's cost by exactly this number on a six-core 2013
|
|
//--- box.
|
|
#define BASELINE_MLP_CV_FOLDS 3
|
|
//--- Below this the fit is not worth reporting: an OOS slice this thin cannot resolve an edge from
|
|
//--- noise at any precision, and a train slice this thin cannot fit anything but its own rows.
|
|
#define BASELINE_MIN_ROWS 200
|
|
//--- Redundancy report: the share of total variance the leading components must reach before the
|
|
//--- count is called the matrix's effective width, and the |r| above which a pair is called a copy.
|
|
#define BASELINE_PCA_VAR_SHARE 0.95
|
|
#define BASELINE_COLLINEAR_ABS_R 0.90
|
|
//--- A regime test needs at least this many DECLUSTERED samples per half to say anything at all.
|
|
#define BASELINE_MIN_REGIME_N 20
|
|
|
|
//+------------------------------------------------------------------+
|
|
//| PUT TWO COMPLETELY DIFFERENT LEARNERS ON THE NET'S OWN MATRIX. |
|
|
//| |
|
|
//| One question, and the whole module exists to answer it: when the |
|
|
//| net scores at chance, is that the MATRIX or is it us? An Alglib |
|
|
//| random forest and an Alglib MLP fit the same rows, the same |
|
|
//| labels and the same split. All at chance means the data is the |
|
|
//| limit. An off-the-shelf MLP well ABOVE our net on those rows |
|
|
//| would mean the limit is our implementation. |
|
|
//| |
|
|
//| It is a CONTROL, not a competitor - which is why the models are |
|
|
//| small, capped and un-tuned on purpose. |
|
|
//| |
|
|
//| It knows nothing about CExpertSignalAIBase. It reads a |
|
|
//| CTrainingDataView and prints; that is the entire contract, and it |
|
|
//| is what lets this be read, changed or dropped on its own. |
|
|
//+------------------------------------------------------------------+
|
|
class CBaselineComparator
|
|
{
|
|
private:
|
|
//--- BORROWED. The signal owns both the view and this object.
|
|
CTrainingDataView *m_data;
|
|
//--- ONCE PER MODEL. Set before the first thing that can fail, so a run that bails does not
|
|
//--- retry every era for the rest of the session.
|
|
bool m_done;
|
|
|
|
int BaselineCandidateBars(const int lo, const int hi, const int bars,
|
|
const int cap, int &rows[]);
|
|
bool BaselineBudgetSpent(const uint startTick, const string nextPhase);
|
|
void ReportCombiningWeights(void);
|
|
double CombinerSSE(const double &w[], const int members, const double &s[],
|
|
const double &t[], const int rows);
|
|
void ReportLinearLagProfile(const int bars, const int oosCutoff);
|
|
void ReportFeatureRedundancy(CMatrixDouble &rows, const int n);
|
|
void ReportRegimeStability(double &margins[], int &marginBars[]);
|
|
void ReportBaselineModel(const string label, const int calls, const int hits,
|
|
const int scored, const double chancePct);
|
|
|
|
public:
|
|
CBaselineComparator(void) : m_data(NULL), m_done(false) { }
|
|
~CBaselineComparator(void) { m_data = NULL; }
|
|
//--- Hand it the view once, when the owner is constructed. Everything else it needs it asks
|
|
//--- for; there is deliberately no setter for anything else.
|
|
void Bind(CTrainingDataView *data) { m_data = data; }
|
|
//--- Already reported? The owner asks so it can say so rather than silently skipping.
|
|
bool Done(void) const { return m_done; }
|
|
void RunBaselineComparison(const int bars, const int totalIter, const int oosCutoff);
|
|
};
|
|
//+------------------------------------------------------------------+
|
|
//| Bar indices in [lo, hi) that carry a label and could carry a full |
|
|
//| feature window, thinned by a uniform stride to at most `cap`. |
|
|
//| |
|
|
//| The window check is a bounds test only - whether BufferTempData |
|
|
//| will actually produce the row is not known until it is asked, and |
|
|
//| asking is the expensive part. So this over-selects slightly and |
|
|
//| the callers report the count they really got. |
|
|
//+------------------------------------------------------------------+
|
|
int CBaselineComparator::BaselineCandidateBars(const int lo, const int hi, const int bars,
|
|
const int cap, int &rows[])
|
|
{
|
|
int all[];
|
|
ArrayResize(all, 0);
|
|
for(int r = MathMax(lo, 0); r < hi; r++)
|
|
{
|
|
//--- The window reaches BACK from r, so once it runs off the deep end every larger r does too.
|
|
if(r + m_data.HistoryBars() - 1 >= bars)
|
|
break;
|
|
//--- HasLabel is the whole resolvedness test: the finality-gated cache never holds an
|
|
//--- unresolved bar.
|
|
if(!m_data.HasLabel(r))
|
|
continue;
|
|
int n = ArraySize(all);
|
|
ArrayResize(all, n + 1);
|
|
all[n] = r;
|
|
}
|
|
int total = ArraySize(all);
|
|
if(total <= cap || cap <= 0)
|
|
{
|
|
ArrayResize(rows, total);
|
|
for(int k = 0; k < total; k++)
|
|
rows[k] = all[k];
|
|
return total;
|
|
}
|
|
//--- UNIFORM stride, not the newest `cap` rows: taking a contiguous block would hand the models a
|
|
//--- narrower stretch of market than the net trained on, and any difference in score would then be
|
|
//--- partly a difference in regime.
|
|
double stride = (double)total / cap;
|
|
ArrayResize(rows, cap);
|
|
for(int k = 0; k < cap; k++)
|
|
rows[k] = all[(int)MathMin(total - 1, (int)(k * stride))];
|
|
return cap;
|
|
}
|
|
|
|
//+------------------------------------------------------------------+
|
|
//| Stop requested, or the pass has spent its wall clock. Replaces |
|
|
//| the bare ShutdownRequested() checks at every phase boundary: a |
|
|
//| stop and a blown budget need the same answer, and only one of |
|
|
//| them used to be asked. |
|
|
//+------------------------------------------------------------------+
|
|
bool CBaselineComparator::BaselineBudgetSpent(const uint startTick, const string nextPhase)
|
|
{
|
|
if(m_data.Stopping())
|
|
return true;
|
|
uint spent = GetTickCount() - startTick;
|
|
if(spent < BASELINE_BUDGET_MS)
|
|
return false;
|
|
Print(m_data.Id() + StringFormat(": Alglib baselines stopping before %s - %.1f s spent of a %.0f s budget."
|
|
" Everything printed above is complete; %s and the phases after it did not"
|
|
" run. The EA is single threaded, so this pass freezes training, the panel"
|
|
" and tick handling for exactly as long as it takes.",
|
|
nextPhase, spent / 1000.0, BASELINE_BUDGET_MS / 1000.0, nextPhase));
|
|
return true;
|
|
}
|
|
|
|
//+------------------------------------------------------------------+
|
|
//| TRAIN AND SCORE THE TWO NON-NN BASELINES. Once per run, at a |
|
|
//| pass-3 completion, and only when the trader asked for it. |
|
|
//+------------------------------------------------------------------+
|
|
void CBaselineComparator::RunBaselineComparison(const int bars, const int totalIter,
|
|
const int oosCutoff)
|
|
{
|
|
//--- THE pointer test for this whole module. Every private helper below is reached only from
|
|
//--- here, so guarding the one public entry guards all of them - and an unbound comparator is a
|
|
//--- wiring bug that must be loud, not a pass that quietly reports nothing.
|
|
if(CheckPointer(m_data) == POINTER_INVALID)
|
|
{
|
|
Print(__FUNCTION__ + ": ERROR - baseline comparator has no data view bound. No baselines this"
|
|
" run; the models are unaffected but the MATRIX-vs-IMPLEMENTATION question goes"
|
|
" unanswered.");
|
|
return;
|
|
}
|
|
if(!Run_Alglib_Baselines || m_done)
|
|
return;
|
|
m_done = true; // set FIRST: a run that bails below must not retry every era
|
|
if(m_data.IsEnsembleMember() && g_ensembleChartBaselinesDone)
|
|
{
|
|
Print(m_data.Id() + ": Alglib baselines already measured by another ensemble member on this chart - same"
|
|
" windows, same labels, same three fits, same answer. Skipped; the first member's report"
|
|
" above is this model's too.");
|
|
return;
|
|
}
|
|
g_ensembleChartBaselinesDone = true;
|
|
int width = m_data.HistoryBars() * m_data.FeaturesPerBar();
|
|
if(width <= 0 || oosCutoff <= 0 || totalIter <= oosCutoff)
|
|
{
|
|
Print(m_data.Id() + ": Alglib baselines skipped - no usable era geometry yet.");
|
|
return;
|
|
}
|
|
uint t0 = GetTickCount();
|
|
if(BaselineBudgetSpent(t0, "the matrix build"))
|
|
return;
|
|
//--- SAME SPLIT AS THE NET, read off the same helpers rather than recomputed here. IS starts one
|
|
//--- purge past the calibration band's far edge, exactly where the backprop queue starts.
|
|
int isLo = m_data.CalibrationHiIndex(totalIter, oosCutoff) + m_data.PurgeBars();
|
|
int trainBars[], scoreBars[];
|
|
//--- Rows shrink as the window widens - see BASELINE_MAX_CELLS. Floored at BASELINE_MIN_ROWS so a
|
|
//--- very wide window skips below on "too few rows to say anything" rather than silently fitting one.
|
|
int rowCap = (int)MathMax(BASELINE_MAX_CELLS / width, BASELINE_MIN_ROWS);
|
|
int trainCap = MathMin(BASELINE_MAX_TRAIN_ROWS, rowCap);
|
|
int scoreCap = MathMin(BASELINE_MAX_SCORE_ROWS, rowCap);
|
|
if(trainCap < BASELINE_MAX_TRAIN_ROWS)
|
|
Print(m_data.Id() + StringFormat(": Alglib baselines - %d-column window caps this run at %d rows (was %d)."
|
|
" Every fit costs rows x columns, and the un-capped design is what got the"
|
|
" EA force-terminated on 2026-08-21.",
|
|
width, trainCap, BASELINE_MAX_TRAIN_ROWS));
|
|
int nTrainWanted = BaselineCandidateBars(isLo, totalIter, bars, trainCap, trainBars);
|
|
int nScoreWanted = BaselineCandidateBars(0, oosCutoff, bars, scoreCap, scoreBars);
|
|
if(nTrainWanted < BASELINE_MIN_ROWS || nScoreWanted < BASELINE_MIN_ROWS)
|
|
{
|
|
Print(m_data.Id() + StringFormat(": Alglib baselines skipped - %d train / %d score bars carry a label,"
|
|
" below the %d needed to say anything.",
|
|
nTrainWanted, nScoreWanted, BASELINE_MIN_ROWS));
|
|
return;
|
|
}
|
|
Print(m_data.Id() + StringFormat(": Alglib baselines starting - %d inputs (%d bars x %d features),"
|
|
" %d train rows, %d OOS rows. Same windows, labels and split as the net.",
|
|
width, m_data.HistoryBars(), m_data.FeaturesPerBar(), nTrainWanted, nScoreWanted));
|
|
//--- BUILD THE TRAIN MATRIX. Last column is the 3-class target the forest wants (0=Sell, 1=Neutral,
|
|
//--- 2=Buy); the linear fit reuses the same allocation with that column shifted to -1/0/+1 below,
|
|
//--- so the two models cannot end up looking at different rows.
|
|
CMatrixDouble xy(nTrainWanted, width + 1);
|
|
//--- ...and, at no extra feature cost, the ANCHOR BAR'S OWN ROW alone. BuildFeatureWindow
|
|
//--- appends deepest-lookback first and lands on the anchor last, so those are the final
|
|
//--- FeaturesPerBar() values of every window.
|
|
CMatrixDouble anchors(nTrainWanted, m_data.FeaturesPerBar());
|
|
double x[];
|
|
ArrayResize(x, width);
|
|
int nTrain = 0;
|
|
for(int k = 0; k < nTrainWanted; k++)
|
|
{
|
|
int r = trainBars[k];
|
|
if(!m_data.RowFeatures(r, width, x))
|
|
continue;
|
|
for(int f = 0; f < width; f++)
|
|
xy.Set(nTrain, f, x[f]);
|
|
for(int f = 0; f < m_data.FeaturesPerBar(); f++)
|
|
anchors.Set(nTrain, f, x[width - m_data.FeaturesPerBar() + f]);
|
|
double cls = m_data.IsBuyLabel(r) ? 2.0 : (m_data.IsSellLabel(r) ? 0.0 : 1.0);
|
|
xy.Set(nTrain, width, cls);
|
|
nTrain++;
|
|
}
|
|
if(nTrain < BASELINE_MIN_ROWS || BaselineBudgetSpent(t0, "the forest"))
|
|
{
|
|
Print(m_data.Id() + StringFormat(": Alglib baselines abandoned - only %d of %d train windows built.",
|
|
nTrain, nTrainWanted));
|
|
return;
|
|
}
|
|
//--- FOREST. Builder object, not the deprecated DFBuildRandomDecisionForest one-shot. Variables per
|
|
//--- split are left on the library's own auto rule: with a window this wide any hand-picked number
|
|
//--- would be a tuned knob, and a baseline that needed tuning to lose would prove nothing.
|
|
CDecisionForestBuilder builder;
|
|
CDecisionForest forest;
|
|
CDFReport frep;
|
|
CDForest::DFBuilderCreate(builder);
|
|
CDForest::DFBuilderSetDataset(builder, xy, nTrain, width, 3);
|
|
CDForest::DFBuilderSetSubsampleRatio(builder, BASELINE_SUBSAMPLE_RATIO);
|
|
CDForest::DFBuilderSetRndVarsAuto(builder);
|
|
CDForest::DFBuilderSetSeed(builder, 1);
|
|
CDForest::DFBuilderSetImportanceNone(builder);
|
|
CDForest::DFBuilderBuildRandomForest(builder, BASELINE_TREES, forest, frep);
|
|
Print(m_data.Id() + StringFormat(": forest built - %d trees, out-of-bag class error %.3f, OOB avg CE %.4f.",
|
|
BASELINE_TREES, frep.m_oobrelclserror, frep.m_oobavgce));
|
|
if(BaselineBudgetSpent(t0, "the MLP"))
|
|
return;
|
|
//--- MLP. ALGLIB's own multilayer perceptron on the identical matrix - the control that
|
|
//--- separates "our architecture is wrong" from "our IMPLEMENTATION is wrong".
|
|
CMultilayerPerceptron mlp;
|
|
CMLPReport mrep;
|
|
int mlpInfo = 0;
|
|
CMLPBase::MLPCreateC1(width, BASELINE_MLP_HIDDEN, 3, mlp);
|
|
uint mlpStart = GetTickCount();
|
|
CMLPTrain::MLPTrainLBFGS(mlp, xy, nTrain, BASELINE_MLP_DECAY, BASELINE_MLP_RESTARTS,
|
|
BASELINE_MLP_WSTEP, BASELINE_MLP_MAXITS, mlpInfo, mrep);
|
|
uint mlpMs = GetTickCount() - mlpStart;
|
|
bool mlpOK = (mlpInfo > 0);
|
|
//--- CAPACITY FIRST, because it decides what the training error below is allowed to mean.
|
|
int mlpWeights = (width + 1) * BASELINE_MLP_HIDDEN + (BASELINE_MLP_HIDDEN + 1) * 3;
|
|
if(mlpOK)
|
|
Print(m_data.Id() + StringFormat(": ALGLIB MLP trained in %.1f s - %d inputs -> %d hidden -> 3, %d gradient"
|
|
" evaluations, training class error %.3f, avg CE %.4f.%s",
|
|
mlpMs / 1000.0, width, BASELINE_MLP_HIDDEN, mrep.m_ngrad,
|
|
mrep.m_RelCLSError, mrep.m_AvgCE,
|
|
(mlpWeights >= nTrain)
|
|
? StringFormat(" NOTE: %d weights against %d rows - this fit can memorise"
|
|
" its own rows outright, so read the OOS score below and"
|
|
" not the training error.", mlpWeights, nTrain)
|
|
: ""));
|
|
else
|
|
Print(m_data.Id() + StringFormat(": ALGLIB MLP did not train (info %d) - skipped below.", mlpInfo));
|
|
if(BaselineBudgetSpent(t0, "the cross-validation"))
|
|
return;
|
|
//--- ...AND AN ERROR BAR ON IT. ReportBaselineModel gives every row a binomial SE, which is the
|
|
//--- sampling error of SCORING a fixed model. PREDICTED, not attempted. K-fold cross-validation
|
|
//--- measures that second variance directly, on the training rows only, so it never touches the
|
|
//--- OOS window the score below is taken on.
|
|
uint cvPredictedMs = mlpMs * BASELINE_MLP_CV_FOLDS;
|
|
uint cvSpentMs = GetTickCount() - t0;
|
|
bool cvFits = (cvSpentMs + cvPredictedMs < BASELINE_BUDGET_MS);
|
|
if(mlpOK && !cvFits)
|
|
Print(m_data.Id() + StringFormat(": ALGLIB MLP cross-validation SKIPPED - %d folds x %.1f s = ~%.1f s"
|
|
" predicted against %.1f s left of the budget. It is an error bar on the"
|
|
" score below, not the score itself, so the pass continues without it.",
|
|
BASELINE_MLP_CV_FOLDS, mlpMs / 1000.0, cvPredictedMs / 1000.0,
|
|
(BASELINE_BUDGET_MS - cvSpentMs) / 1000.0));
|
|
if(mlpOK && cvFits)
|
|
{
|
|
CMultilayerPerceptron cvNet;
|
|
CMLPReport cvTrainRep;
|
|
CMLPCVReport cvRep;
|
|
int cvInfo = 0;
|
|
CMLPBase::MLPCreateC1(width, BASELINE_MLP_HIDDEN, 3, cvNet);
|
|
CMLPTrain::MLPKFoldCVLBFGS(cvNet, xy, nTrain, BASELINE_MLP_DECAY, BASELINE_MLP_RESTARTS,
|
|
BASELINE_MLP_WSTEP, BASELINE_MLP_MAXITS, BASELINE_MLP_CV_FOLDS,
|
|
cvInfo, cvTrainRep, cvRep);
|
|
if(cvInfo > 0)
|
|
Print(m_data.Id() + StringFormat(": ALGLIB MLP %d-fold CV - held-out class error %.3f vs %.3f in"
|
|
" training (avg CE %.4f vs %.4f). A large gap is the fit memorising"
|
|
" its own rows, and it caps what the OOS row below can mean.",
|
|
BASELINE_MLP_CV_FOLDS, cvRep.m_RelCLSError, mrep.m_RelCLSError,
|
|
cvRep.m_AvgCE, mrep.m_AvgCE));
|
|
else
|
|
Print(m_data.Id() + StringFormat(": ALGLIB MLP cross-validation did not run (info %d).", cvInfo));
|
|
if(BaselineBudgetSpent(t0, "the linear fit"))
|
|
return;
|
|
}
|
|
//--- LINEAR. Same rows, same columns, target shifted from the class index to a signed direction.
|
|
double lrPredictedMs = BASELINE_LR_MS_AT_800 * MathPow(width / 800.0, 3.0);
|
|
double lrLeftMs = (double)BASELINE_BUDGET_MS - (double)(GetTickCount() - t0);
|
|
bool lrIdentified = (nTrain > width);
|
|
bool lrAffordable = (lrPredictedMs < lrLeftMs);
|
|
CLinearModel linear;
|
|
CLRReport lrep;
|
|
int lrInfo = 0;
|
|
bool linearOK = false;
|
|
if(!lrIdentified || !lrAffordable)
|
|
Print(m_data.Id() + StringFormat(": linear baseline SKIPPED - %s. The forest and the MLP below are"
|
|
" unaffected; only the linear row is missing.",
|
|
!lrIdentified
|
|
? StringFormat("%d rows against %d columns, so the normal equations are"
|
|
" singular and any coefficients returned would be one"
|
|
" arbitrary solution of infinitely many", nTrain, width)
|
|
: StringFormat("~%.0f s predicted at %d columns (cost grows as width^3)"
|
|
" against %.0f s left of the budget", lrPredictedMs / 1000.0,
|
|
width, lrLeftMs / 1000.0)));
|
|
else
|
|
{
|
|
for(int k = 0; k < nTrain; k++)
|
|
xy.Set(k, width, xy.Get(k, width) - 1.0);
|
|
CLinReg::LRBuild(xy, nTrain, width, lrInfo, linear, lrep);
|
|
linearOK = (lrInfo > 0);
|
|
if(!linearOK)
|
|
Print(m_data.Id() + StringFormat(": linear baseline did not fit (Alglib info %d) - forest only below.",
|
|
lrInfo));
|
|
}
|
|
if(BaselineBudgetSpent(t0, "the OOS scoring"))
|
|
return;
|
|
//--- SCORE BOTH ON THE OOS SLICE, in the deploy gate's own currency: a call is a directional
|
|
//--- prediction, a hit is that direction WINNING at the measured geometry (the view's Outcome),
|
|
//--- and the benchmark is the always-call-one-direction win rate over the same bars. See the
|
|
//--- chancePrecPct derivation in Train() for why that benchmark and not the label base rate.
|
|
int scored = 0, winLongN = 0, winShortN = 0;
|
|
int fCalls = 0, fHits = 0, lCalls = 0, lHits = 0, mCalls = 0, mHits = 0;
|
|
double y[], my[];
|
|
//--- THE NET'S OWN per-bar outcome, collected on the same walk for the regime test below. Paired
|
|
//--- with the bar index so the test can decluster before it computes a p-value.
|
|
double margins[];
|
|
int marginBars[];
|
|
ArrayResize(margins, 0);
|
|
ArrayResize(marginBars, 0);
|
|
for(int k = 0; k < nScoreWanted; k++)
|
|
{
|
|
int r = scoreBars[k];
|
|
if(!m_data.RowFeatures(r, width, x))
|
|
continue;
|
|
scored++;
|
|
bool wl = m_data.IsBuyLabel(r), ws = m_data.IsSellLabel(r);
|
|
if(wl) winLongN++;
|
|
if(ws) winShortN++;
|
|
//--- The arrow cache holds pass 3's ADJUSTED decision per scored bar, and pass 3 has just
|
|
//--- finished, so it is complete for this era at exactly this moment (it is wiped at the NEXT
|
|
//--- era start).
|
|
bool calledBuy = false;
|
|
double calledMag = 0.0;
|
|
if(m_data.DirectionalCall(r, calledBuy, calledMag))
|
|
{
|
|
int mi = ArraySize(margins);
|
|
ArrayResize(margins, mi + 1);
|
|
ArrayResize(marginBars, mi + 1);
|
|
//--- SIGNED BY WHETHER IT WAS RIGHT, not by which way it pointed: the regime test asks
|
|
//--- whether the model's confidence still tracks its outcomes, so a confident loser has
|
|
//--- to score negative.
|
|
margins[mi] = calledMag * ((calledBuy ? wl : ws) ? 1.0 : -1.0);
|
|
marginBars[mi] = r;
|
|
}
|
|
//--- Forest: argmax over the three class probabilities. A Neutral argmax is an abstention and
|
|
//--- enters neither the numerator nor the denominator, exactly as the net's own Neutral does.
|
|
CDForest::DFProcess(forest, x, y);
|
|
if(ArraySize(y) == 3)
|
|
{
|
|
int arg = 0;
|
|
for(int c = 1; c < 3; c++)
|
|
if(y[c] > y[arg])
|
|
arg = c;
|
|
if(arg != 1) // 1 == Neutral == abstain
|
|
{
|
|
fCalls++;
|
|
if(arg == 2 ? wl : ws)
|
|
fHits++;
|
|
}
|
|
}
|
|
//--- MLP: argmax over the softmax outputs, Neutral abstains - scored exactly as the forest is.
|
|
if(mlpOK)
|
|
{
|
|
CMLPBase::MLPProcess(mlp, x, my);
|
|
if(ArraySize(my) == 3)
|
|
{
|
|
int marg = 0;
|
|
for(int c = 1; c < 3; c++)
|
|
if(my[c] > my[marg])
|
|
marg = c;
|
|
if(marg != 1)
|
|
{
|
|
mCalls++;
|
|
if(marg == 2 ? wl : ws)
|
|
mHits++;
|
|
}
|
|
}
|
|
}
|
|
//--- Linear: the sign of the fitted value. No threshold sweep - a threshold fitted on this same
|
|
//--- slice is the calibration leak this codebase carved a purged band to avoid, and a baseline
|
|
//--- allowed one operating point per candidate would be the best-of-N problem all over again.
|
|
if(linearOK)
|
|
{
|
|
double pred = CLinReg::LRProcess(linear, x);
|
|
if(MathIsValidNumber(pred) && pred != 0.0)
|
|
{
|
|
lCalls++;
|
|
bool won = (pred > 0.0) ? wl : ws;
|
|
if(won)
|
|
lHits++;
|
|
}
|
|
}
|
|
}
|
|
if(scored < BASELINE_MIN_ROWS)
|
|
{
|
|
Print(m_data.Id() + StringFormat(": Alglib baselines scored only %d OOS windows - nothing to report.",
|
|
scored));
|
|
return;
|
|
}
|
|
double chancePct = 100.0 * MathMax(winLongN, winShortN) / scored;
|
|
ReportBaselineModel("forest", fCalls, fHits, scored, chancePct);
|
|
if(mlpOK)
|
|
ReportBaselineModel("alglib-mlp", mCalls, mHits, scored, chancePct);
|
|
if(linearOK)
|
|
ReportBaselineModel("linear", lCalls, lHits, scored, chancePct);
|
|
double gatePrec = -1.0, gateChance = -1.0;
|
|
int gateCalls = 0;
|
|
m_data.GateReference(gatePrec, gateCalls, gateChance);
|
|
Print(m_data.Id() + StringFormat(": baseline reference - the net's own gate on this chart last read"
|
|
" %.1f%% precision on %d calls against %.1f%% chance. Read the rows above"
|
|
" against it: all at chance means the MATRIX is the limit, not the topology;"
|
|
" alglib-mlp well ABOVE our own net on the same rows would point at our"
|
|
" IMPLEMENTATION rather than at the data.",
|
|
gatePrec, gateCalls, gateChance));
|
|
if(BaselineBudgetSpent(t0, "the redundancy / regime / lag reports"))
|
|
return;
|
|
ReportFeatureRedundancy(anchors, nTrain);
|
|
ReportRegimeStability(margins, marginBars);
|
|
ReportLinearLagProfile(bars, oosCutoff);
|
|
ReportCombiningWeights();
|
|
}
|
|
|
|
//+------------------------------------------------------------------+
|
|
//| WHAT WOULD THE BEST MIX OF MEMBERS HAVE BEEN? |
|
|
//+------------------------------------------------------------------+
|
|
void CBaselineComparator::ReportCombiningWeights(void)
|
|
{
|
|
//--- ONE MEMBER RUNS THIS. The vote rows are ensemble-global, so every member would otherwise fit
|
|
//--- and print the identical answer once each.
|
|
if(!m_data.IsEnsembleMember() || m_data.EnsembleIndex() != 0 || g_ensVoteRows < BASELINE_MIN_ROWS)
|
|
return;
|
|
//--- Which slots actually carry a member this run - the registry can be sparse.
|
|
int slot[ENS_MAX_MEMBERS], members = 0;
|
|
for(int m = 0; m < ENS_MAX_MEMBERS; m++)
|
|
{
|
|
bool seen = false;
|
|
for(int r = 0; r < g_ensVoteRows && !seen; r++)
|
|
if(g_ensVoteMember[r * ENS_MAX_MEMBERS + m] != 0.0)
|
|
seen = true;
|
|
if(seen)
|
|
slot[members++] = m;
|
|
}
|
|
if(members < 2)
|
|
return; // nothing to combine
|
|
//--- Declustered, for the reason set out in ReportRegimeStability(): overlapping labels are not
|
|
//--- independent draws, and a mixture fitted on all of them is fitted to ~L copies of each bar.
|
|
int gap = (int)MathMax(m_data.LabelResolutionBars(), 1);
|
|
double sMat[], tVec[];
|
|
ArrayResize(sMat, g_ensVoteRows * members);
|
|
ArrayResize(tVec, g_ensVoteRows);
|
|
int rows = 0, lastRow = -1;
|
|
for(int r = 0; r < g_ensVoteRows; r++)
|
|
{
|
|
if(lastRow >= 0 && (r - lastRow) < gap)
|
|
continue;
|
|
bool wl = g_ensVoteLabelBuy[r], ws = g_ensVoteLabelSell[r];
|
|
double t = (wl && !ws) ? 1.0 : ((ws && !wl) ? -1.0 : 0.0);
|
|
for(int m = 0; m < members; m++)
|
|
sMat[rows * members + m] = g_ensVoteMember[r * ENS_MAX_MEMBERS + slot[m]] / 100.0;
|
|
tVec[rows++] = t;
|
|
lastRow = r;
|
|
}
|
|
if(rows < BASELINE_MIN_ROWS)
|
|
{
|
|
Print(m_data.Id() + StringFormat(": combining weights - %d vote rows decluster to %d independent bars at"
|
|
" a %d-bar horizon, too few to fit %d weights on.",
|
|
g_ensVoteRows, rows, gap, members));
|
|
return;
|
|
}
|
|
//--- Start at the equal mix, which is also the honest null: "no member deserves more than another".
|
|
double w[], bndl[], bndu[];
|
|
ArrayResize(w, members);
|
|
ArrayResize(bndl, members);
|
|
ArrayResize(bndu, members);
|
|
for(int m = 0; m < members; m++)
|
|
{
|
|
w[m] = 1.0 / members;
|
|
bndl[m] = 0.0;
|
|
bndu[m] = 1.0;
|
|
}
|
|
double sseEqual = CombinerSSE(w, members, sMat, tVec, rows);
|
|
//--- sum(w) == 1, as one linear equality row: [1 1 ... 1 | 1] with ct = 0 meaning "=".
|
|
CMatrixDouble lc(1, members + 1);
|
|
int ct[];
|
|
ArrayResize(ct, 1);
|
|
ct[0] = 0;
|
|
for(int m = 0; m < members; m++)
|
|
lc.Set(0, m, 1.0);
|
|
lc.Set(0, members, 1.0);
|
|
CMinBLEICState state;
|
|
CMinBLEICReport rep;
|
|
//--- CreateF, i.e. NUMERICAL gradients: with at most ENS_MAX_MEMBERS unknowns a finite-difference
|
|
//--- gradient is members+1 evaluations, which is nothing against the alternative of hand-deriving
|
|
//--- and maintaining an analytic one for a diagnostic.
|
|
CMinBLEIC::MinBLEICCreateF(members, w, 1.0e-6, state);
|
|
CMinBLEIC::MinBLEICSetBC(state, bndl, bndu);
|
|
CMinBLEIC::MinBLEICSetLC(state, lc, ct, 1);
|
|
CMinBLEIC::MinBLEICSetCond(state, 0.0, 0.0, 1.0e-8, 200);
|
|
//--- Driven by hand rather than through MinBLEICOptimize's delegate: the objective needs the two
|
|
//--- local matrices above, and a CNDimensional_Func subclass would have to smuggle them through a
|
|
//--- CObject to reach them.
|
|
double trial[];
|
|
ArrayResize(trial, members);
|
|
while(CMinBLEIC::MinBLEICIteration(state))
|
|
{
|
|
if(!state.m_needf)
|
|
continue;
|
|
for(int m = 0; m < members; m++)
|
|
trial[m] = state.m_x[m];
|
|
state.m_f = CombinerSSE(trial, members, sMat, tVec, rows);
|
|
}
|
|
double fitted[];
|
|
CMinBLEIC::MinBLEICResults(state, fitted, rep);
|
|
if(rep.m_terminationtype <= 0 || ArraySize(fitted) < members)
|
|
{
|
|
Print(m_data.Id() + StringFormat(": combining weights - MinBLEIC did not converge (termination %d).",
|
|
rep.m_terminationtype));
|
|
return;
|
|
}
|
|
double sseFit = CombinerSSE(fitted, members, sMat, tVec, rows);
|
|
string line = "";
|
|
for(int m = 0; m < members; m++)
|
|
//--- Slot index, not a name: g_aiSignals[] is declared in Warrior_EA.mq5, which is compiled
|
|
//--- AFTER this partial, so the registry is not reachable from here. The startup census line
|
|
//--- maps index to model id.
|
|
line += StringFormat(" m%d=%.3f", slot[m], fitted[m]);
|
|
//--- ...and the same mixture judged on the decision the EA actually makes: sign of the blend
|
|
//--- against which direction paid. This is the number that decides whether the fit meant anything.
|
|
int fitCalls = 0, fitHits = 0, eqCalls = 0, eqHits = 0;
|
|
for(int r = 0; r < rows; r++)
|
|
{
|
|
double bf = 0.0, be = 0.0;
|
|
for(int m = 0; m < members; m++)
|
|
{
|
|
bf += fitted[m] * sMat[r * members + m];
|
|
be += sMat[r * members + m] / members;
|
|
}
|
|
if(bf != 0.0 && tVec[r] != 0.0)
|
|
{
|
|
fitCalls++;
|
|
if((bf > 0.0) == (tVec[r] > 0.0))
|
|
fitHits++;
|
|
}
|
|
if(be != 0.0 && tVec[r] != 0.0)
|
|
{
|
|
eqCalls++;
|
|
if((be > 0.0) == (tVec[r] > 0.0))
|
|
eqHits++;
|
|
}
|
|
}
|
|
Print(m_data.Id() + StringFormat(": combining weights over %d independent bars -%s | SSE %.4f fitted vs"
|
|
" %.4f at the equal mix | directional hit rate %.1f%% (%d) vs %.1f%% (%d)"
|
|
" equal-weighted. %s",
|
|
rows, line, sseFit, sseEqual,
|
|
(fitCalls > 0 ? 100.0 * fitHits / fitCalls : 0.0), fitCalls,
|
|
(eqCalls > 0 ? 100.0 * eqHits / eqCalls : 0.0), eqCalls,
|
|
"Reported only - the live weights are unchanged, because moving them moves"
|
|
" the vote the deploy gate certifies."));
|
|
}
|
|
|
|
//+------------------------------------------------------------------+
|
|
//| Mean squared error of a weighted blend against the signed outcome.|
|
|
//| The objective MinBLEIC minimises above, and the number the equal |
|
|
//| mix is scored on, so both readings come from one expression. |
|
|
//+------------------------------------------------------------------+
|
|
double CBaselineComparator::CombinerSSE(const double &w[], const int members, const double &s[],
|
|
const double &t[], const int rows)
|
|
{
|
|
if(rows <= 0 || members <= 0)
|
|
return 0.0;
|
|
double acc = 0.0;
|
|
for(int r = 0; r < rows; r++)
|
|
{
|
|
double blend = 0.0;
|
|
for(int m = 0; m < members; m++)
|
|
blend += w[m] * s[r * members + m];
|
|
double e = blend - t[r];
|
|
acc += e * e;
|
|
}
|
|
return acc / rows;
|
|
}
|
|
|
|
//+------------------------------------------------------------------+
|
|
//| EVERY LAG AT ONCE, via FFT cross-correlation. |
|
|
//+------------------------------------------------------------------+
|
|
void CBaselineComparator::ReportLinearLagProfile(const int bars, const int oosCutoff)
|
|
{
|
|
int nvars = m_data.FeaturesPerBar();
|
|
int width = m_data.HistoryBars() * m_data.FeaturesPerBar();
|
|
if(nvars < 1 || width <= 0)
|
|
return;
|
|
//--- Newest-first walk over a contiguous OOS stretch, capped so one FFT stays cheap.
|
|
int want = (int)MathMin(BASELINE_MAX_SCORE_ROWS, oosCutoff);
|
|
double x[], series[], target[];
|
|
ArrayResize(x, width);
|
|
ArrayResize(series, want * nvars);
|
|
ArrayResize(target, want);
|
|
int n = 0;
|
|
//--- SKIP THE LEADING UNRESOLVED EDGE BEFORE DEMANDING CONTIGUITY. r=0 is the NEWEST bar, and a
|
|
//--- forward-looking pivot label cannot be resolved there by construction - the opposite pivot has
|
|
//--- not committed yet. Breaking on the first hole therefore ended the walk at r=0 every single
|
|
//--- time, and this diagnostic has never once run: it reported "only 0 contiguous OOS bars" on
|
|
//--- every chart, every era, which reads as "not enough data" rather than "never started".
|
|
//--- The leading gap is SYSTEMATIC (always ~the label resolution), not a hole in the middle, so
|
|
//--- stepping over it splices nothing - contiguity is still enforced from the first labelled row on.
|
|
int start = 0;
|
|
while(start < oosCutoff && start + m_data.HistoryBars() - 1 < bars && !m_data.HasLabel(start))
|
|
start++;
|
|
if(start > 0)
|
|
PrintVerbose(m_data.Id() + StringFormat(": lag profile - skipped %d unresolved bar(s) at the"
|
|
" newest edge before starting the contiguous walk.", start));
|
|
for(int r = start; r < oosCutoff && n < want; r++)
|
|
{
|
|
if(r + m_data.HistoryBars() - 1 >= bars)
|
|
break;
|
|
if(!m_data.HasLabel(r))
|
|
break; // a hole breaks contiguity - stop rather than splice across it
|
|
if(!m_data.RowFeatures(r, width, x))
|
|
break;
|
|
for(int f = 0; f < nvars; f++)
|
|
series[n * nvars + f] = x[width - nvars + f];
|
|
//--- Signed label. Neutral bars carry no direction and enter as 0 rather than being dropped,
|
|
//--- which would break the contiguity a lag index depends on.
|
|
bool wl = m_data.IsBuyLabel(r), ws = m_data.IsSellLabel(r);
|
|
target[n] = (wl && !ws) ? 1.0 : ((ws && !wl) ? -1.0 : 0.0);
|
|
n++;
|
|
}
|
|
if(n < BASELINE_MIN_ROWS)
|
|
{
|
|
//--- Say WHERE it ran out, not just that it did. "0 from r=0" is a walk that never started
|
|
//--- (the old, permanent failure); "0 from r=37" is a genuinely short or hole-ridden window.
|
|
Print(m_data.Id() + StringFormat(": linear lag profile skipped - only %d contiguous OOS bar(s)"
|
|
" from r=%d (need %d; OOS cutoff %d, %d bars, history %d).",
|
|
n, start, BASELINE_MIN_ROWS, oosCutoff, bars, m_data.HistoryBars()));
|
|
return;
|
|
}
|
|
//--- The band a correlation of exactly nothing still fluctuates inside, ~1/sqrt(n) per lag. Read
|
|
//--- against the MAXIMUM over all lags and every column, so it is a family of n*nvars draws - the
|
|
//--- best-of-N null, not a single one. Sidak over that family, not a bare 2-sigma line.
|
|
//--- CAP THE LAG RANGE TO THE QUESTION BEING ASKED. This profile exists to say how deep a lookback
|
|
//--- carries linear structure - the input window is HistoryBars, so lags an order of magnitude past
|
|
//--- it answer nothing and are where this estimator is at its worst (see the per-lag band below).
|
|
//--- Uncapped, the first run of this diagnostic reported its maximum at lag +3900 of 3999 on BOTH
|
|
//--- XTIUSD and SP500, same column, |r| within 0.01 of each other - two independent instruments
|
|
//--- cannot agree to that precision at a 3900-bar lag, which is what identified it as an artifact.
|
|
int maxLag = (int)MathMin(n - BASELINE_MIN_ROWS, MathMax(4 * m_data.HistoryBars(), 50));
|
|
if(maxLag < 1)
|
|
maxLag = 1;
|
|
//--- THE NO-INFORMATION BAND IS PER-LAG, NOT GLOBAL. CorrR1D is NON-circular, so lag k is computed
|
|
//--- from only (n - |k|) overlapping terms while the normalising denominator uses all n. Scoring
|
|
//--- every lag against a single 1/sqrt(n) band therefore understates the band at large |k| by
|
|
//--- sqrt(n/(n-|k|)) - a factor of 6 at the edge - which is exactly how a marginal correlation on a
|
|
//--- 100-sample overlap was reported as "32.4 SE". Each lag now gets its own band.
|
|
//--- AND the overlap deflation applies here as well: labels persist ~LabelResolutionBars bars, so
|
|
//--- neighbouring rows are not independent draws - see [[project_label_overlap_effective_n]].
|
|
double lifespan = MathMax(1.0, (double)m_data.LabelResolutionBars());
|
|
double bestAbs = 0.0, bestZ = 0.0, bestNoise = 0.0;
|
|
int bestLag = 0, bestCol = -1;
|
|
double tgt[], col[], corr[];
|
|
ArrayResize(tgt, n);
|
|
for(int i = 0; i < n; i++)
|
|
tgt[i] = target[i];
|
|
ArrayResize(col, n);
|
|
for(int f = 0; f < nvars; f++)
|
|
{
|
|
for(int i = 0; i < n; i++)
|
|
col[i] = series[i * nvars + f];
|
|
//--- Both series are mean-centred first: CorrR1D is a raw sum of products, so a column with a
|
|
//--- non-zero mean would return its own mean times the target's at every lag and swamp the
|
|
//--- structure being looked for.
|
|
double mc = MathMean(col), mt = MathMean(tgt);
|
|
double cc[], tc[];
|
|
ArrayResize(cc, n);
|
|
ArrayResize(tc, n);
|
|
double sc = 0.0, st = 0.0;
|
|
for(int i = 0; i < n; i++)
|
|
{
|
|
cc[i] = col[i] - mc;
|
|
tc[i] = tgt[i] - mt;
|
|
sc += cc[i] * cc[i];
|
|
st += tc[i] * tc[i];
|
|
}
|
|
if(sc <= 0.0 || st <= 0.0)
|
|
continue; // a constant column correlates with nothing
|
|
double denom = MathSqrt(sc * st);
|
|
CCorr::CorrR1D(cc, n, tc, n, corr);
|
|
int cn = ArraySize(corr);
|
|
for(int k = 0; k < cn; k++)
|
|
{
|
|
//--- CorrR1D returns the non-circular correlation with negative lags first; index n-1 is
|
|
//--- lag 0, so this is the offset in bars with its sign.
|
|
int lag = k - (n - 1);
|
|
if(MathAbs(lag) > maxLag)
|
|
continue;
|
|
double rr = MathAbs(corr[k]) / denom;
|
|
//--- Overlapping terms at this lag, deflated for label persistence -> this lag's own band.
|
|
double effN = MathMax(2.0, (double)(n - MathAbs(lag)) / lifespan);
|
|
double band = 1.0 / MathSqrt(effN);
|
|
double zk = rr / band;
|
|
//--- RANKED ON z, NOT ON |r|. The whole point of a per-lag band is that two correlations of
|
|
//--- equal size are NOT equally surprising when they rest on different numbers of terms.
|
|
if(zk > bestZ)
|
|
{
|
|
bestZ = zk;
|
|
bestAbs = rr;
|
|
bestNoise = band;
|
|
bestLag = lag;
|
|
bestCol = f;
|
|
}
|
|
}
|
|
if(m_data.Stopping())
|
|
return;
|
|
}
|
|
if(bestCol < 0)
|
|
{
|
|
Print(m_data.Id() + ": linear lag profile - every column was constant over the window.");
|
|
return;
|
|
}
|
|
//--- The family is what was actually SEARCHED: every column at every lag inside the cap.
|
|
int lagsTried = 2 * maxLag + 1;
|
|
double pFamily = SidakFamilyP(bestZ, lagsTried * nvars);
|
|
Print(m_data.Id() + StringFormat(": linear lag profile - strongest of %d columns x %d lags (|lag|<=%d,"
|
|
" %d rows, label lifespan %.1f) is column %d at lag %+d, |r| %.4f"
|
|
" (%.1f SE of that lag's own %.4f band). Family-wise p %.4f -> %s.",
|
|
nvars, lagsTried, maxLag, n, lifespan, bestCol, bestLag, bestAbs,
|
|
bestZ, bestNoise, pFamily,
|
|
(pFamily <= DEPLOY_FAMILY_WISE_ALPHA
|
|
? "SURVIVES the best-of-N null - worth a look"
|
|
: "consistent with no linear structure at any lag in the window")));
|
|
}
|
|
|
|
//+------------------------------------------------------------------+
|
|
//| HOW WIDE IS THE MATRIX REALLY? Two readings of the same |
|
|
//| question, both on the anchor bar's own feature row. |
|
|
//+------------------------------------------------------------------+
|
|
void CBaselineComparator::ReportFeatureRedundancy(CMatrixDouble &rows, const int n)
|
|
{
|
|
int nvars = m_data.FeaturesPerBar();
|
|
if(n < BASELINE_MIN_ROWS || nvars < 2)
|
|
return;
|
|
//--- PAIRWISE. The worst offender by absolute correlation, plus how many pairs are effectively
|
|
//--- one column wearing two names.
|
|
CMatrixDouble corr;
|
|
if(CBaseStat::PearsonCorrM(rows, n, nvars, corr))
|
|
{
|
|
int pairs = 0, wi = -1, wj = -1;
|
|
double worst = 0.0;
|
|
for(int i = 0; i < nvars; i++)
|
|
for(int j = i + 1; j < nvars; j++)
|
|
{
|
|
double r = MathAbs(corr.Get(i, j));
|
|
if(!MathIsValidNumber(r))
|
|
continue;
|
|
if(r >= BASELINE_COLLINEAR_ABS_R)
|
|
pairs++;
|
|
if(r > worst)
|
|
{
|
|
worst = r;
|
|
wi = i;
|
|
wj = j;
|
|
}
|
|
}
|
|
Print(m_data.Id() + StringFormat(": matrix redundancy - %d of %d column pairs correlate above %.2f;"
|
|
" worst is columns %d/%d at |r| %.3f.",
|
|
pairs, nvars * (nvars - 1) / 2, BASELINE_COLLINEAR_ABS_R, wi, wj, worst));
|
|
}
|
|
//--- EFFECTIVE WIDTH. PCABuildBasis returns the component variances in decreasing order, so the
|
|
//--- count needed to reach BASELINE_PCA_VAR_SHARE is the number of directions the columns really
|
|
//--- span. Far cheaper here than on the full window: nvars^2, not (historyBars*nvars)^2.
|
|
int info = 0;
|
|
double s2[];
|
|
CMatrixDouble basis;
|
|
CPCAnalysis::PCABuildBasis(rows, n, nvars, info, s2, basis);
|
|
if(info <= 0 || ArraySize(s2) < nvars)
|
|
{
|
|
Print(m_data.Id() + StringFormat(": matrix redundancy - PCA did not converge (info %d).", info));
|
|
return;
|
|
}
|
|
double total = 0.0;
|
|
for(int i = 0; i < nvars; i++)
|
|
total += MathMax(0.0, s2[i]);
|
|
if(total <= 0.0)
|
|
return;
|
|
double cum = 0.0;
|
|
int need = nvars;
|
|
for(int i = 0; i < nvars; i++)
|
|
{
|
|
cum += MathMax(0.0, s2[i]);
|
|
if(cum / total >= BASELINE_PCA_VAR_SHARE)
|
|
{
|
|
need = i + 1;
|
|
break;
|
|
}
|
|
}
|
|
Print(m_data.Id() + StringFormat(": matrix redundancy - %d of %d columns carry %.0f%% of the variance"
|
|
" (top component alone %.0f%%). The window multiplies this by %d lags,"
|
|
" so the net sees ~%d effective inputs, not %d.",
|
|
need, nvars, BASELINE_PCA_VAR_SHARE * 100.0, 100.0 * s2[0] / total,
|
|
m_data.HistoryBars(), need * m_data.HistoryBars(),
|
|
m_data.HistoryBars() * nvars));
|
|
}
|
|
|
|
//+------------------------------------------------------------------+
|
|
//| HAS THE EDGE MOVED? Mann-Whitney U on the net's own signed |
|
|
//| margin (conviction, signed by whether the conviction paid), |
|
|
//| older half of the OOS window against newer. |
|
|
//+------------------------------------------------------------------+
|
|
void CBaselineComparator::ReportRegimeStability(double &margins[], int &marginBars[])
|
|
{
|
|
int n = ArraySize(margins);
|
|
if(n < 2 * BASELINE_MIN_REGIME_N)
|
|
return;
|
|
//--- One sample per label resolution lag. These are MQL5 timeseries indices, so the array ascends
|
|
//--- in index while descending in TIME - element 0 is the newest scored call.
|
|
int gap = (int)MathMax(m_data.LabelResolutionBars(), 1);
|
|
double kept[];
|
|
ArrayResize(kept, n);
|
|
int m = 0, lastBar = -1;
|
|
for(int i = 0; i < n; i++)
|
|
{
|
|
if(lastBar >= 0 && MathAbs(marginBars[i] - lastBar) < gap)
|
|
continue;
|
|
kept[m++] = margins[i];
|
|
lastBar = marginBars[i];
|
|
}
|
|
if(m < 2 * BASELINE_MIN_REGIME_N)
|
|
{
|
|
Print(m_data.Id() + StringFormat(": regime test skipped - %d scored calls decluster to only %d"
|
|
" independent samples at a %d-bar label horizon, below the %d per half"
|
|
" the test needs.", n, m, gap, BASELINE_MIN_REGIME_N));
|
|
return;
|
|
}
|
|
//--- kept[] is newest-first, so the SECOND half is the older window.
|
|
int half = m / 2;
|
|
double recent[], older[];
|
|
ArrayResize(recent, half);
|
|
ArrayResize(older, m - half);
|
|
for(int i = 0; i < half; i++)
|
|
recent[i] = kept[i];
|
|
for(int i = half; i < m; i++)
|
|
older[i - half] = kept[i];
|
|
double both = 1.0, left = 1.0, right = 1.0;
|
|
CMannWhitneyU::CMannWhitneyUTest(older, m - half, recent, half, both, left, right);
|
|
double meanOld = MathMean(older), meanNew = MathMean(recent);
|
|
Print(m_data.Id() + StringFormat(": regime test - signed margin over %d independent OOS calls"
|
|
" (%d older / %d newer, declustered from %d at a %d-bar horizon):"
|
|
" mean %+.3f -> %+.3f, Mann-Whitney p %.4f -> %s.",
|
|
m, m - half, half, n, gap, meanOld, meanNew, both,
|
|
(both <= 0.05 ? (meanNew < meanOld ? "DISTRIBUTION SHIFTED, and downward"
|
|
: "distribution shifted, upward")
|
|
: "no detectable shift")));
|
|
}
|
|
|
|
//+------------------------------------------------------------------+
|
|
//| One baseline's verdict, through the deploy gate's arithmetic. |
|
|
//+------------------------------------------------------------------+
|
|
void CBaselineComparator::ReportBaselineModel(const string label, const int calls, const int hits,
|
|
const int scored, const double chancePct)
|
|
{
|
|
if(calls <= 0)
|
|
{
|
|
Print(m_data.Id() + StringFormat(": baseline %s - abstained on all %d OOS windows. No edge measurable.",
|
|
label, scored));
|
|
return;
|
|
}
|
|
double precPct = 100.0 * hits / calls;
|
|
//--- DEFLATED, like the NN's own DEPLOY BAR is. These calls are consecutive bars sharing
|
|
//--- overlapping labels, not independent trades, so raw n understates the SE by ~sqrt(L) - and
|
|
//--- this line exists to be compared against the net measured on the SAME windows. Judging the
|
|
//--- baseline on the more permissive standard is how a forest at +1.0 SE reads as +3.6.
|
|
double effN = m_data.EffectiveSampleSize((double)calls);
|
|
double se = BinomialSEPct(chancePct / 100.0, effN);
|
|
if(se <= 0.0)
|
|
{
|
|
Print(m_data.Id() + StringFormat(": baseline %s - %.1f%% on %d calls, but chance %.1f%% is degenerate"
|
|
" so no test applies.", label, precPct, calls, chancePct));
|
|
return;
|
|
}
|
|
double z = (precPct - chancePct) / se;
|
|
double pFamily = SidakFamilyP(z, BASELINE_MODELS_TRIED);
|
|
Print(m_data.Id() + StringFormat(": baseline %s - %.1f%% precision on %d calls (%.0f%% coverage, worth"
|
|
" %.0f INDEPENDENT ones after the label overlap) vs %.1f%% chance | edge"
|
|
" %+.1fpp = %+.2f SE | family-wise p %.4f -> %s.",
|
|
label, precPct, calls, 100.0 * calls / scored, effN, chancePct,
|
|
precPct - chancePct, z, pFamily,
|
|
(pFamily <= DEPLOY_FAMILY_WISE_ALPHA ? "CLEARS the deploy gate"
|
|
: "does not clear")));
|
|
}
|
|
|
|
#endif
|