forked from animatedread/Warrior_EA
STrainPoolHeader wrote its fingerprint into a FILE_BIN stream as FileWriteString(h, fingerprint + "\n") and read it back with FileReadString(h) - no length argument. In binary mode FileWriteString emits the characters raw: no length prefix, no terminator, and "\n" is just another character rather than a delimiter anything honours. The reader had nothing to stop at, over-read into the float rows that follow, and returned the fingerprint plus a few bytes of binary garbage - so `fingerprint != wantFp` could never succeed between two genuinely identical models. Verified in the bytes rather than inferred: xxd on a v1 file shows three ints then the fingerprint starting immediately at offset 12 with no count in front of it, and EURUSD/USDJPY/USDCAD all stored width 624 with byte-identical fingerprints while each one's log rejected the other two as "different model fingerprint". The StringReplace on "\n" is the tell that a delimiter was intended. Cross-asset-class peers really are incompatible and always will be - FX majors carry XA:6, indices/metals/oil carry XA:6:IDX2, giving widths 600/612/624 - which is why the reject list looked plausible and this went unread. The three FX majors were always poolable and never pooled. Length-prefixes the string, bounds-checks the count before sizing a read from it, and bumps TRAINPOOL_RECORD_VERSION 1 -> 2 so existing files are refused by the version gate with a reason instead of being misread. Also documents, without changing, why Signal_ThresholdOpen is now a unanimity rule: the vote is a weighted mean of tier weights, those fell from ~70 to ~30 with the pivot-event label, so PCT_25 went from ~36% of the reachable ceiling to ~83%. Measured: all 6 symbols clear their precision bar, 4 of 6 fail only on coverage, and coverage decays 6.8% -> 2.2% over 35 eras as the models specialise - which shrinks effN and so RAISES the deploy bar at flat precision. PCT_20 (a 3-of-4 quorum) is the indicated change but is left unmade: MT5 stores input values per chart in profiles\Charts\*\chart*.chr, so an already-attached EA ignores this default entirely - confirmed by a full close/recompile/relaunch cycle after which the log still read "fired at vote>=25%". Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
530 lines
25 KiB
MQL5
530 lines
25 KiB
MQL5
//+------------------------------------------------------------------+
|
|
//| TrainingPool.mqh |
|
|
//| Cross-instrument TRAINING rows, not just a cross-instrument |
|
|
//| gate. PooledGate.mqh pools the DECISION; this pools the DATA. |
|
|
//| |
|
|
//| WHY, MEASURED |
|
|
//| Measured 2026-08-24 in research/edge.py on 15 instruments over |
|
|
//| 7 distinct markets, both arms sharing calendar folds, exit-time |
|
|
//| purge, benchmark and scoring so TRAINING BREADTH is the only |
|
|
//| variable: |
|
|
//| |
|
|
//| H4 k=2 pooled +1.18pp vs per-instrument -0.62pp, gap +2.02pp |
|
|
//| t_mkt 3.97 -> CLEARS the Sidak bar of 3.69 over the |
|
|
//| five feature sets at df=6 |
|
|
//| D1 k=1 gap +2.03pp, t_mkt 2.79 - an independent replication |
|
|
//| on a different timeframe AND a different barrier |
|
|
//| |
|
|
//| The per-instrument arm was NEGATIVE on every feature set at both |
|
|
//| timeframes: it does not merely fail to beat "always take the |
|
|
//| drift side", it loses to it. This EA trains one net per chart, |
|
|
//| which is that arm. |
|
|
//| |
|
|
//| WHY ROWS AND NOT SYMBOLS |
|
|
//| Pointing the feature stack at another symbol needs per-symbol |
|
|
//| indicator handles, and this project has been bitten there twice |
|
|
//| (the handle leak that never released the old handle, and the |
|
|
//| twelve "dead" handles that were one shared refcounted iMA). |
|
|
//| Instead each chart computes its OWN features with its OWN handles |
|
|
//| and shares the resulting NUMBERS - the idiom PooledGate already |
|
|
//| uses for gate statistics. Nothing here touches an indicator. |
|
|
//| |
|
|
//| Sound only because the feature vector is already scale-free: |
|
|
//| FeatureBuilder ATR-normalises every price-unit feature (see its |
|
|
//| "instead of feeding e.g. 0.0005 on EURUSD" comment) and scales |
|
|
//| volume against the prior bar. A XAUUSD row and a USDJPY row are |
|
|
//| therefore in the same units. |
|
|
//| |
|
|
//| NOT A FINGERPRINT PARTICIPANT. Pooling changes what the model is |
|
|
//| TRAINED ON, not what it IS - topology, feature layout and input |
|
|
//| width are untouched - so adding it to BuildModelFingerprint would |
|
|
//| re-key every .nnw to record something outside the model's |
|
|
//| identity. The fingerprint instead GATES adoption: it is precisely |
|
|
//| the assertion "column k means the same thing in your file as in |
|
|
//| mine". |
|
|
//+------------------------------------------------------------------+
|
|
#property strict
|
|
#ifndef WARRIOR_TRAINING_TRAININGPOOL_MQH
|
|
#define WARRIOR_TRAINING_TRAININGPOOL_MQH
|
|
|
|
#include <Arrays\ArrayDouble.mqh>
|
|
#include "..\..\System\AtomicFile.mqh"
|
|
|
|
//--- Bumped whenever the record layout changes. A reader finding a different version SKIPS the file
|
|
//--- rather than misreading its columns - the same discipline as POOL_RECORD_VERSION, and the same
|
|
//--- reason: a silently misread column is indistinguishable from a feature that stopped working.
|
|
//--- 1 -> 2 (2026-08-26): the header's fingerprint is now LENGTH-PREFIXED. See STrainPoolHeader::
|
|
//--- Write() for why version 1's unprefixed write could not be read back exactly, and therefore why
|
|
//--- every v1 file on disk holds a fingerprint this build must not compare against. The version gate
|
|
//--- in MismatchReason() is what makes that safe: a v1 file is refused with a reason, not misread.
|
|
#define TRAINPOOL_RECORD_VERSION 2
|
|
#define TRAINPOOL_DIR "Warrior_EA\\TrainPool"
|
|
//--- Ceiling on rows adopted from ALL peers combined. Training time is linear in this and the MQL5
|
|
//--- trainer is single-threaded, so an uncapped pool turns a 20-minute era into an overnight one.
|
|
//--- Rows past the cap are dropped BY UNIFORM STRIDE, never truncated: taking the first N would
|
|
//--- adopt only the oldest slice of the alphabetically-first peers and call it a pool.
|
|
#define TRAINPOOL_MAX_ROWS 60000
|
|
//--- A peer file older than this is ignored. Generous, unlike the gate's 12 hours, because a training
|
|
//--- row does not go stale the way a live gate record does - a bar from last week is still a valid
|
|
//--- observation. This exists to stop a decommissioned symbol training this model forever.
|
|
#define TRAINPOOL_MAX_AGE_DAYS 30
|
|
//--- Minimum wall-clock gap between two publishes of the same chart's corpus. An era over a WARM
|
|
//--- feature cache can finish in a fraction of a second (observed 2026-08-24: ~0.3s/era on H4 SP500),
|
|
//--- and every era re-derives the SAME rows from the same in-sample span - so publishing per era would
|
|
//--- rewrite a multi-megabyte file continuously for no new information. Peers re-read on their own era
|
|
//--- boundary, so a few minutes of staleness costs a pool nothing.
|
|
#define TRAINPOOL_MIN_PUBLISH_SEC 300
|
|
|
|
//--- One place that knows the file layout. Both sides use it, so a layout change cannot be applied to
|
|
//--- the writer and missed in the reader - which is the failure this codebase has hit three times
|
|
//--- under different names (renamed field, moved scope, deleted variable).
|
|
struct STrainPoolHeader
|
|
{
|
|
int version;
|
|
int width;
|
|
int rows;
|
|
string fingerprint;
|
|
|
|
//--- LENGTH-PREFIXED, because this is a BINARY file and a binary file has no line ends.
|
|
//---
|
|
//--- Version 1 wrote `fingerprint + "\n"` and read it back with FileReadString(h) - no length
|
|
//--- argument. In a FILE_BIN stream FileWriteString emits the characters RAW: no length prefix,
|
|
//--- no terminator, and the "\n" is just another character, not a delimiter anything honours.
|
|
//--- The reader therefore had nothing to stop at and over-read into the float rows that follow,
|
|
//--- so the fingerprint came back as itself PLUS a few characters of binary garbage and the
|
|
//--- `fingerprint != wantFp` test below could never succeed between two genuinely identical
|
|
//--- models. Verified in the bytes, not inferred: at offset 12 of every v1 file the fingerprint
|
|
//--- starts immediately after the three ints with no count in front of it.
|
|
//---
|
|
//--- IT MADE THE WHOLE CROSS-INSTRUMENT POOL INERT for the case it exists to serve. Measured
|
|
//--- 2026-08-26: EURUSD, USDJPY and USDCAD all stored width 624 and BYTE-IDENTICAL fingerprints,
|
|
//--- and every one of them rejected the other two as "different model fingerprint". The
|
|
//--- StringReplace on "\n" is the tell that a delimiter was intended; binary mode never gave one.
|
|
void Write(const int h) const
|
|
{
|
|
FileWriteInteger(h, version, INT_VALUE);
|
|
FileWriteInteger(h, width, INT_VALUE);
|
|
FileWriteInteger(h, rows, INT_VALUE);
|
|
FileWriteInteger(h, StringLen(fingerprint), INT_VALUE);
|
|
FileWriteString(h, fingerprint);
|
|
}
|
|
|
|
void Read(const int h)
|
|
{
|
|
version = FileReadInteger(h, INT_VALUE);
|
|
width = FileReadInteger(h, INT_VALUE);
|
|
rows = FileReadInteger(h, INT_VALUE);
|
|
//--- Read the count BEFORE trusting it: a truncated or foreign file must not size a read from
|
|
//--- a garbage length. The cap is generous against the ~200-character fingerprints this
|
|
//--- project builds and still bounds the damage.
|
|
int fpLen = FileReadInteger(h, INT_VALUE);
|
|
if(fpLen <= 0 || fpLen > 4096)
|
|
{
|
|
fingerprint = "";
|
|
//--- Force MismatchReason() to refuse rather than compare against a half-read string. A
|
|
//--- version that no build writes is the honest description of a header this one cannot
|
|
//--- parse, and it is the first thing that predicate tests.
|
|
version = -1;
|
|
return;
|
|
}
|
|
fingerprint = FileReadString(h, fpLen);
|
|
}
|
|
|
|
//--- The whole adoption gate in one predicate, and it says WHY when it refuses. "" means adopt.
|
|
//--- Width is checked as well as fingerprint because a fingerprint match with a width mismatch
|
|
//--- means one side's .nnw pinned an older layout - the persisted-architecture trap.
|
|
//---
|
|
//--- The reason is not decoration. Two charts of DIFFERENT SYMBOLS do not generally share a
|
|
//--- fingerprint here: NeuronsCount counts the alt-data columns, which are per-symbol (SP500
|
|
//--- carries cot_spec_net, the FX majors carry cot_idx_1y/3y/chg_4w), and the cross-asset block
|
|
//--- appends ":IDX2" when base currency == profit currency, which is true of an index and false
|
|
//--- of a pair. So the first thing an operator needs from a pool that adopted nothing is which
|
|
//--- of those it was - and a silent skip is exactly the failure this project has paid for under
|
|
//--- three other names.
|
|
string MismatchReason(const string wantFp, const int wantWidth) const
|
|
{
|
|
if(version != TRAINPOOL_RECORD_VERSION)
|
|
return StringFormat("record version %d, this build writes %d", version,
|
|
TRAINPOOL_RECORD_VERSION);
|
|
if(rows <= 0)
|
|
return "no rows";
|
|
if(width != wantWidth)
|
|
return StringFormat("%d features per row, this model needs %d (different feature layout -"
|
|
" usually a different alt-data column set)", width, wantWidth);
|
|
if(fingerprint != wantFp)
|
|
return "different model fingerprint " + fingerprint;
|
|
return "";
|
|
}
|
|
|
|
bool Compatible(const string wantFp, const int wantWidth) const
|
|
{ return MismatchReason(wantFp, wantWidth) == ""; }
|
|
};
|
|
|
|
string TrainPoolPath(const string symbol, const int period)
|
|
{
|
|
return TRAINPOOL_DIR + "\\" + symbol + "_" + IntegerToString(period) + ".bin";
|
|
}
|
|
|
|
int TrainPoolOpenRead(const string path)
|
|
{
|
|
//--- Both share flags, every time. One missing flag is what produced the zero-trade optimizer-cache
|
|
//--- corruption, and the symptom there was silence rather than an error.
|
|
return FileOpen(path, FILE_COMMON | FILE_READ | FILE_BIN | FILE_SHARE_READ | FILE_SHARE_WRITE);
|
|
}
|
|
|
|
//+------------------------------------------------------------------+
|
|
//| HEADER-ONLY CENSUS for capacity sizing (CTopology:: |
|
|
//| EstimatedInSampleBars), never for adoption - reads only the fixed- |
|
|
//| size STrainPoolHeader from each COMPATIBLE peer file (same |
|
|
//| MismatchReason() gate CTrainPoolReader::Adopt() uses), never the |
|
|
//| rows themselves. A topology decision that runs once at model |
|
|
//| creation should cost a handful of FileOpen calls, not the |
|
|
//| multi-hundred-MB read Adopt() does per era. |
|
|
//| |
|
|
//| Mirrors Adopt()'s own file scan (same glob, own-file exclusion, |
|
|
//| age filter, compatibility gate) so this predicts what Adopt() |
|
|
//| would actually adopt for the SAME (fingerprint, expectWidth) - an |
|
|
//| estimator that used a different rule than the real reader would |
|
|
//| just be a second, disagreeing opinion about the same number. |
|
|
//+------------------------------------------------------------------+
|
|
int TrainPoolEstimateAvailableRows(const string fingerprint, const string selfSymbol,
|
|
const int selfPeriod, const int expectWidth, int &peerFilesOut)
|
|
{
|
|
peerFilesOut = 0;
|
|
string selfPath = TrainPoolPath(selfSymbol, selfPeriod);
|
|
long minStamp = (long)TimeCurrent() - (long)TRAINPOOL_MAX_AGE_DAYS * 86400;
|
|
string name;
|
|
long find = FileFindFirst(TRAINPOOL_DIR + "\\*.bin", name, FILE_COMMON);
|
|
if(find == INVALID_HANDLE)
|
|
return 0;
|
|
int total = 0;
|
|
do
|
|
{
|
|
string full = TRAINPOOL_DIR + "\\" + name;
|
|
if(full == selfPath)
|
|
continue;
|
|
if((long)FileGetInteger(full, FILE_MODIFY_DATE, true) < minStamp)
|
|
continue;
|
|
int h = TrainPoolOpenRead(full);
|
|
if(h == INVALID_HANDLE)
|
|
continue;
|
|
STrainPoolHeader hdr;
|
|
hdr.Read(h);
|
|
FileClose(h);
|
|
if(hdr.Compatible(fingerprint, expectWidth) && hdr.rows > 0)
|
|
{
|
|
total += hdr.rows;
|
|
peerFilesOut++;
|
|
}
|
|
}
|
|
while(FileFindNext(find, name));
|
|
FileFindClose(find);
|
|
return MathMin(total, TRAINPOOL_MAX_ROWS);
|
|
}
|
|
|
|
//+------------------------------------------------------------------+
|
|
//| WRITER: buffers this chart's rows and publishes them once. |
|
|
//| |
|
|
//| Separate from the reader because they change for different |
|
|
//| reasons and run at different times - the writer during training, |
|
|
//| the reader before it. Sharing one class would mean every era's |
|
|
//| read carried the export buffers around with it. |
|
|
//+------------------------------------------------------------------+
|
|
class CTrainPoolWriter
|
|
{
|
|
private:
|
|
string m_fingerprint;
|
|
string m_path;
|
|
double m_rows[]; // flattened: row r at [r*m_width .. r*m_width + m_width)
|
|
int m_labels[];
|
|
long m_resolvedMs[];
|
|
int m_count;
|
|
int m_width;
|
|
long m_lastPublishSec;
|
|
|
|
public:
|
|
CTrainPoolWriter(void) : m_count(0), m_width(0), m_lastPublishSec(0) {}
|
|
|
|
int Count(void) const { return m_count; }
|
|
|
|
void Begin(const string fingerprint, const string symbol, const int period)
|
|
{
|
|
m_fingerprint = fingerprint;
|
|
m_path = TrainPoolPath(symbol, period);
|
|
m_count = 0;
|
|
m_width = 0;
|
|
}
|
|
|
|
bool Add(const CArrayDouble &features, const int label, const long resolvedMs);
|
|
bool Publish(void);
|
|
};
|
|
|
|
//+------------------------------------------------------------------+
|
|
//| Buffer one row. Called where TempData is already built, so this |
|
|
//| costs a copy and no recomputation. |
|
|
//+------------------------------------------------------------------+
|
|
bool CTrainPoolWriter::Add(const CArrayDouble &features, const int label, const long resolvedMs)
|
|
{
|
|
int w = features.Total();
|
|
if(w <= 0 || label < 0 || resolvedMs <= 0)
|
|
return false;
|
|
if(m_count == 0)
|
|
m_width = w;
|
|
//--- A width change mid-training means the feature layout moved under us. Refusing is the only safe
|
|
//--- answer: a short row in a flat array shifts every subsequent row's columns.
|
|
if(w != m_width)
|
|
return false;
|
|
int need = (m_count + 1) * m_width;
|
|
if(ArraySize(m_rows) < need)
|
|
ArrayResize(m_rows, need, 16384 * m_width);
|
|
for(int k = 0; k < m_width; k++)
|
|
m_rows[m_count * m_width + k] = features.At(k);
|
|
if(ArraySize(m_labels) <= m_count)
|
|
{
|
|
ArrayResize(m_labels, m_count + 1, 16384);
|
|
ArrayResize(m_resolvedMs, m_count + 1, 16384);
|
|
}
|
|
m_labels[m_count] = label;
|
|
m_resolvedMs[m_count] = resolvedMs;
|
|
m_count++;
|
|
return true;
|
|
}
|
|
|
|
//+------------------------------------------------------------------+
|
|
bool CTrainPoolWriter::Publish(void)
|
|
{
|
|
if(m_count <= 0 || m_width <= 0)
|
|
return false;
|
|
//--- Every era rebuilds the same rows from the same span; only the clock decides a republish is
|
|
//--- worth the write. The FIRST publish is never delayed - a peer waiting on this chart's corpus
|
|
//--- must not sit idle for the interval before the pool exists at all.
|
|
long nowSec = (long)TimeCurrent();
|
|
if(m_lastPublishSec > 0 && nowSec - m_lastPublishSec < TRAINPOOL_MIN_PUBLISH_SEC)
|
|
return true;
|
|
FolderCreate(TRAINPOOL_DIR, FILE_COMMON);
|
|
ResetLastError();
|
|
//--- Staged and renamed via the shared helper rather than hand-rolled here. A peer reading a
|
|
//--- half-written corpus would adopt a truncated final row and train on shifted columns, and the
|
|
//--- file would look perfectly valid while doing it.
|
|
string tmp;
|
|
int h = AtomicWriteBegin(m_path, FILE_COMMON, tmp);
|
|
if(h == INVALID_HANDLE)
|
|
{
|
|
Print(__FUNCTION__, ": cannot stage ", m_path, " err=", GetLastError());
|
|
return false;
|
|
}
|
|
STrainPoolHeader hdr;
|
|
hdr.version = TRAINPOOL_RECORD_VERSION;
|
|
hdr.width = m_width;
|
|
hdr.rows = m_count;
|
|
hdr.fingerprint = m_fingerprint;
|
|
hdr.Write(h);
|
|
for(int r = 0; r < m_count; r++)
|
|
{
|
|
FileWriteLong(h, m_resolvedMs[r]);
|
|
FileWriteInteger(h, m_labels[r], INT_VALUE);
|
|
for(int k = 0; k < m_width; k++)
|
|
FileWriteDouble(h, m_rows[r * m_width + k]);
|
|
}
|
|
bool ok = AtomicWriteEnd(h, m_path, tmp, FILE_COMMON, true, __FUNCTION__);
|
|
if(ok)
|
|
{
|
|
m_lastPublishSec = nowSec;
|
|
Print(__FUNCTION__, ": contributed ", m_count, " rows x ", m_width,
|
|
" features to the training pool as ", m_path);
|
|
}
|
|
return ok;
|
|
}
|
|
|
|
//+------------------------------------------------------------------+
|
|
//| READER: the peer rows this model may legitimately train on. |
|
|
//+------------------------------------------------------------------+
|
|
class CTrainPoolReader
|
|
{
|
|
private:
|
|
double m_rows[];
|
|
int m_labels[];
|
|
int m_rowCount;
|
|
int m_width;
|
|
int m_peerCount;
|
|
string m_peerNames;
|
|
//--- Adoption runs once per era and an era on a warm feature cache can be a fraction of a second,
|
|
//--- so an unconditional line would bury the journal (the same reason ReportTrainStall exists).
|
|
//--- The verdict is printed when it CHANGES, which is the only time it carries information.
|
|
string m_lastVerdict;
|
|
|
|
void Take(const int h, const int label);
|
|
void Announce(const string verdict)
|
|
{
|
|
if(verdict == m_lastVerdict)
|
|
return;
|
|
m_lastVerdict = verdict;
|
|
Print(__FUNCTION__, ": ", verdict);
|
|
}
|
|
|
|
public:
|
|
CTrainPoolReader(void) : m_rowCount(0), m_width(0), m_peerCount(0),
|
|
m_peerNames(""), m_lastVerdict("") {}
|
|
|
|
int RowCount(void) const { return m_rowCount; }
|
|
int PeerCount(void) const { return m_peerCount; }
|
|
string PeerNames(void) const { return m_peerNames; }
|
|
double At(const int r, const int k) const { return m_rows[r * m_width + k]; }
|
|
int LabelAt(const int r) const { return m_labels[r]; }
|
|
|
|
void Clear(void);
|
|
int Adopt(const string fingerprint, const string selfSymbol, const int selfPeriod,
|
|
const long trainCutoffMs, const int expectWidth);
|
|
};
|
|
|
|
//+------------------------------------------------------------------+
|
|
void CTrainPoolReader::Clear(void)
|
|
{
|
|
ArrayFree(m_rows);
|
|
ArrayFree(m_labels);
|
|
m_rowCount = 0;
|
|
m_peerCount = 0;
|
|
m_peerNames = "";
|
|
}
|
|
|
|
//+------------------------------------------------------------------+
|
|
void CTrainPoolReader::Take(const int h, const int label)
|
|
{
|
|
int need = (m_rowCount + 1) * m_width;
|
|
if(ArraySize(m_rows) < need)
|
|
ArrayResize(m_rows, need, 16384 * m_width);
|
|
if(ArraySize(m_labels) <= m_rowCount)
|
|
ArrayResize(m_labels, m_rowCount + 1, 16384);
|
|
for(int k = 0; k < m_width; k++)
|
|
m_rows[m_rowCount * m_width + k] = FileReadDouble(h);
|
|
m_labels[m_rowCount] = label;
|
|
m_rowCount++;
|
|
}
|
|
|
|
//+------------------------------------------------------------------+
|
|
//| -> rows adopted. Three gates, each with a failure mode already |
|
|
//| paid for here: |
|
|
//| |
|
|
//| FINGERPRINT+WIDTH STrainPoolHeader::Compatible - column k must |
|
|
//| mean the same thing in both files. |
|
|
//| AGE a decommissioned symbol stops voting. |
|
|
//| RESOLVED a peer row whose label resolved at or after |
|
|
//| this model's training cutoff is LOOKAHEAD - |
|
|
//| the same exit-time purge the research arms |
|
|
//| used. Purging on BAR INDEX would be wrong |
|
|
//| across instruments: every symbol has its own |
|
|
//| calendar of weekends, holidays and sessions. |
|
|
//+------------------------------------------------------------------+
|
|
int CTrainPoolReader::Adopt(const string fingerprint, const string selfSymbol, const int selfPeriod,
|
|
const long trainCutoffMs, const int expectWidth)
|
|
{
|
|
Clear();
|
|
m_width = expectWidth;
|
|
string selfPath = TrainPoolPath(selfSymbol, selfPeriod);
|
|
long minStamp = (long)TimeCurrent() - (long)TRAINPOOL_MAX_AGE_DAYS * 86400;
|
|
|
|
//--- Headers first, rows second. The stride needs the TOTAL before any row is adopted; adopting
|
|
//--- greedily and trimming afterwards would keep whichever peers were read first, which is
|
|
//--- alphabetical order, not a sample.
|
|
string paths[];
|
|
int counts[];
|
|
string rejected = "";
|
|
string name;
|
|
long find = FileFindFirst(TRAINPOOL_DIR + "\\*.bin", name, FILE_COMMON);
|
|
if(find == INVALID_HANDLE)
|
|
{
|
|
Announce("no peer files in the pool yet - this chart publishes into it, and adopts once a"
|
|
" SECOND chart with a matching model fingerprint has finished an era");
|
|
return 0;
|
|
}
|
|
do
|
|
{
|
|
string full = TRAINPOOL_DIR + "\\" + name;
|
|
if(full == selfPath)
|
|
continue;
|
|
if((long)FileGetInteger(full, FILE_MODIFY_DATE, true) < minStamp)
|
|
{
|
|
rejected += (rejected == "" ? "" : "; ") + name + ": older than " +
|
|
IntegerToString(TRAINPOOL_MAX_AGE_DAYS) + " days";
|
|
continue;
|
|
}
|
|
int h = TrainPoolOpenRead(full);
|
|
if(h == INVALID_HANDLE)
|
|
{
|
|
rejected += (rejected == "" ? "" : "; ") + name + ": cannot open, err=" +
|
|
IntegerToString(GetLastError());
|
|
continue;
|
|
}
|
|
STrainPoolHeader hdr;
|
|
hdr.Read(h);
|
|
FileClose(h);
|
|
string why = hdr.MismatchReason(fingerprint, expectWidth);
|
|
if(why != "")
|
|
{
|
|
rejected += (rejected == "" ? "" : "; ") + name + ": " + why;
|
|
continue;
|
|
}
|
|
int idx = ArraySize(paths);
|
|
ArrayResize(paths, idx + 1);
|
|
ArrayResize(counts, idx + 1);
|
|
paths[idx] = full;
|
|
counts[idx] = hdr.rows;
|
|
}
|
|
while(FileFindNext(find, name));
|
|
FileFindClose(find);
|
|
|
|
int files = ArraySize(paths);
|
|
if(files <= 0)
|
|
{
|
|
Announce(rejected == ""
|
|
? "no peer files in the pool yet - this chart publishes into it, and adopts once a"
|
|
" SECOND chart with a matching model fingerprint has finished an era"
|
|
: "EVERY peer file was REJECTED, so this chart is training alone. This model wants "
|
|
+ IntegerToString(expectWidth) + " features per row under fingerprint " + fingerprint
|
|
+ ". Rejected: " + rejected);
|
|
return 0;
|
|
}
|
|
int total = 0;
|
|
for(int i = 0; i < files; i++)
|
|
total += counts[i];
|
|
//--- Uniform stride across the WHOLE pool, so every peer thins in the same proportion and the cap
|
|
//--- cannot silently drop an instrument entirely.
|
|
int stride = (total > TRAINPOOL_MAX_ROWS) ? (total / TRAINPOOL_MAX_ROWS) + 1 : 1;
|
|
|
|
for(int i = 0; i < files; i++)
|
|
{
|
|
int h = TrainPoolOpenRead(paths[i]);
|
|
if(h == INVALID_HANDLE)
|
|
continue;
|
|
STrainPoolHeader hdr;
|
|
hdr.Read(h);
|
|
int before = m_rowCount;
|
|
for(int r = 0; r < hdr.rows && !FileIsEnding(h); r++)
|
|
{
|
|
long resolvedMs = FileReadLong(h);
|
|
int label = FileReadInteger(h, INT_VALUE);
|
|
if((r % stride) == 0 && resolvedMs < trainCutoffMs)
|
|
Take(h, label);
|
|
else
|
|
for(int k = 0; k < m_width; k++) // consume, to stay on the record boundary
|
|
FileReadDouble(h);
|
|
}
|
|
FileClose(h);
|
|
int got = m_rowCount - before;
|
|
if(got > 0)
|
|
{
|
|
m_peerCount++;
|
|
m_peerNames += (m_peerNames == "" ? "" : ", ") +
|
|
StringSubstr(paths[i], StringLen(TRAINPOOL_DIR) + 1) +
|
|
"(" + IntegerToString(got) + ")";
|
|
}
|
|
}
|
|
//--- One place reports pool state, adopted or not, so "it is off", "it is on but alone" and "it is
|
|
//--- pooling" are never three different silences. Rejections are appended even on a partial success:
|
|
//--- adopting two peers out of three still means one instrument is missing from the pool.
|
|
Announce(StringFormat("adopted %d rows from %d peer(s): %s%s", m_rowCount, m_peerCount, m_peerNames,
|
|
(rejected == "" ? "" : " | REJECTED: " + rejected)));
|
|
return m_rowCount;
|
|
}
|
|
|
|
#endif // WARRIOR_TRAINING_TRAININGPOOL_MQH
|