Warrior_EA/System/AtomicFile.mqh

160 lines
8.4 KiB
MQL5
Raw Permalink Normal View History

fix: make sidecar writes atomic; extract shared AtomicFile helper FileOpen(FILE_WRITE) truncates its target on open. CNet::Save already staged the .nnw through a temp file + rename for that reason, but the three sidecars written beside it did not: .stats ExpertSignalAIBase.mqh:5918 .arrows ExpertSignalAIBase.mqh:6224 .cfg ExpertSignalAIBase.mqh:7329 Two defects followed. 1. An interrupted write published a truncated sidecar. For .cfg that is the worst case: LoadAndCompareTopologyConfiguration() reads a short file as a mismatch, which discards the trained model and restarts from era 0. 2. Windows file sharing is a mutual contract - a writer opened with no FILE_SHARE_* blocks every concurrent open regardless of the reader's flags. All three read paths carry FILE_SHARE_READ|FILE_SHARE_WRITE specifically so a tester agent can read them while a live chart runs; an exclusive writer on the same path defeated that. Extracted CNet::Save's proven pattern into System\AtomicFile.mqh (AtomicWriteBegin/AtomicWriteEnd) and routed all four writers through it. This also encodes the FileMove gotcha once instead of per call site: the destination location comes from FILE_COMMON inside the 4th arg, NOT inherited from the source, and getting it wrong moves the file to the wrong sandbox silently. Also fixed while in these functions: - SaveTopologyConfiguration had 13 copy-pasted 6-line error blocks that each returned WITHOUT FileClose(handle), leaking the handle on every write failure. Collapsed to one ok-chain that closes exactly once. The on-disk field order and types are unchanged (asserted during the rewrite) so existing .cfg files still load. - SaveChartSignals documented that pruning runs only after a successful write ("a failed write above leaves both the file AND the chart untouched") but never checked any write result, so a partial write still deleted the chart objects. Results are checked now, making the existing comment true. Compiles 0 errors, 0 warnings. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 00:31:29 -04:00
//+------------------------------------------------------------------+
//| Warrior_EA |
//| AnimateDread |
//| |
//| Crash-safe file writes: temp file + atomic rename. |
fix: make sidecar writes atomic; extract shared AtomicFile helper FileOpen(FILE_WRITE) truncates its target on open. CNet::Save already staged the .nnw through a temp file + rename for that reason, but the three sidecars written beside it did not: .stats ExpertSignalAIBase.mqh:5918 .arrows ExpertSignalAIBase.mqh:6224 .cfg ExpertSignalAIBase.mqh:7329 Two defects followed. 1. An interrupted write published a truncated sidecar. For .cfg that is the worst case: LoadAndCompareTopologyConfiguration() reads a short file as a mismatch, which discards the trained model and restarts from era 0. 2. Windows file sharing is a mutual contract - a writer opened with no FILE_SHARE_* blocks every concurrent open regardless of the reader's flags. All three read paths carry FILE_SHARE_READ|FILE_SHARE_WRITE specifically so a tester agent can read them while a live chart runs; an exclusive writer on the same path defeated that. Extracted CNet::Save's proven pattern into System\AtomicFile.mqh (AtomicWriteBegin/AtomicWriteEnd) and routed all four writers through it. This also encodes the FileMove gotcha once instead of per call site: the destination location comes from FILE_COMMON inside the 4th arg, NOT inherited from the source, and getting it wrong moves the file to the wrong sandbox silently. Also fixed while in these functions: - SaveTopologyConfiguration had 13 copy-pasted 6-line error blocks that each returned WITHOUT FileClose(handle), leaking the handle on every write failure. Collapsed to one ok-chain that closes exactly once. The on-disk field order and types are unchanged (asserted during the rewrite) so existing .cfg files still load. - SaveChartSignals documented that pruning runs only after a successful write ("a failed write above leaves both the file AND the chart untouched") but never checked any write result, so a partial write still deleted the chart objects. Results are checked now, making the existing comment true. Compiles 0 errors, 0 warnings. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 00:31:29 -04:00
//+------------------------------------------------------------------+
#ifndef WARRIOR_ATOMIC_FILE_MQH
#define WARRIOR_ATOMIC_FILE_MQH
fix(pool): defer the atomic promotion to the timer instead of spinning on the tick REPLACES the in-line retry from ad4ae58, which was the wrong shape and did not work. Measured after deploying it: atomic rename ... failed (error 5004) after 4 attempts MQL5 exposes no FILE_SHARE_DELETE, so a rename CANNOT succeed while any reader holds the destination open - it is not a lock that waiting longer wins. The retry assumed a peer holds a pool file for "tens of ms"; USDCAD_16388.bin is 134 MB and a peer reading it holds the handle for SECONDS. The loop lost every time and bought nothing but 75ms of tick latency on the failure path. The content is already written and correct - only the SWAP is blocked. So try the rename once, and on failure remember the temp and promote it from OnTimer, where I/O belongs. Once the reader closes, a single FileMove lands it. That beats the old fallback of waiting for the next full publish, which rewrites all 134 MB and may be an era away. * pending list is bounded (8) and deduplicated - AtomicWriteBegin reuses one temp name per file, so a second failure for the same file must not take a second slot. A full list falls back to the previous next-publish behaviour. * a successful write FORGETS any queued promotion for that name, so a stale temp can never overwrite fresher content. * a vanished temp (a later publish succeeded outright) is dropped, not retried. * a landed promotion is LOGGED. Silence is what made me misread the last attempt as working when there had simply been no contention in the window. Compiled clean; NOT yet run - and note that verification needs a collision to occur, which happened ~27 times across a whole day. Absence of the message in any one window is not evidence either way. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 13:41:35 -04:00
//--- DEFERRED PROMOTION, and the in-line retry that preceded it was WRONG. A rename fails when a peer
//--- chart holds the destination open, and MQL5 has no FILE_SHARE_DELETE, so the rename simply cannot
//--- succeed for as long as that reader is open. The first attempt at this assumed a reader holds a
//--- pool file for "tens of ms" and spun 4x25ms; measured, `USDCAD_16388.bin` is 134 MB and a peer
//--- reading it holds the handle for SECONDS. The loop lost every time (`failed ... after 4 attempts`)
//--- and all it bought was 75ms of tick latency on the failure path.
//---
//--- So: try ONCE, and if the destination is busy, REMEMBER the temp and promote it from the timer.
//--- The content is already written and correct - only the swap is blocked - so once the reader closes,
//--- a single FileMove lands it. That beats waiting for the next full publish, which would re-write all
//--- 134 MB and might be an era away.
fix(io): stage atomic writes to a PER-CHART temp, not a shared one A DATA-INTEGRITY BUG, pre-existing, surfaced by the clearer failure message in 869cd1b putting two identical timestamps next to each other: 13:42:36.584 (EURUSD) AltRawSave: atomic rename raw_EIA_WPSR.csv.savetmp -> ... failed 13:42:36.584 (XTIUSD) AltRawSave: atomic rename raw_EIA_WPSR.csv.savetmp -> ... failed Same file, same millisecond, two charts, a third winning the race. That is not reader/writer contention - it is THREE WRITERS on one destination, and AtomicWriteBegin derived the staging name from the destination alone: tmpName = finalName + ".savetmp" So all three opened the SAME temp with FILE_WRITE and wrote it from offset 0 at once. The published file could be an interleaved mixture of two charts' output, and the atomic rename publishes that mixture faithfully - the swap guarantees a reader never sees a HALF-WRITTEN file, and does nothing about a HALF-CORRECT one. Alt-data is the exposed case: several charts fetch the same series and write the same Common file. Keying the temp on symbol+period makes staging private. The rename stays the only contended operation, and a rename IS atomic, so a loser now publishes nothing rather than half of itself. It also makes deferred promotion sound for the first time: the temp promoted later is THIS chart's complete content, never a fragment of someone else's. SharedFileCopy.mqh uses the same shape but its destination is agent/terminal-local and keyed by symbol+fingerprint, so charts cannot collide there. Left alone. Note the two bugs are independent and both fixes are real. Confirmed in situ at 13:45:45, on the reader/writer one: CTrainPoolWriter::Publish: atomic rename TrainPool\USDCAD_16388.bin failed (5004) Warrior: deferred promotion of TrainPool\USDCAD_16388.bin succeeded - the peer chart that held it has closed it, and the content written earlier is now live without rewriting the file. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 13:48:43 -04:00
//--- THE STAGING FILE MUST BE PRIVATE TO THE WRITER. It used to be `finalName + ".savetmp"`, derived
//--- from the DESTINATION alone - which is fine for a chart-local file and a genuine data-integrity
//--- bug for a shared one in Common\Files. Several charts fetch the same alt-data series and write
//--- the same destination, so they all opened ONE temp with FILE_WRITE and wrote it from offset 0 at
//--- once; whichever renamed first published whatever mixture of two charts' output the interleaving
//--- happened to leave. Observed 2026-08-26: EURUSD and XTIUSD both failed to rename
//--- `raw_EIA_WPSR.csv` in the SAME MILLISECOND, with a third writer winning the race.
//---
//--- Keying the temp on symbol+period makes staging private: the rename stays the only contended
//--- operation, and a rename is atomic, so a loser now publishes nothing rather than half of itself.
//--- It is also what makes deferred promotion sound - the temp promoted later is THIS chart's
//--- complete content, never a fragment of someone else's.
string AtomicTempName(const string finalName)
{
return StringFormat("%s.%s_%d.savetmp", finalName, _Symbol, (int)_Period);
}
fix(pool): defer the atomic promotion to the timer instead of spinning on the tick REPLACES the in-line retry from ad4ae58, which was the wrong shape and did not work. Measured after deploying it: atomic rename ... failed (error 5004) after 4 attempts MQL5 exposes no FILE_SHARE_DELETE, so a rename CANNOT succeed while any reader holds the destination open - it is not a lock that waiting longer wins. The retry assumed a peer holds a pool file for "tens of ms"; USDCAD_16388.bin is 134 MB and a peer reading it holds the handle for SECONDS. The loop lost every time and bought nothing but 75ms of tick latency on the failure path. The content is already written and correct - only the SWAP is blocked. So try the rename once, and on failure remember the temp and promote it from OnTimer, where I/O belongs. Once the reader closes, a single FileMove lands it. That beats the old fallback of waiting for the next full publish, which rewrites all 134 MB and may be an era away. * pending list is bounded (8) and deduplicated - AtomicWriteBegin reuses one temp name per file, so a second failure for the same file must not take a second slot. A full list falls back to the previous next-publish behaviour. * a successful write FORGETS any queued promotion for that name, so a stale temp can never overwrite fresher content. * a vanished temp (a later publish succeeded outright) is dropped, not retried. * a landed promotion is LOGGED. Silence is what made me misread the last attempt as working when there had simply been no contention in the window. Compiled clean; NOT yet run - and note that verification needs a collision to occur, which happened ~27 times across a whole day. Absence of the message in any one window is not evidence either way. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 13:41:35 -04:00
#define ATOMIC_PENDING_MAX 8
string g_atomicPendingFinal[ATOMIC_PENDING_MAX];
int g_atomicPendingFlag[ATOMIC_PENDING_MAX];
int g_atomicPendingCount = 0;
fix(io): stage atomic writes to a PER-CHART temp, not a shared one A DATA-INTEGRITY BUG, pre-existing, surfaced by the clearer failure message in 869cd1b putting two identical timestamps next to each other: 13:42:36.584 (EURUSD) AltRawSave: atomic rename raw_EIA_WPSR.csv.savetmp -> ... failed 13:42:36.584 (XTIUSD) AltRawSave: atomic rename raw_EIA_WPSR.csv.savetmp -> ... failed Same file, same millisecond, two charts, a third winning the race. That is not reader/writer contention - it is THREE WRITERS on one destination, and AtomicWriteBegin derived the staging name from the destination alone: tmpName = finalName + ".savetmp" So all three opened the SAME temp with FILE_WRITE and wrote it from offset 0 at once. The published file could be an interleaved mixture of two charts' output, and the atomic rename publishes that mixture faithfully - the swap guarantees a reader never sees a HALF-WRITTEN file, and does nothing about a HALF-CORRECT one. Alt-data is the exposed case: several charts fetch the same series and write the same Common file. Keying the temp on symbol+period makes staging private. The rename stays the only contended operation, and a rename IS atomic, so a loser now publishes nothing rather than half of itself. It also makes deferred promotion sound for the first time: the temp promoted later is THIS chart's complete content, never a fragment of someone else's. SharedFileCopy.mqh uses the same shape but its destination is agent/terminal-local and keyed by symbol+fingerprint, so charts cannot collide there. Left alone. Note the two bugs are independent and both fixes are real. Confirmed in situ at 13:45:45, on the reader/writer one: CTrainPoolWriter::Publish: atomic rename TrainPool\USDCAD_16388.bin failed (5004) Warrior: deferred promotion of TrainPool\USDCAD_16388.bin succeeded - the peer chart that held it has closed it, and the content written earlier is now live without rewriting the file. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 13:48:43 -04:00
//--- Bounded and deduplicated: there is only ever one temp per (final name, THIS chart), because
//--- AtomicTempName() reuses it, so a second failure for the same file must not take a second slot.
fix(pool): defer the atomic promotion to the timer instead of spinning on the tick REPLACES the in-line retry from ad4ae58, which was the wrong shape and did not work. Measured after deploying it: atomic rename ... failed (error 5004) after 4 attempts MQL5 exposes no FILE_SHARE_DELETE, so a rename CANNOT succeed while any reader holds the destination open - it is not a lock that waiting longer wins. The retry assumed a peer holds a pool file for "tens of ms"; USDCAD_16388.bin is 134 MB and a peer reading it holds the handle for SECONDS. The loop lost every time and bought nothing but 75ms of tick latency on the failure path. The content is already written and correct - only the SWAP is blocked. So try the rename once, and on failure remember the temp and promote it from OnTimer, where I/O belongs. Once the reader closes, a single FileMove lands it. That beats the old fallback of waiting for the next full publish, which rewrites all 134 MB and may be an era away. * pending list is bounded (8) and deduplicated - AtomicWriteBegin reuses one temp name per file, so a second failure for the same file must not take a second slot. A full list falls back to the previous next-publish behaviour. * a successful write FORGETS any queued promotion for that name, so a stale temp can never overwrite fresher content. * a vanished temp (a later publish succeeded outright) is dropped, not retried. * a landed promotion is LOGGED. Silence is what made me misread the last attempt as working when there had simply been no contention in the window. Compiled clean; NOT yet run - and note that verification needs a collision to occur, which happened ~27 times across a whole day. Absence of the message in any one window is not evidence either way. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 13:41:35 -04:00
void AtomicRememberPending(const string finalName, const int commonFlag)
{
for(int i = 0; i < g_atomicPendingCount; i++)
if(g_atomicPendingFinal[i] == finalName && g_atomicPendingFlag[i] == commonFlag)
return;
if(g_atomicPendingCount >= ATOMIC_PENDING_MAX)
return; // full: this one falls back to the next publish, as it always did
g_atomicPendingFinal[g_atomicPendingCount] = finalName;
g_atomicPendingFlag[g_atomicPendingCount] = commonFlag;
g_atomicPendingCount++;
}
void AtomicForgetPending(const string finalName, const int commonFlag)
{
for(int i = 0; i < g_atomicPendingCount; i++)
if(g_atomicPendingFinal[i] == finalName && g_atomicPendingFlag[i] == commonFlag)
{
g_atomicPendingFinal[i] = g_atomicPendingFinal[g_atomicPendingCount - 1];
g_atomicPendingFlag[i] = g_atomicPendingFlag[g_atomicPendingCount - 1];
g_atomicPendingCount--;
return;
}
}
//--- Called from the TIMER, not the tick: this is I/O, and the whole point is to stop paying for it
//--- inside a quote. Returns how many it managed to land. Cheap when there is nothing pending, which
//--- is the overwhelmingly common case.
int AtomicPromotePending(void)
{
int promoted = 0;
for(int i = g_atomicPendingCount - 1; i >= 0; i--)
{
string finalName = g_atomicPendingFinal[i];
fix(io): stage atomic writes to a PER-CHART temp, not a shared one A DATA-INTEGRITY BUG, pre-existing, surfaced by the clearer failure message in 869cd1b putting two identical timestamps next to each other: 13:42:36.584 (EURUSD) AltRawSave: atomic rename raw_EIA_WPSR.csv.savetmp -> ... failed 13:42:36.584 (XTIUSD) AltRawSave: atomic rename raw_EIA_WPSR.csv.savetmp -> ... failed Same file, same millisecond, two charts, a third winning the race. That is not reader/writer contention - it is THREE WRITERS on one destination, and AtomicWriteBegin derived the staging name from the destination alone: tmpName = finalName + ".savetmp" So all three opened the SAME temp with FILE_WRITE and wrote it from offset 0 at once. The published file could be an interleaved mixture of two charts' output, and the atomic rename publishes that mixture faithfully - the swap guarantees a reader never sees a HALF-WRITTEN file, and does nothing about a HALF-CORRECT one. Alt-data is the exposed case: several charts fetch the same series and write the same Common file. Keying the temp on symbol+period makes staging private. The rename stays the only contended operation, and a rename IS atomic, so a loser now publishes nothing rather than half of itself. It also makes deferred promotion sound for the first time: the temp promoted later is THIS chart's complete content, never a fragment of someone else's. SharedFileCopy.mqh uses the same shape but its destination is agent/terminal-local and keyed by symbol+fingerprint, so charts cannot collide there. Left alone. Note the two bugs are independent and both fixes are real. Confirmed in situ at 13:45:45, on the reader/writer one: CTrainPoolWriter::Publish: atomic rename TrainPool\USDCAD_16388.bin failed (5004) Warrior: deferred promotion of TrainPool\USDCAD_16388.bin succeeded - the peer chart that held it has closed it, and the content written earlier is now live without rewriting the file. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 13:48:43 -04:00
string tmpName = AtomicTempName(finalName);
fix(pool): defer the atomic promotion to the timer instead of spinning on the tick REPLACES the in-line retry from ad4ae58, which was the wrong shape and did not work. Measured after deploying it: atomic rename ... failed (error 5004) after 4 attempts MQL5 exposes no FILE_SHARE_DELETE, so a rename CANNOT succeed while any reader holds the destination open - it is not a lock that waiting longer wins. The retry assumed a peer holds a pool file for "tens of ms"; USDCAD_16388.bin is 134 MB and a peer reading it holds the handle for SECONDS. The loop lost every time and bought nothing but 75ms of tick latency on the failure path. The content is already written and correct - only the SWAP is blocked. So try the rename once, and on failure remember the temp and promote it from OnTimer, where I/O belongs. Once the reader closes, a single FileMove lands it. That beats the old fallback of waiting for the next full publish, which rewrites all 134 MB and may be an era away. * pending list is bounded (8) and deduplicated - AtomicWriteBegin reuses one temp name per file, so a second failure for the same file must not take a second slot. A full list falls back to the previous next-publish behaviour. * a successful write FORGETS any queued promotion for that name, so a stale temp can never overwrite fresher content. * a vanished temp (a later publish succeeded outright) is dropped, not retried. * a landed promotion is LOGGED. Silence is what made me misread the last attempt as working when there had simply been no contention in the window. Compiled clean; NOT yet run - and note that verification needs a collision to occur, which happened ~27 times across a whole day. Absence of the message in any one window is not evidence either way. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 13:41:35 -04:00
int flag = g_atomicPendingFlag[i];
//--- Gone means someone else already resolved it (a later publish succeeded outright). Drop it.
if(!FileIsExist(tmpName, flag))
{
AtomicForgetPending(finalName, flag);
continue;
}
ResetLastError();
if(!FileMove(tmpName, flag, finalName, flag | FILE_REWRITE))
continue; // still held; try again on the next timer
Print("Warrior: deferred promotion of ", finalName, " succeeded - the peer chart that held it"
" has closed it, and the content written earlier is now live without rewriting the file.");
AtomicForgetPending(finalName, flag);
promoted++;
}
return promoted;
}
feat(vote): exit-on-reversal boolean, pin the threshold, retry the atomic rename THE EXIT KNOB. Exit_On_Reversal_Vote (default false) replaces the deleted Signal_ThresholdClose with one boolean: false pins the close threshold to an arithmetically unreachable 101, true pins it to the SAME threshold the entry uses - the seed at first, then the derived value, republished together whenever it moves. A second threshold was always redundant; "the bot now says the other way" is one question. It also arms CExpertSignalCustom::m_holdToBarrier, which was DEAD CODE: HoldToBarrier(bool) had no caller anywhere in the build, so the flag had been permanently false and the disabled close threshold was carrying the whole hold-to-barrier policy alone. Both halves now move together. Default stays false because the reason is statistical: the gate certifies P(label agrees | vote fired) against a label that runs to the barrier, so an early close trades something never measured. Turning it on is a different strategy, not a tightening of this one. THE PIN. The live threshold now moves only when an era's weights become the checkpoint, and freezes once g_ensDeployApproved. Every era still derives its own rung - that is how the best one is found - but the rung that TRADES belongs to the checkpoint, exactly as the weights do. Two reasons, one measured and one structural: the per-era rung moves on 6-34% of steps (the live run flapped SP500 15 -> 10 -> 15 within a minute of starting), and without the pin a later era's rung could end up applied to an earlier era's deployed model. A ladder restart releases the pin, since clearing the checkpoint clears what it pinned. The era line now prints the rung its own numbers came from, so it stays honest when that differs from the pinned one. THE ATOMIC RENAME retried zero times. Six charts share the TrainPool and AltData directories, so a publish regularly lands while a peer chart holds the destination open and FileMove returns 5004 - 27 times in one day on the live fleet. Nothing was lost (the temp keeps the new content, the old file stays intact) but the row did not update until the next publish. Now four attempts at 25ms, on the FAILURE PATH ONLY - a successful rename never sleeps - and skipped in the tester, where the contention cannot happen and Sleep would distort a pass. A rescued retry is logged, so worsening contention is visible. Retrain-neutral. Compiled clean; NOT yet run. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 09:53:56 -04:00
fix: make sidecar writes atomic; extract shared AtomicFile helper FileOpen(FILE_WRITE) truncates its target on open. CNet::Save already staged the .nnw through a temp file + rename for that reason, but the three sidecars written beside it did not: .stats ExpertSignalAIBase.mqh:5918 .arrows ExpertSignalAIBase.mqh:6224 .cfg ExpertSignalAIBase.mqh:7329 Two defects followed. 1. An interrupted write published a truncated sidecar. For .cfg that is the worst case: LoadAndCompareTopologyConfiguration() reads a short file as a mismatch, which discards the trained model and restarts from era 0. 2. Windows file sharing is a mutual contract - a writer opened with no FILE_SHARE_* blocks every concurrent open regardless of the reader's flags. All three read paths carry FILE_SHARE_READ|FILE_SHARE_WRITE specifically so a tester agent can read them while a live chart runs; an exclusive writer on the same path defeated that. Extracted CNet::Save's proven pattern into System\AtomicFile.mqh (AtomicWriteBegin/AtomicWriteEnd) and routed all four writers through it. This also encodes the FileMove gotcha once instead of per call site: the destination location comes from FILE_COMMON inside the 4th arg, NOT inherited from the source, and getting it wrong moves the file to the wrong sandbox silently. Also fixed while in these functions: - SaveTopologyConfiguration had 13 copy-pasted 6-line error blocks that each returned WITHOUT FileClose(handle), leaking the handle on every write failure. Collapsed to one ok-chain that closes exactly once. The on-disk field order and types are unchanged (asserted during the rewrite) so existing .cfg files still load. - SaveChartSignals documented that pruning runs only after a successful write ("a failed write above leaves both the file AND the chart untouched") but never checked any write result, so a partial write still deleted the chart objects. Results are checked now, making the existing comment true. Compiles 0 errors, 0 warnings. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 00:31:29 -04:00
//+------------------------------------------------------------------+
//| Open a temp file to stage an atomic write. modeFlags defaults to |
//| FILE_BIN (every original caller wrote binary payloads); a text |
//| writer passes FILE_TXT|FILE_ANSI so FileWriteString still emits |
//| plain lines instead of length-prefixed binary records. |
fix: make sidecar writes atomic; extract shared AtomicFile helper FileOpen(FILE_WRITE) truncates its target on open. CNet::Save already staged the .nnw through a temp file + rename for that reason, but the three sidecars written beside it did not: .stats ExpertSignalAIBase.mqh:5918 .arrows ExpertSignalAIBase.mqh:6224 .cfg ExpertSignalAIBase.mqh:7329 Two defects followed. 1. An interrupted write published a truncated sidecar. For .cfg that is the worst case: LoadAndCompareTopologyConfiguration() reads a short file as a mismatch, which discards the trained model and restarts from era 0. 2. Windows file sharing is a mutual contract - a writer opened with no FILE_SHARE_* blocks every concurrent open regardless of the reader's flags. All three read paths carry FILE_SHARE_READ|FILE_SHARE_WRITE specifically so a tester agent can read them while a live chart runs; an exclusive writer on the same path defeated that. Extracted CNet::Save's proven pattern into System\AtomicFile.mqh (AtomicWriteBegin/AtomicWriteEnd) and routed all four writers through it. This also encodes the FileMove gotcha once instead of per call site: the destination location comes from FILE_COMMON inside the 4th arg, NOT inherited from the source, and getting it wrong moves the file to the wrong sandbox silently. Also fixed while in these functions: - SaveTopologyConfiguration had 13 copy-pasted 6-line error blocks that each returned WITHOUT FileClose(handle), leaking the handle on every write failure. Collapsed to one ok-chain that closes exactly once. The on-disk field order and types are unchanged (asserted during the rewrite) so existing .cfg files still load. - SaveChartSignals documented that pruning runs only after a successful write ("a failed write above leaves both the file AND the chart untouched") but never checked any write result, so a partial write still deleted the chart objects. Results are checked now, making the existing comment true. Compiles 0 errors, 0 warnings. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 00:31:29 -04:00
//+------------------------------------------------------------------+
int AtomicWriteBegin(const string finalName, const int commonFlag, string &tmpName, const int modeFlags = FILE_BIN)
fix: make sidecar writes atomic; extract shared AtomicFile helper FileOpen(FILE_WRITE) truncates its target on open. CNet::Save already staged the .nnw through a temp file + rename for that reason, but the three sidecars written beside it did not: .stats ExpertSignalAIBase.mqh:5918 .arrows ExpertSignalAIBase.mqh:6224 .cfg ExpertSignalAIBase.mqh:7329 Two defects followed. 1. An interrupted write published a truncated sidecar. For .cfg that is the worst case: LoadAndCompareTopologyConfiguration() reads a short file as a mismatch, which discards the trained model and restarts from era 0. 2. Windows file sharing is a mutual contract - a writer opened with no FILE_SHARE_* blocks every concurrent open regardless of the reader's flags. All three read paths carry FILE_SHARE_READ|FILE_SHARE_WRITE specifically so a tester agent can read them while a live chart runs; an exclusive writer on the same path defeated that. Extracted CNet::Save's proven pattern into System\AtomicFile.mqh (AtomicWriteBegin/AtomicWriteEnd) and routed all four writers through it. This also encodes the FileMove gotcha once instead of per call site: the destination location comes from FILE_COMMON inside the 4th arg, NOT inherited from the source, and getting it wrong moves the file to the wrong sandbox silently. Also fixed while in these functions: - SaveTopologyConfiguration had 13 copy-pasted 6-line error blocks that each returned WITHOUT FileClose(handle), leaking the handle on every write failure. Collapsed to one ok-chain that closes exactly once. The on-disk field order and types are unchanged (asserted during the rewrite) so existing .cfg files still load. - SaveChartSignals documented that pruning runs only after a successful write ("a failed write above leaves both the file AND the chart untouched") but never checked any write result, so a partial write still deleted the chart objects. Results are checked now, making the existing comment true. Compiles 0 errors, 0 warnings. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 00:31:29 -04:00
{
fix(io): stage atomic writes to a PER-CHART temp, not a shared one A DATA-INTEGRITY BUG, pre-existing, surfaced by the clearer failure message in 869cd1b putting two identical timestamps next to each other: 13:42:36.584 (EURUSD) AltRawSave: atomic rename raw_EIA_WPSR.csv.savetmp -> ... failed 13:42:36.584 (XTIUSD) AltRawSave: atomic rename raw_EIA_WPSR.csv.savetmp -> ... failed Same file, same millisecond, two charts, a third winning the race. That is not reader/writer contention - it is THREE WRITERS on one destination, and AtomicWriteBegin derived the staging name from the destination alone: tmpName = finalName + ".savetmp" So all three opened the SAME temp with FILE_WRITE and wrote it from offset 0 at once. The published file could be an interleaved mixture of two charts' output, and the atomic rename publishes that mixture faithfully - the swap guarantees a reader never sees a HALF-WRITTEN file, and does nothing about a HALF-CORRECT one. Alt-data is the exposed case: several charts fetch the same series and write the same Common file. Keying the temp on symbol+period makes staging private. The rename stays the only contended operation, and a rename IS atomic, so a loser now publishes nothing rather than half of itself. It also makes deferred promotion sound for the first time: the temp promoted later is THIS chart's complete content, never a fragment of someone else's. SharedFileCopy.mqh uses the same shape but its destination is agent/terminal-local and keyed by symbol+fingerprint, so charts cannot collide there. Left alone. Note the two bugs are independent and both fixes are real. Confirmed in situ at 13:45:45, on the reader/writer one: CTrainPoolWriter::Publish: atomic rename TrainPool\USDCAD_16388.bin failed (5004) Warrior: deferred promotion of TrainPool\USDCAD_16388.bin succeeded - the peer chart that held it has closed it, and the content written earlier is now live without rewriting the file. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 13:48:43 -04:00
tmpName = AtomicTempName(finalName);
fix: make sidecar writes atomic; extract shared AtomicFile helper FileOpen(FILE_WRITE) truncates its target on open. CNet::Save already staged the .nnw through a temp file + rename for that reason, but the three sidecars written beside it did not: .stats ExpertSignalAIBase.mqh:5918 .arrows ExpertSignalAIBase.mqh:6224 .cfg ExpertSignalAIBase.mqh:7329 Two defects followed. 1. An interrupted write published a truncated sidecar. For .cfg that is the worst case: LoadAndCompareTopologyConfiguration() reads a short file as a mismatch, which discards the trained model and restarts from era 0. 2. Windows file sharing is a mutual contract - a writer opened with no FILE_SHARE_* blocks every concurrent open regardless of the reader's flags. All three read paths carry FILE_SHARE_READ|FILE_SHARE_WRITE specifically so a tester agent can read them while a live chart runs; an exclusive writer on the same path defeated that. Extracted CNet::Save's proven pattern into System\AtomicFile.mqh (AtomicWriteBegin/AtomicWriteEnd) and routed all four writers through it. This also encodes the FileMove gotcha once instead of per call site: the destination location comes from FILE_COMMON inside the 4th arg, NOT inherited from the source, and getting it wrong moves the file to the wrong sandbox silently. Also fixed while in these functions: - SaveTopologyConfiguration had 13 copy-pasted 6-line error blocks that each returned WITHOUT FileClose(handle), leaking the handle on every write failure. Collapsed to one ok-chain that closes exactly once. The on-disk field order and types are unchanged (asserted during the rewrite) so existing .cfg files still load. - SaveChartSignals documented that pruning runs only after a successful write ("a failed write above leaves both the file AND the chart untouched") but never checked any write result, so a partial write still deleted the chart objects. Results are checked now, making the existing comment true. Compiles 0 errors, 0 warnings. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 00:31:29 -04:00
ResetLastError();
return FileOpen(tmpName, commonFlag | modeFlags | FILE_WRITE | FILE_SHARE_READ | FILE_SHARE_WRITE);
fix: make sidecar writes atomic; extract shared AtomicFile helper FileOpen(FILE_WRITE) truncates its target on open. CNet::Save already staged the .nnw through a temp file + rename for that reason, but the three sidecars written beside it did not: .stats ExpertSignalAIBase.mqh:5918 .arrows ExpertSignalAIBase.mqh:6224 .cfg ExpertSignalAIBase.mqh:7329 Two defects followed. 1. An interrupted write published a truncated sidecar. For .cfg that is the worst case: LoadAndCompareTopologyConfiguration() reads a short file as a mismatch, which discards the trained model and restarts from era 0. 2. Windows file sharing is a mutual contract - a writer opened with no FILE_SHARE_* blocks every concurrent open regardless of the reader's flags. All three read paths carry FILE_SHARE_READ|FILE_SHARE_WRITE specifically so a tester agent can read them while a live chart runs; an exclusive writer on the same path defeated that. Extracted CNet::Save's proven pattern into System\AtomicFile.mqh (AtomicWriteBegin/AtomicWriteEnd) and routed all four writers through it. This also encodes the FileMove gotcha once instead of per call site: the destination location comes from FILE_COMMON inside the 4th arg, NOT inherited from the source, and getting it wrong moves the file to the wrong sandbox silently. Also fixed while in these functions: - SaveTopologyConfiguration had 13 copy-pasted 6-line error blocks that each returned WITHOUT FileClose(handle), leaking the handle on every write failure. Collapsed to one ok-chain that closes exactly once. The on-disk field order and types are unchanged (asserted during the rewrite) so existing .cfg files still load. - SaveChartSignals documented that pruning runs only after a successful write ("a failed write above leaves both the file AND the chart untouched") but never checked any write result, so a partial write still deleted the chart objects. Results are checked now, making the existing comment true. Compiles 0 errors, 0 warnings. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 00:31:29 -04:00
}
//+------------------------------------------------------------------+
//| Close the staged write and either publish it or discard it. |
fix: make sidecar writes atomic; extract shared AtomicFile helper FileOpen(FILE_WRITE) truncates its target on open. CNet::Save already staged the .nnw through a temp file + rename for that reason, but the three sidecars written beside it did not: .stats ExpertSignalAIBase.mqh:5918 .arrows ExpertSignalAIBase.mqh:6224 .cfg ExpertSignalAIBase.mqh:7329 Two defects followed. 1. An interrupted write published a truncated sidecar. For .cfg that is the worst case: LoadAndCompareTopologyConfiguration() reads a short file as a mismatch, which discards the trained model and restarts from era 0. 2. Windows file sharing is a mutual contract - a writer opened with no FILE_SHARE_* blocks every concurrent open regardless of the reader's flags. All three read paths carry FILE_SHARE_READ|FILE_SHARE_WRITE specifically so a tester agent can read them while a live chart runs; an exclusive writer on the same path defeated that. Extracted CNet::Save's proven pattern into System\AtomicFile.mqh (AtomicWriteBegin/AtomicWriteEnd) and routed all four writers through it. This also encodes the FileMove gotcha once instead of per call site: the destination location comes from FILE_COMMON inside the 4th arg, NOT inherited from the source, and getting it wrong moves the file to the wrong sandbox silently. Also fixed while in these functions: - SaveTopologyConfiguration had 13 copy-pasted 6-line error blocks that each returned WITHOUT FileClose(handle), leaking the handle on every write failure. Collapsed to one ok-chain that closes exactly once. The on-disk field order and types are unchanged (asserted during the rewrite) so existing .cfg files still load. - SaveChartSignals documented that pruning runs only after a successful write ("a failed write above leaves both the file AND the chart untouched") but never checked any write result, so a partial write still deleted the chart objects. Results are checked now, making the existing comment true. Compiles 0 errors, 0 warnings. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 00:31:29 -04:00
//+------------------------------------------------------------------+
bool AtomicWriteEnd(const int handle, const string finalName, const string tmpName,
const int commonFlag, const bool ok, const string context)
{
if(handle != INVALID_HANDLE)
{
FileFlush(handle);
FileClose(handle);
}
//--- A failed/partial write must NEVER touch the real file - discard the temp, keep the last good copy.
if(!ok)
{
FileDelete(tmpName, commonFlag);
Print(context, ": write to ", tmpName, " failed - kept the existing ", finalName,
" intact (no atomic swap performed)");
return false;
}
//--- Atomic swap: FILE_REWRITE lets FileMove replace an existing destination in one operation.
feat(vote): exit-on-reversal boolean, pin the threshold, retry the atomic rename THE EXIT KNOB. Exit_On_Reversal_Vote (default false) replaces the deleted Signal_ThresholdClose with one boolean: false pins the close threshold to an arithmetically unreachable 101, true pins it to the SAME threshold the entry uses - the seed at first, then the derived value, republished together whenever it moves. A second threshold was always redundant; "the bot now says the other way" is one question. It also arms CExpertSignalCustom::m_holdToBarrier, which was DEAD CODE: HoldToBarrier(bool) had no caller anywhere in the build, so the flag had been permanently false and the disabled close threshold was carrying the whole hold-to-barrier policy alone. Both halves now move together. Default stays false because the reason is statistical: the gate certifies P(label agrees | vote fired) against a label that runs to the barrier, so an early close trades something never measured. Turning it on is a different strategy, not a tightening of this one. THE PIN. The live threshold now moves only when an era's weights become the checkpoint, and freezes once g_ensDeployApproved. Every era still derives its own rung - that is how the best one is found - but the rung that TRADES belongs to the checkpoint, exactly as the weights do. Two reasons, one measured and one structural: the per-era rung moves on 6-34% of steps (the live run flapped SP500 15 -> 10 -> 15 within a minute of starting), and without the pin a later era's rung could end up applied to an earlier era's deployed model. A ladder restart releases the pin, since clearing the checkpoint clears what it pinned. The era line now prints the rung its own numbers came from, so it stays honest when that differs from the pinned one. THE ATOMIC RENAME retried zero times. Six charts share the TrainPool and AltData directories, so a publish regularly lands while a peer chart holds the destination open and FileMove returns 5004 - 27 times in one day on the live fleet. Nothing was lost (the temp keeps the new content, the old file stays intact) but the row did not update until the next publish. Now four attempts at 25ms, on the FAILURE PATH ONLY - a successful rename never sleeps - and skipped in the tester, where the contention cannot happen and Sleep would distort a pass. A rescued retry is logged, so worsening contention is visible. Retrain-neutral. Compiled clean; NOT yet run. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 09:53:56 -04:00
//---
//--- RETRIED, BECAUSE THE COMMON FAILURE IS TRANSIENT AND NOT OURS. Six charts share the TrainPool
//--- and AltData directories, so a publish regularly lands while a PEER CHART has the destination
//--- open for reading, and the rename comes back 5004 (which MQL5 uses for both "no such file" and
//--- "locked by another process"). Measured 27 times in one day on the live fleet before this loop
//--- existed. Nothing was lost - the temp keeps the new content and the old file stays intact - but
//--- the pool row simply did not update until the next publish, which on a chronically busy
//--- directory could be a long time.
//---
//--- The sleep is on the FAILURE PATH ONLY: a successful rename does not sleep at all, so the common
//--- case costs one extra comparison. Skipped in the tester (where Sleep distorts a pass and the
//--- contention cannot happen - one process, no peers) and when the program is stopping, where
//--- Sleep returns immediately anyway and the remaining attempts should just be spent at once.
fix(pool): defer the atomic promotion to the timer instead of spinning on the tick REPLACES the in-line retry from ad4ae58, which was the wrong shape and did not work. Measured after deploying it: atomic rename ... failed (error 5004) after 4 attempts MQL5 exposes no FILE_SHARE_DELETE, so a rename CANNOT succeed while any reader holds the destination open - it is not a lock that waiting longer wins. The retry assumed a peer holds a pool file for "tens of ms"; USDCAD_16388.bin is 134 MB and a peer reading it holds the handle for SECONDS. The loop lost every time and bought nothing but 75ms of tick latency on the failure path. The content is already written and correct - only the SWAP is blocked. So try the rename once, and on failure remember the temp and promote it from OnTimer, where I/O belongs. Once the reader closes, a single FileMove lands it. That beats the old fallback of waiting for the next full publish, which rewrites all 134 MB and may be an era away. * pending list is bounded (8) and deduplicated - AtomicWriteBegin reuses one temp name per file, so a second failure for the same file must not take a second slot. A full list falls back to the previous next-publish behaviour. * a successful write FORGETS any queued promotion for that name, so a stale temp can never overwrite fresher content. * a vanished temp (a later publish succeeded outright) is dropped, not retried. * a landed promotion is LOGGED. Silence is what made me misread the last attempt as working when there had simply been no contention in the window. Compiled clean; NOT yet run - and note that verification needs a collision to occur, which happened ~27 times across a whole day. Absence of the message in any one window is not evidence either way. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 13:41:35 -04:00
ResetLastError();
if(FileMove(tmpName, commonFlag, finalName, commonFlag | FILE_REWRITE))
fix: make sidecar writes atomic; extract shared AtomicFile helper FileOpen(FILE_WRITE) truncates its target on open. CNet::Save already staged the .nnw through a temp file + rename for that reason, but the three sidecars written beside it did not: .stats ExpertSignalAIBase.mqh:5918 .arrows ExpertSignalAIBase.mqh:6224 .cfg ExpertSignalAIBase.mqh:7329 Two defects followed. 1. An interrupted write published a truncated sidecar. For .cfg that is the worst case: LoadAndCompareTopologyConfiguration() reads a short file as a mismatch, which discards the trained model and restarts from era 0. 2. Windows file sharing is a mutual contract - a writer opened with no FILE_SHARE_* blocks every concurrent open regardless of the reader's flags. All three read paths carry FILE_SHARE_READ|FILE_SHARE_WRITE specifically so a tester agent can read them while a live chart runs; an exclusive writer on the same path defeated that. Extracted CNet::Save's proven pattern into System\AtomicFile.mqh (AtomicWriteBegin/AtomicWriteEnd) and routed all four writers through it. This also encodes the FileMove gotcha once instead of per call site: the destination location comes from FILE_COMMON inside the 4th arg, NOT inherited from the source, and getting it wrong moves the file to the wrong sandbox silently. Also fixed while in these functions: - SaveTopologyConfiguration had 13 copy-pasted 6-line error blocks that each returned WITHOUT FileClose(handle), leaking the handle on every write failure. Collapsed to one ok-chain that closes exactly once. The on-disk field order and types are unchanged (asserted during the rewrite) so existing .cfg files still load. - SaveChartSignals documented that pruning runs only after a successful write ("a failed write above leaves both the file AND the chart untouched") but never checked any write result, so a partial write still deleted the chart objects. Results are checked now, making the existing comment true. Compiles 0 errors, 0 warnings. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 00:31:29 -04:00
{
fix(pool): defer the atomic promotion to the timer instead of spinning on the tick REPLACES the in-line retry from ad4ae58, which was the wrong shape and did not work. Measured after deploying it: atomic rename ... failed (error 5004) after 4 attempts MQL5 exposes no FILE_SHARE_DELETE, so a rename CANNOT succeed while any reader holds the destination open - it is not a lock that waiting longer wins. The retry assumed a peer holds a pool file for "tens of ms"; USDCAD_16388.bin is 134 MB and a peer reading it holds the handle for SECONDS. The loop lost every time and bought nothing but 75ms of tick latency on the failure path. The content is already written and correct - only the SWAP is blocked. So try the rename once, and on failure remember the temp and promote it from OnTimer, where I/O belongs. Once the reader closes, a single FileMove lands it. That beats the old fallback of waiting for the next full publish, which rewrites all 134 MB and may be an era away. * pending list is bounded (8) and deduplicated - AtomicWriteBegin reuses one temp name per file, so a second failure for the same file must not take a second slot. A full list falls back to the previous next-publish behaviour. * a successful write FORGETS any queued promotion for that name, so a stale temp can never overwrite fresher content. * a vanished temp (a later publish succeeded outright) is dropped, not retried. * a landed promotion is LOGGED. Silence is what made me misread the last attempt as working when there had simply been no contention in the window. Compiled clean; NOT yet run - and note that verification needs a collision to occur, which happened ~27 times across a whole day. Absence of the message in any one window is not evidence either way. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 13:41:35 -04:00
//--- A fresh write supersedes anything queued for this name.
AtomicForgetPending(finalName, commonFlag);
return true;
fix: make sidecar writes atomic; extract shared AtomicFile helper FileOpen(FILE_WRITE) truncates its target on open. CNet::Save already staged the .nnw through a temp file + rename for that reason, but the three sidecars written beside it did not: .stats ExpertSignalAIBase.mqh:5918 .arrows ExpertSignalAIBase.mqh:6224 .cfg ExpertSignalAIBase.mqh:7329 Two defects followed. 1. An interrupted write published a truncated sidecar. For .cfg that is the worst case: LoadAndCompareTopologyConfiguration() reads a short file as a mismatch, which discards the trained model and restarts from era 0. 2. Windows file sharing is a mutual contract - a writer opened with no FILE_SHARE_* blocks every concurrent open regardless of the reader's flags. All three read paths carry FILE_SHARE_READ|FILE_SHARE_WRITE specifically so a tester agent can read them while a live chart runs; an exclusive writer on the same path defeated that. Extracted CNet::Save's proven pattern into System\AtomicFile.mqh (AtomicWriteBegin/AtomicWriteEnd) and routed all four writers through it. This also encodes the FileMove gotcha once instead of per call site: the destination location comes from FILE_COMMON inside the 4th arg, NOT inherited from the source, and getting it wrong moves the file to the wrong sandbox silently. Also fixed while in these functions: - SaveTopologyConfiguration had 13 copy-pasted 6-line error blocks that each returned WITHOUT FileClose(handle), leaking the handle on every write failure. Collapsed to one ok-chain that closes exactly once. The on-disk field order and types are unchanged (asserted during the rewrite) so existing .cfg files still load. - SaveChartSignals documented that pruning runs only after a successful write ("a failed write above leaves both the file AND the chart untouched") but never checked any write result, so a partial write still deleted the chart objects. Results are checked now, making the existing comment true. Compiles 0 errors, 0 warnings. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 00:31:29 -04:00
}
fix(pool): defer the atomic promotion to the timer instead of spinning on the tick REPLACES the in-line retry from ad4ae58, which was the wrong shape and did not work. Measured after deploying it: atomic rename ... failed (error 5004) after 4 attempts MQL5 exposes no FILE_SHARE_DELETE, so a rename CANNOT succeed while any reader holds the destination open - it is not a lock that waiting longer wins. The retry assumed a peer holds a pool file for "tens of ms"; USDCAD_16388.bin is 134 MB and a peer reading it holds the handle for SECONDS. The loop lost every time and bought nothing but 75ms of tick latency on the failure path. The content is already written and correct - only the SWAP is blocked. So try the rename once, and on failure remember the temp and promote it from OnTimer, where I/O belongs. Once the reader closes, a single FileMove lands it. That beats the old fallback of waiting for the next full publish, which rewrites all 134 MB and may be an era away. * pending list is bounded (8) and deduplicated - AtomicWriteBegin reuses one temp name per file, so a second failure for the same file must not take a second slot. A full list falls back to the previous next-publish behaviour. * a successful write FORGETS any queued promotion for that name, so a stale temp can never overwrite fresher content. * a vanished temp (a later publish succeeded outright) is dropped, not retried. * a landed promotion is LOGGED. Silence is what made me misread the last attempt as working when there had simply been no contention in the window. Compiled clean; NOT yet run - and note that verification needs a collision to occur, which happened ~27 times across a whole day. Absence of the message in any one window is not evidence either way. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 13:41:35 -04:00
int lastErr = GetLastError();
AtomicRememberPending(finalName, commonFlag);
feat(vote): exit-on-reversal boolean, pin the threshold, retry the atomic rename THE EXIT KNOB. Exit_On_Reversal_Vote (default false) replaces the deleted Signal_ThresholdClose with one boolean: false pins the close threshold to an arithmetically unreachable 101, true pins it to the SAME threshold the entry uses - the seed at first, then the derived value, republished together whenever it moves. A second threshold was always redundant; "the bot now says the other way" is one question. It also arms CExpertSignalCustom::m_holdToBarrier, which was DEAD CODE: HoldToBarrier(bool) had no caller anywhere in the build, so the flag had been permanently false and the disabled close threshold was carrying the whole hold-to-barrier policy alone. Both halves now move together. Default stays false because the reason is statistical: the gate certifies P(label agrees | vote fired) against a label that runs to the barrier, so an early close trades something never measured. Turning it on is a different strategy, not a tightening of this one. THE PIN. The live threshold now moves only when an era's weights become the checkpoint, and freezes once g_ensDeployApproved. Every era still derives its own rung - that is how the best one is found - but the rung that TRADES belongs to the checkpoint, exactly as the weights do. Two reasons, one measured and one structural: the per-era rung moves on 6-34% of steps (the live run flapped SP500 15 -> 10 -> 15 within a minute of starting), and without the pin a later era's rung could end up applied to an earlier era's deployed model. A ladder restart releases the pin, since clearing the checkpoint clears what it pinned. The era line now prints the rung its own numbers came from, so it stays honest when that differs from the pinned one. THE ATOMIC RENAME retried zero times. Six charts share the TrainPool and AltData directories, so a publish regularly lands while a peer chart holds the destination open and FileMove returns 5004 - 27 times in one day on the live fleet. Nothing was lost (the temp keeps the new content, the old file stays intact) but the row did not update until the next publish. Now four attempts at 25ms, on the FAILURE PATH ONLY - a successful rename never sleeps - and skipped in the tester, where the contention cannot happen and Sleep would distort a pass. A rescued retry is logged, so worsening contention is visible. Retrain-neutral. Compiled clean; NOT yet run. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 09:53:56 -04:00
Print(context, ": atomic rename ", tmpName, " -> ", finalName, " failed (error ",
fix(pool): defer the atomic promotion to the timer instead of spinning on the tick REPLACES the in-line retry from ad4ae58, which was the wrong shape and did not work. Measured after deploying it: atomic rename ... failed (error 5004) after 4 attempts MQL5 exposes no FILE_SHARE_DELETE, so a rename CANNOT succeed while any reader holds the destination open - it is not a lock that waiting longer wins. The retry assumed a peer holds a pool file for "tens of ms"; USDCAD_16388.bin is 134 MB and a peer reading it holds the handle for SECONDS. The loop lost every time and bought nothing but 75ms of tick latency on the failure path. The content is already written and correct - only the SWAP is blocked. So try the rename once, and on failure remember the temp and promote it from OnTimer, where I/O belongs. Once the reader closes, a single FileMove lands it. That beats the old fallback of waiting for the next full publish, which rewrites all 134 MB and may be an era away. * pending list is bounded (8) and deduplicated - AtomicWriteBegin reuses one temp name per file, so a second failure for the same file must not take a second slot. A full list falls back to the previous next-publish behaviour. * a successful write FORGETS any queued promotion for that name, so a stale temp can never overwrite fresher content. * a vanished temp (a later publish succeeded outright) is dropped, not retried. * a landed promotion is LOGGED. Silence is what made me misread the last attempt as working when there had simply been no contention in the window. Compiled clean; NOT yet run - and note that verification needs a collision to occur, which happened ~27 times across a whole day. Absence of the message in any one window is not evidence either way. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 13:41:35 -04:00
IntegerToString(lastErr), ") - a peer chart is holding it open. The temp holds the new"
" content and ", finalName, " is unchanged; the timer will promote it as soon as the"
" reader closes, without rewriting the file.");
feat(vote): exit-on-reversal boolean, pin the threshold, retry the atomic rename THE EXIT KNOB. Exit_On_Reversal_Vote (default false) replaces the deleted Signal_ThresholdClose with one boolean: false pins the close threshold to an arithmetically unreachable 101, true pins it to the SAME threshold the entry uses - the seed at first, then the derived value, republished together whenever it moves. A second threshold was always redundant; "the bot now says the other way" is one question. It also arms CExpertSignalCustom::m_holdToBarrier, which was DEAD CODE: HoldToBarrier(bool) had no caller anywhere in the build, so the flag had been permanently false and the disabled close threshold was carrying the whole hold-to-barrier policy alone. Both halves now move together. Default stays false because the reason is statistical: the gate certifies P(label agrees | vote fired) against a label that runs to the barrier, so an early close trades something never measured. Turning it on is a different strategy, not a tightening of this one. THE PIN. The live threshold now moves only when an era's weights become the checkpoint, and freezes once g_ensDeployApproved. Every era still derives its own rung - that is how the best one is found - but the rung that TRADES belongs to the checkpoint, exactly as the weights do. Two reasons, one measured and one structural: the per-era rung moves on 6-34% of steps (the live run flapped SP500 15 -> 10 -> 15 within a minute of starting), and without the pin a later era's rung could end up applied to an earlier era's deployed model. A ladder restart releases the pin, since clearing the checkpoint clears what it pinned. The era line now prints the rung its own numbers came from, so it stays honest when that differs from the pinned one. THE ATOMIC RENAME retried zero times. Six charts share the TrainPool and AltData directories, so a publish regularly lands while a peer chart holds the destination open and FileMove returns 5004 - 27 times in one day on the live fleet. Nothing was lost (the temp keeps the new content, the old file stays intact) but the row did not update until the next publish. Now four attempts at 25ms, on the FAILURE PATH ONLY - a successful rename never sleeps - and skipped in the tester, where the contention cannot happen and Sleep would distort a pass. A rescued retry is logged, so worsening contention is visible. Retrain-neutral. Compiled clean; NOT yet run. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 09:53:56 -04:00
ResetLastError();
return false;
fix: make sidecar writes atomic; extract shared AtomicFile helper FileOpen(FILE_WRITE) truncates its target on open. CNet::Save already staged the .nnw through a temp file + rename for that reason, but the three sidecars written beside it did not: .stats ExpertSignalAIBase.mqh:5918 .arrows ExpertSignalAIBase.mqh:6224 .cfg ExpertSignalAIBase.mqh:7329 Two defects followed. 1. An interrupted write published a truncated sidecar. For .cfg that is the worst case: LoadAndCompareTopologyConfiguration() reads a short file as a mismatch, which discards the trained model and restarts from era 0. 2. Windows file sharing is a mutual contract - a writer opened with no FILE_SHARE_* blocks every concurrent open regardless of the reader's flags. All three read paths carry FILE_SHARE_READ|FILE_SHARE_WRITE specifically so a tester agent can read them while a live chart runs; an exclusive writer on the same path defeated that. Extracted CNet::Save's proven pattern into System\AtomicFile.mqh (AtomicWriteBegin/AtomicWriteEnd) and routed all four writers through it. This also encodes the FileMove gotcha once instead of per call site: the destination location comes from FILE_COMMON inside the 4th arg, NOT inherited from the source, and getting it wrong moves the file to the wrong sandbox silently. Also fixed while in these functions: - SaveTopologyConfiguration had 13 copy-pasted 6-line error blocks that each returned WITHOUT FileClose(handle), leaking the handle on every write failure. Collapsed to one ok-chain that closes exactly once. The on-disk field order and types are unchanged (asserted during the rewrite) so existing .cfg files still load. - SaveChartSignals documented that pruning runs only after a successful write ("a failed write above leaves both the file AND the chart untouched") but never checked any write result, so a partial write still deleted the chart objects. Results are checked now, making the existing comment true. Compiles 0 errors, 0 warnings. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 00:31:29 -04:00
}
//+------------------------------------------------------------------+
#endif // WARRIOR_ATOMIC_FILE_MQH