feat: mini-batch gradient accumulation (F4), front-end-aware capacity budget (F6), split Wyckoff categoricals (N1)
Completes the 2026-08-09 training audit. FORCES A RETRAIN of every
Wyckoff-enabled config (N1 re-keys the fingerprint), and BOTH DLLs must be
redeployed alongside the .ex5 - they carry new exports.
F4 - mini-batch accumulation, TRAIN_BATCH_SIZE=32. Training was pure online
SGD (one weight update per bar), which is the mechanical source of the
era-to-era whipsaw every downstream guard was built to cope with. The O(n^2)
outer product is native - AccumulateWeightGrad / AccumulateWeightGradConv /
AccumulateBufferInto in Network.cl, WarriorCPU and WarriorDML - while the
optimizer step is host-side MQL5 shared by all tiers (ApplyAccumToBlock), so
there is one Adam/SGD implementation instead of four that can drift.
- the LSTM needs no outer-product kernel (WeightsGradient already holds the
sample's full dW) but could NOT simply be left un-zeroed between samples:
CPU_LSTMSeqBackward/DML_LSTMSeqBackward memset it on entry. Hence a
separate accumulator plus an elementwise add.
- batch-norm gamma/beta accumulate in host arrays, not new BatchOptions
slots - BN_OPT_STRIDE is baked into every persisted .nnw.
- scoped to pass 2; online learning keeps immediate updates. Every save /
checkpoint / scoring boundary flushes, scaling by the real sample count.
- degrades to per-sample updates (one log line) on a tier that cannot
accumulate, so old devices and DLL-free builds are unaffected.
- verified offline: DirectML/batch_accum_check.cpp drives the real exports
against an independent reference; at B=1 the accumulator matches the
shipped unbatched kernel's own gradient to 1.1e-16. Math only - the
in-situ check remains the per-layer dW/W report on a real era.
F6 - ComputeFirstLayerWidth budgeted against the RAW input width even where a
conv/LSTM front end had already reduced it, so an LSTM's dense stack was
charged for 1,280 inputs when it receives 64. Confirmed from the deployed
.cfg files: CONV, LSTM and HYBRID were all pinned at the 16-unit floor. Now
budgeted against the front-end output and capped at it (never fan out), with
the derivation reordered so both stages settle first.
N1 - EventCode/EventPhase/StructuralPhase are signed categoricals packing
direction and Wyckoff stage into one scalar across a sign discontinuity. Split
into direction + [0,1] magnitude, the same convention the base OHLC block uses.
Information-preserving; 13 readings now occupy 16 inputs.
Compiled clean (0 errors, 0 warnings); both DLLs rebuilt.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 11:48:03 -04:00
|
|
|
//+------------------------------------------------------------------+
|
|
|
|
|
//| batch_accum_check.cpp |
|
|
|
|
|
//| |
|
|
|
|
|
//| Offline math check for the mini-batch accumulation exports added |
|
|
|
|
|
//| by the 2026-08-09 audit (F4). Same shape as lstm_seq_gradcheck |
|
|
|
|
|
//| beside it: link straight against WarriorCPU.dll, drive the real |
|
|
|
|
|
//| exports, compare against an independent reference computed here. |
|
|
|
|
|
//| |
|
|
|
|
|
//| WHAT THIS PROVES, precisely: |
|
|
|
|
|
//| 1. CPU_AccumulateWeightGrad over B samples equals the sum of |
|
|
|
|
|
//| the per-sample outer products g (x) x - i.e. the accumulator |
|
|
|
|
|
//| really is Sum(g_s (x) x_s), which is the whole claim the |
|
|
|
|
|
//| batched path rests on. |
|
|
|
|
|
//| 2. At B == 1 the accumulated gradient equals the gradient the |
|
|
|
|
|
//| UNBATCHED kernel forms internally, so batch size 1 is |
|
|
|
|
|
//| genuinely the old behaviour and not a near-miss. Checked by |
|
|
|
|
|
//| running CPU_UpdateWeightsAdam from a zeroed m/v/weight state,|
|
|
|
|
|
//| where its first step reduces to a known function of grad. |
|
|
|
|
|
//| 3. The conv accumulator matches a direct transcription of the |
|
|
|
|
|
//| sliding-window gradient, including the bias row. |
|
|
|
|
|
//| 4. CPU_AccumulateBufferInto is an exact elementwise add. |
|
|
|
|
|
//| |
|
|
|
|
|
//| WHAT IT DOES NOT PROVE - and this matters, see |
|
|
|
|
|
//| [[feedback_verify_in_situ_not_offline]]: that a layer TRAINS in |
|
|
|
|
|
//| the assembled network. It is a math check on four functions, not |
|
|
|
|
|
//| a training run. The in-situ check is the per-layer dW/W report on |
|
|
|
|
|
//| a real era (CNet::LayerLearningReport). |
|
|
|
|
|
//| |
|
|
|
|
|
//| Build (from a plain cmd, after build_cpu.bat): |
|
|
|
|
|
//| cl /nologo /EHsc /O2 /std:c++17 batch_accum_check.cpp WarriorCPU.lib
|
|
|
|
|
//+------------------------------------------------------------------+
|
|
|
|
|
#include <cstdio>
|
|
|
|
|
#include <cmath>
|
|
|
|
|
#include <vector>
|
|
|
|
|
#include <random>
|
|
|
|
|
|
|
|
|
|
// WarriorCPU.h is deliberately NOT included: it hardcodes WARRIORCPU_API to __declspec(dllexport),
|
|
|
|
|
// which is right for building the DLL and wrong for consuming it (a consumer would try to re-export
|
|
|
|
|
// every symbol instead of importing it, and the link fails). Redeclaring just the entry points this
|
|
|
|
|
// check drives keeps the shared header untouched. Signatures must match WarriorCPU.h exactly.
|
|
|
|
|
typedef long long CpuHandle;
|
|
|
|
|
extern "C"
|
|
|
|
|
{
|
|
|
|
|
__declspec(dllimport) CpuHandle __stdcall CPU_Init(int threads);
|
|
|
|
|
__declspec(dllimport) void __stdcall CPU_Shutdown(CpuHandle ctx);
|
|
|
|
|
__declspec(dllimport) int __stdcall CPU_BufferCreate(CpuHandle ctx, int elementCount);
|
|
|
|
|
__declspec(dllimport) int __stdcall CPU_BufferWrite(CpuHandle ctx, int handle, const double *data, int count);
|
|
|
|
|
__declspec(dllimport) int __stdcall CPU_BufferRead(CpuHandle ctx, int handle, double *data, int count);
|
|
|
|
|
__declspec(dllimport) void __stdcall CPU_BufferFree(CpuHandle ctx, int handle);
|
|
|
|
|
__declspec(dllimport) int __stdcall CPU_UpdateWeightsAdam(CpuHandle ctx, int wHandle, int gHandle, int iHandle,
|
|
|
|
|
int mHandle, int vHandle, int inputs, double lt, double b1, double b2, int neurons);
|
|
|
|
|
__declspec(dllimport) int __stdcall CPU_AccumulateWeightGrad(CpuHandle ctx, int accHandle, int gHandle, int iHandle,
|
|
|
|
|
int inputs, int neurons);
|
|
|
|
|
__declspec(dllimport) int __stdcall CPU_AccumulateWeightGradConv(CpuHandle ctx, int accHandle, int gHandle, int iHandle,
|
|
|
|
|
int inputs, int windowIn, int windowOut, int step);
|
|
|
|
|
__declspec(dllimport) int __stdcall CPU_AccumulateBufferInto(CpuHandle ctx, int dstHandle, int srcHandle, int count);
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
namespace
|
|
|
|
|
{
|
|
|
|
|
int g_failures = 0;
|
|
|
|
|
|
|
|
|
|
void Check(const char *what, double got, double want, double tol = 1e-9)
|
|
|
|
|
{
|
|
|
|
|
double diff = std::fabs(got - want);
|
|
|
|
|
if(!(diff <= tol))
|
|
|
|
|
{
|
|
|
|
|
std::printf(" FAIL %-38s got %.17g want %.17g (diff %.3g)\n", what, got, want, diff);
|
|
|
|
|
++g_failures;
|
|
|
|
|
}
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
void CheckMax(const char *what, double maxDiff, double tol = 1e-9)
|
|
|
|
|
{
|
|
|
|
|
if(!(maxDiff <= tol))
|
|
|
|
|
{
|
|
|
|
|
std::printf(" FAIL %-38s max |diff| %.3g > %.3g\n", what, maxDiff, tol);
|
|
|
|
|
++g_failures;
|
|
|
|
|
}
|
|
|
|
|
else
|
|
|
|
|
std::printf(" ok %-38s max |diff| %.3g\n", what, maxDiff);
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
int MakeBuffer(CpuHandle ctx, const std::vector<double> &v)
|
|
|
|
|
{
|
|
|
|
|
int h = CPU_BufferCreate(ctx, (int)v.size());
|
|
|
|
|
if(h >= 0 && !v.empty())
|
|
|
|
|
CPU_BufferWrite(ctx, h, v.data(), (int)v.size());
|
|
|
|
|
return h;
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
std::vector<double> ReadBuffer(CpuHandle ctx, int handle, int count)
|
|
|
|
|
{
|
|
|
|
|
std::vector<double> out((size_t)count, 0.0);
|
|
|
|
|
CPU_BufferRead(ctx, handle, out.data(), count);
|
|
|
|
|
return out;
|
|
|
|
|
}
|
|
|
|
|
} // namespace
|
|
|
|
|
|
|
|
|
|
//+------------------------------------------------------------------+
|
|
|
|
|
//| 1 + 2: dense accumulation equals the summed outer product. |
|
|
|
|
|
//+------------------------------------------------------------------+
|
|
|
|
|
static void TestDense(CpuHandle ctx, int neurons, int inputs, int batch)
|
|
|
|
|
{
|
|
|
|
|
std::mt19937 rng(1234u);
|
|
|
|
|
std::uniform_real_distribution<double> dist(-1.5, 1.5);
|
|
|
|
|
|
|
|
|
|
const int weightCount = neurons * (inputs + 1);
|
|
|
|
|
std::vector<double> acc((size_t)weightCount, 0.0);
|
|
|
|
|
int accH = MakeBuffer(ctx, acc);
|
|
|
|
|
int gH = CPU_BufferCreate(ctx, neurons);
|
|
|
|
|
int iH = CPU_BufferCreate(ctx, inputs);
|
|
|
|
|
|
|
|
|
|
// Independent reference: plain double-precision sum of outer products, bias slot fed a constant 1.
|
|
|
|
|
std::vector<double> reference((size_t)weightCount, 0.0);
|
|
|
|
|
for(int s = 0; s < batch; ++s)
|
|
|
|
|
{
|
|
|
|
|
std::vector<double> g((size_t)neurons), x((size_t)inputs);
|
|
|
|
|
for(auto &val : g)
|
|
|
|
|
val = dist(rng);
|
|
|
|
|
for(auto &val : x)
|
|
|
|
|
val = dist(rng);
|
|
|
|
|
CPU_BufferWrite(ctx, gH, g.data(), neurons);
|
|
|
|
|
CPU_BufferWrite(ctx, iH, x.data(), inputs);
|
|
|
|
|
if(!CPU_AccumulateWeightGrad(ctx, accH, gH, iH, inputs, neurons))
|
|
|
|
|
{
|
|
|
|
|
std::printf(" FAIL CPU_AccumulateWeightGrad returned 0\n");
|
|
|
|
|
++g_failures;
|
|
|
|
|
return;
|
|
|
|
|
}
|
|
|
|
|
for(int i = 0; i < neurons; ++i)
|
|
|
|
|
for(int j = 0; j <= inputs; ++j)
|
|
|
|
|
reference[(size_t)i * (inputs + 1) + j] += g[i] * (j < inputs ? x[j] : 1.0);
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
std::vector<double> got = ReadBuffer(ctx, accH, weightCount);
|
|
|
|
|
double maxDiff = 0.0;
|
|
|
|
|
for(int k = 0; k < weightCount; ++k)
|
|
|
|
|
maxDiff = std::fmax(maxDiff, std::fabs(got[k] - reference[k]));
|
|
|
|
|
char label[96];
|
|
|
|
|
std::snprintf(label, sizeof(label), "dense accum %dx%d over %d samples", neurons, inputs, batch);
|
|
|
|
|
CheckMax(label, maxDiff);
|
|
|
|
|
|
|
|
|
|
CPU_BufferFree(ctx, accH);
|
|
|
|
|
CPU_BufferFree(ctx, gH);
|
|
|
|
|
CPU_BufferFree(ctx, iH);
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
//+------------------------------------------------------------------+
|
|
|
|
|
//| Batch size 1 must reproduce the unbatched kernel's own gradient. |
|
|
|
|
|
//| |
|
|
|
|
|
//| From m = v = 0 the Adam step is |
|
|
|
|
|
//| mt = (1-b1) g, vt = sqrt((1-b2) g^2) = sqrt(1-b2) |g| |
|
|
|
|
|
//| delta = lt*mt/vt - lt*decay*w (clamped) |
|
|
|
|
|
//| so with w = 0 the applied delta pins down |g| exactly. Comparing |
|
|
|
|
|
//| that against the accumulator proves both paths form the SAME |
|
|
|
|
|
//| gradient, which is the property "batch size 1 == old behaviour" |
|
|
|
|
|
//| actually depends on. |
|
|
|
|
|
//+------------------------------------------------------------------+
|
|
|
|
|
static void TestDenseMatchesUnbatched(CpuHandle ctx)
|
|
|
|
|
{
|
|
|
|
|
const int neurons = 3, inputs = 5;
|
|
|
|
|
const int weightCount = neurons * (inputs + 1);
|
|
|
|
|
const double b1 = 0.9, b2 = 0.999, lt = 1e-3;
|
|
|
|
|
|
|
|
|
|
std::mt19937 rng(99u);
|
|
|
|
|
std::uniform_real_distribution<double> dist(-1.2, 1.2);
|
|
|
|
|
std::vector<double> g((size_t)neurons), x((size_t)inputs);
|
|
|
|
|
for(auto &val : g)
|
|
|
|
|
val = dist(rng);
|
|
|
|
|
for(auto &val : x)
|
|
|
|
|
val = dist(rng);
|
|
|
|
|
|
|
|
|
|
int gH = MakeBuffer(ctx, g);
|
|
|
|
|
int iH = MakeBuffer(ctx, x);
|
|
|
|
|
|
|
|
|
|
// Path A - the accumulator.
|
|
|
|
|
std::vector<double> zeros((size_t)weightCount, 0.0);
|
|
|
|
|
int accH = MakeBuffer(ctx, zeros);
|
|
|
|
|
CPU_AccumulateWeightGrad(ctx, accH, gH, iH, inputs, neurons);
|
|
|
|
|
std::vector<double> accumulated = ReadBuffer(ctx, accH, weightCount);
|
|
|
|
|
|
|
|
|
|
// Path B - the shipped unbatched Adam kernel, from a zeroed state.
|
|
|
|
|
int wH = MakeBuffer(ctx, zeros);
|
|
|
|
|
int mH = MakeBuffer(ctx, zeros);
|
|
|
|
|
int vH = MakeBuffer(ctx, zeros);
|
|
|
|
|
CPU_UpdateWeightsAdam(ctx, wH, gH, iH, mH, vH, inputs, lt, b1, b2, neurons);
|
|
|
|
|
std::vector<double> mAfter = ReadBuffer(ctx, mH, weightCount);
|
|
|
|
|
|
|
|
|
|
// m after one step is (1-b1)*grad, so grad is recoverable exactly.
|
|
|
|
|
double maxDiff = 0.0;
|
|
|
|
|
for(int k = 0; k < weightCount; ++k)
|
|
|
|
|
{
|
|
|
|
|
double kernelGrad = mAfter[k] / (1.0 - b1);
|
|
|
|
|
maxDiff = std::fmax(maxDiff, std::fabs(kernelGrad - accumulated[k]));
|
|
|
|
|
}
|
|
|
|
|
CheckMax("B=1 accum == unbatched kernel gradient", maxDiff, 1e-12);
|
|
|
|
|
|
|
|
|
|
CPU_BufferFree(ctx, gH);
|
|
|
|
|
CPU_BufferFree(ctx, iH);
|
|
|
|
|
CPU_BufferFree(ctx, accH);
|
|
|
|
|
CPU_BufferFree(ctx, wH);
|
|
|
|
|
CPU_BufferFree(ctx, mH);
|
|
|
|
|
CPU_BufferFree(ctx, vH);
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
//+------------------------------------------------------------------+
|
|
|
|
|
//| 3: conv accumulation against a direct transcription. |
|
|
|
|
|
//+------------------------------------------------------------------+
|
|
|
|
|
static void TestConv(CpuHandle ctx, int windowIn, int windowOut, int step, int inputs, int batch)
|
|
|
|
|
{
|
|
|
|
|
std::mt19937 rng(4242u);
|
|
|
|
|
std::uniform_real_distribution<double> dist(-1.0, 1.0);
|
|
|
|
|
|
|
|
|
|
int total = (windowIn + 1) * windowOut;
|
|
|
|
|
int positions = (inputs - (windowIn - step)) % step;
|
|
|
|
|
positions = (inputs - (windowIn - step) - positions) / step + (positions > 0 ? 1 : 0);
|
|
|
|
|
int gradCount = positions * windowOut;
|
|
|
|
|
|
|
|
|
|
std::vector<double> acc((size_t)total, 0.0);
|
|
|
|
|
int accH = MakeBuffer(ctx, acc);
|
|
|
|
|
int gH = CPU_BufferCreate(ctx, gradCount);
|
|
|
|
|
int iH = CPU_BufferCreate(ctx, inputs);
|
|
|
|
|
|
|
|
|
|
std::vector<double> reference((size_t)total, 0.0);
|
|
|
|
|
for(int s = 0; s < batch; ++s)
|
|
|
|
|
{
|
|
|
|
|
std::vector<double> g((size_t)gradCount), x((size_t)inputs);
|
|
|
|
|
for(auto &val : g)
|
|
|
|
|
val = dist(rng);
|
|
|
|
|
for(auto &val : x)
|
|
|
|
|
val = dist(rng);
|
|
|
|
|
CPU_BufferWrite(ctx, gH, g.data(), gradCount);
|
|
|
|
|
CPU_BufferWrite(ctx, iH, x.data(), inputs);
|
|
|
|
|
if(!CPU_AccumulateWeightGradConv(ctx, accH, gH, iH, inputs, windowIn, windowOut, step))
|
|
|
|
|
{
|
|
|
|
|
std::printf(" FAIL CPU_AccumulateWeightGradConv returned 0\n");
|
|
|
|
|
++g_failures;
|
|
|
|
|
return;
|
|
|
|
|
}
|
|
|
|
|
for(int w = 0; w < total; ++w)
|
|
|
|
|
{
|
|
|
|
|
int shift = w % (windowIn + 1);
|
|
|
|
|
int shiftOut = (w - shift) / (windowIn + 1);
|
|
|
|
|
double grad = 0.0;
|
|
|
|
|
for(int t = 0; t < positions; ++t)
|
|
|
|
|
{
|
|
|
|
|
if(shift != windowIn && (shift + t * step) >= inputs)
|
|
|
|
|
break;
|
|
|
|
|
int gi = t * windowOut + shiftOut;
|
|
|
|
|
int ii = shift + t * step;
|
|
|
|
|
if(gi >= gradCount || (shift != windowIn && ii >= inputs))
|
|
|
|
|
break;
|
|
|
|
|
grad += g[(size_t)gi] * (shift == windowIn ? 1.0 : x[(size_t)ii]);
|
|
|
|
|
}
|
|
|
|
|
reference[(size_t)w] += grad;
|
|
|
|
|
}
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
std::vector<double> got = ReadBuffer(ctx, accH, total);
|
|
|
|
|
double maxDiff = 0.0;
|
|
|
|
|
for(int k = 0; k < total; ++k)
|
|
|
|
|
maxDiff = std::fmax(maxDiff, std::fabs(got[k] - reference[k]));
|
|
|
|
|
char label[96];
|
|
|
|
|
std::snprintf(label, sizeof(label), "conv accum w%d/o%d/s%d over %d samples",
|
|
|
|
|
windowIn, windowOut, step, batch);
|
|
|
|
|
CheckMax(label, maxDiff);
|
|
|
|
|
|
|
|
|
|
CPU_BufferFree(ctx, accH);
|
|
|
|
|
CPU_BufferFree(ctx, gH);
|
|
|
|
|
CPU_BufferFree(ctx, iH);
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
//+------------------------------------------------------------------+
|
|
|
|
|
//| 4: elementwise add (the LSTM's batch path). |
|
|
|
|
|
//+------------------------------------------------------------------+
|
|
|
|
|
static void TestBufferAdd(CpuHandle ctx)
|
|
|
|
|
{
|
|
|
|
|
const int count = 257; // deliberately not a multiple of any thread-block size
|
|
|
|
|
std::mt19937 rng(7u);
|
|
|
|
|
std::uniform_real_distribution<double> dist(-5.0, 5.0);
|
|
|
|
|
std::vector<double> dst((size_t)count), src((size_t)count), reference((size_t)count);
|
|
|
|
|
for(int i = 0; i < count; ++i)
|
|
|
|
|
{
|
|
|
|
|
dst[(size_t)i] = dist(rng);
|
|
|
|
|
src[(size_t)i] = dist(rng);
|
|
|
|
|
}
|
|
|
|
|
int dstH = MakeBuffer(ctx, dst);
|
|
|
|
|
int srcH = MakeBuffer(ctx, src);
|
|
|
|
|
|
|
|
|
|
const int rounds = 4;
|
|
|
|
|
reference = dst;
|
|
|
|
|
for(int r = 0; r < rounds; ++r)
|
|
|
|
|
{
|
|
|
|
|
CPU_AccumulateBufferInto(ctx, dstH, srcH, count);
|
|
|
|
|
for(int i = 0; i < count; ++i)
|
|
|
|
|
reference[(size_t)i] += src[(size_t)i];
|
|
|
|
|
}
|
|
|
|
|
std::vector<double> got = ReadBuffer(ctx, dstH, count);
|
|
|
|
|
double maxDiff = 0.0;
|
|
|
|
|
for(int i = 0; i < count; ++i)
|
|
|
|
|
maxDiff = std::fmax(maxDiff, std::fabs(got[(size_t)i] - reference[(size_t)i]));
|
|
|
|
|
CheckMax("buffer add, 4 rounds", maxDiff);
|
|
|
|
|
|
|
|
|
|
CPU_BufferFree(ctx, dstH);
|
|
|
|
|
CPU_BufferFree(ctx, srcH);
|
|
|
|
|
}
|
|
|
|
|
|
fix: the Adam second moment was never Adam - all four tiers
Root cause of the B=32 regression, and it predates F4 entirely. Every Adam
kernel stored v already square-rooted and then fed that stored value back in
as if it were the variance:
v_new = sqrt(b2 * v_old + (1 - b2) * g^2)
That recursion has a fixed point at v ~= b2 = 0.999 for ANY gradient below
unit scale, so the denominator stops tracking the gradient and Adam degrades
into plain SGD with lr = lt. Measured against the shipped WarriorCPU.dll
(batch_accum_check.cpp, TestOptimizerScaleInvariance), 4000 steps of a
constant gradient: 3285x less displacement at |g|=1e-5 than at |g|=1, where
a scale-invariant optimizer gives the same distance for both. After the fix
all six magnitudes read 1.199 and v tracks |g| exactly.
It hit conv/LSTM specifically because they sit behind a batch-norm with
running variance ~2.6e+05, so their gradients arrive divided by ~500 - deep
in the degraded regime - while the dense stack near the loss stayed in the
working one. In situ on SP500 H1: lstm1 dW/W 2.62/10.0/7.14% -> 0.024/0.022/
0.003%, conv1 decaying to 0.000% by era 30. NeuronBatchNorm.mqh already
squared v back for gamma/beta and its comment named the kernels as wrong,
which is exactly why gamma/beta kept training while the stages behind froze.
Persisted .nnw needs no migration - v keeps its std-dev meaning.
Also, the two ways F4 exposed it, both mine:
- No LR compensation for B fewer steps per era. sqrt(B) for adaptive methods
(Krizhevsky 2014; Granziol et al. 2022), applied once in
InitialEtaForOptimizer(). Linear scaling (Goyal et al. 2017) is for SGD.
- Plateau patience denominated in eras, so raising B made the ladder 32x more
impatient in its only unit. PAI converged at era 41 on ~49k updates where
the same config had been finding new bests at era 1028.
TrainPlateauPatienceEras() stretches it by the same sqrt(B).
TRAIN_BATCH_SIZE 32 -> 8 so the patience stretch stays affordable (8 -> 23
eras per stage, not 8 -> 45). Both helpers are identities at B=1.
Deploy gate: DEPLOY_MIN_SIDE_RECALL_PCT (10%) folded into tradeableOK. The
perceptron reported Sell:0% recall in all 41 eras, cleared the floor on Buy
alone at 36.6% vs 34% chance, deployed, and sprayed buy arrows. Folded into
the ranking key rather than checked at deploy time so a one-sided era cannot
become best-so-far in the first place.
Deinit: the arrow purge now runs BEFORE ExtPanel.Destroy(), an unbounded
CAppDialog teardown that sat ahead of it - the same ordering inversion the
rule there exists to prevent. CONV was force-terminated 4.8 s into OnDeinit
(vs ~1.1 s for the three that finished) having reached none of its cleanup,
so its arrows stayed on the chart. Steps are now timed in the log.
PurgeChart's verification rescan filtered on OBJ_ARROW, the same blind spot
as the bulk delete, so "persisted 10 ... cleared 0" passed silently. It now
walks every object type and reports the object counts when both are zero.
Both build variants compile 0 errors / 0 warnings; both DLLs rebuilt.
FORCES A RETRAIN (already forced by N1) and both DLLs must ship with the .ex5.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 14:02:35 -04:00
|
|
|
//+------------------------------------------------------------------+
|
|
|
|
|
//| 5: IS THE OPTIMIZER SCALE-INVARIANT? (2026-08-09, F4 regression) |
|
|
|
|
|
//| |
|
|
|
|
|
//| Textbook Adam moves the same distance per step whatever the |
|
|
|
|
|
//| gradient's magnitude - that invariance is the whole reason it is |
|
|
|
|
|
//| usable on a net whose stages see wildly different gradient scales. |
|
|
|
|
|
//| This drives the SHIPPED kernel directly, at six magnitudes, and |
|
|
|
|
|
//| prints the displacement after a fixed number of steps. Invariance |
|
|
|
|
|
//| means one number repeated; anything proportional to |g| means the |
|
|
|
|
|
//| layers behind a batch-norm (whose gradients arrive divided by |
|
|
|
|
|
//| sqrt(var), ~500x here) are effectively on plain SGD - and that |
|
|
|
|
|
//| every gradient-shrinking change, mini-batching included, costs |
|
|
|
|
|
//| them progress in direct proportion. |
|
|
|
|
|
//| |
|
|
|
|
|
//| Diagnostic, not pass/fail: it reports, and asserts only the |
|
|
|
|
|
//| invariance claim itself so a future fix is what makes it pass. |
|
|
|
|
|
//+------------------------------------------------------------------+
|
|
|
|
|
static void TestOptimizerScaleInvariance(CpuHandle ctx)
|
|
|
|
|
{
|
|
|
|
|
const int neurons = 1, inputs = 1;
|
|
|
|
|
const int weightCount = neurons * (inputs + 1);
|
|
|
|
|
const double b1 = 0.9, b2 = 0.999, eta = 3e-4;
|
|
|
|
|
const int steps = 4000;
|
|
|
|
|
const double mags[] = { 1e+0, 1e-1, 1e-2, 1e-3, 1e-4, 1e-5 };
|
|
|
|
|
|
|
|
|
|
std::printf(" -- optimizer scale invariance (%d steps of a CONSTANT gradient, eta %.0e)\n", steps, eta);
|
|
|
|
|
double first = 0.0, worstRatio = 1.0;
|
|
|
|
|
for(int mi = 0; mi < (int)(sizeof(mags) / sizeof(mags[0])); ++mi)
|
|
|
|
|
{
|
|
|
|
|
// A constant gradient of exactly mags[mi] reaches the weight as g[0]*x[0].
|
|
|
|
|
std::vector<double> g((size_t)neurons, mags[mi]), x((size_t)inputs, 1.0);
|
|
|
|
|
std::vector<double> zeros((size_t)weightCount, 0.0);
|
|
|
|
|
int gH = MakeBuffer(ctx, g), iH = MakeBuffer(ctx, x);
|
|
|
|
|
int wH = MakeBuffer(ctx, zeros), mH = MakeBuffer(ctx, zeros), vH = MakeBuffer(ctx, zeros);
|
|
|
|
|
|
|
|
|
|
for(int t = 1; t <= steps; ++t)
|
|
|
|
|
{
|
|
|
|
|
// Bias correction exactly as CNeuronBaseOCL::updateInputWeights forms it.
|
|
|
|
|
double lt = eta * std::sqrt(1.0 - std::pow(b2, t)) / (1.0 - std::pow(b1, t));
|
|
|
|
|
CPU_UpdateWeightsAdam(ctx, wH, gH, iH, mH, vH, inputs, lt, b1, b2, neurons);
|
|
|
|
|
}
|
|
|
|
|
std::vector<double> w = ReadBuffer(ctx, wH, weightCount);
|
|
|
|
|
std::vector<double> v = ReadBuffer(ctx, vH, weightCount);
|
|
|
|
|
if(mi == 0)
|
|
|
|
|
first = std::fabs(w[0]);
|
|
|
|
|
double ratio = (std::fabs(w[0]) > 0.0) ? first / std::fabs(w[0]) : 1e30;
|
|
|
|
|
worstRatio = std::fmax(worstRatio, ratio);
|
|
|
|
|
std::printf(" |g| %7.0e -> displacement %10.3e stored v %.6f %.0fx less than |g|=1\n",
|
|
|
|
|
mags[mi], std::fabs(w[0]), v[0], ratio);
|
|
|
|
|
|
|
|
|
|
CPU_BufferFree(ctx, gH);
|
|
|
|
|
CPU_BufferFree(ctx, iH);
|
|
|
|
|
CPU_BufferFree(ctx, wH);
|
|
|
|
|
CPU_BufferFree(ctx, mH);
|
|
|
|
|
CPU_BufferFree(ctx, vH);
|
|
|
|
|
}
|
|
|
|
|
// Scale-invariant means the displacement barely moves across five decades of |g|.
|
|
|
|
|
CheckMax("optimizer is scale-invariant across 5 decades of |g|", worstRatio - 1.0, 1.0);
|
|
|
|
|
}
|
|
|
|
|
|
feat: mini-batch gradient accumulation (F4), front-end-aware capacity budget (F6), split Wyckoff categoricals (N1)
Completes the 2026-08-09 training audit. FORCES A RETRAIN of every
Wyckoff-enabled config (N1 re-keys the fingerprint), and BOTH DLLs must be
redeployed alongside the .ex5 - they carry new exports.
F4 - mini-batch accumulation, TRAIN_BATCH_SIZE=32. Training was pure online
SGD (one weight update per bar), which is the mechanical source of the
era-to-era whipsaw every downstream guard was built to cope with. The O(n^2)
outer product is native - AccumulateWeightGrad / AccumulateWeightGradConv /
AccumulateBufferInto in Network.cl, WarriorCPU and WarriorDML - while the
optimizer step is host-side MQL5 shared by all tiers (ApplyAccumToBlock), so
there is one Adam/SGD implementation instead of four that can drift.
- the LSTM needs no outer-product kernel (WeightsGradient already holds the
sample's full dW) but could NOT simply be left un-zeroed between samples:
CPU_LSTMSeqBackward/DML_LSTMSeqBackward memset it on entry. Hence a
separate accumulator plus an elementwise add.
- batch-norm gamma/beta accumulate in host arrays, not new BatchOptions
slots - BN_OPT_STRIDE is baked into every persisted .nnw.
- scoped to pass 2; online learning keeps immediate updates. Every save /
checkpoint / scoring boundary flushes, scaling by the real sample count.
- degrades to per-sample updates (one log line) on a tier that cannot
accumulate, so old devices and DLL-free builds are unaffected.
- verified offline: DirectML/batch_accum_check.cpp drives the real exports
against an independent reference; at B=1 the accumulator matches the
shipped unbatched kernel's own gradient to 1.1e-16. Math only - the
in-situ check remains the per-layer dW/W report on a real era.
F6 - ComputeFirstLayerWidth budgeted against the RAW input width even where a
conv/LSTM front end had already reduced it, so an LSTM's dense stack was
charged for 1,280 inputs when it receives 64. Confirmed from the deployed
.cfg files: CONV, LSTM and HYBRID were all pinned at the 16-unit floor. Now
budgeted against the front-end output and capped at it (never fan out), with
the derivation reordered so both stages settle first.
N1 - EventCode/EventPhase/StructuralPhase are signed categoricals packing
direction and Wyckoff stage into one scalar across a sign discontinuity. Split
into direction + [0,1] magnitude, the same convention the base OHLC block uses.
Information-preserving; 13 readings now occupy 16 inputs.
Compiled clean (0 errors, 0 warnings); both DLLs rebuilt.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 11:48:03 -04:00
|
|
|
int main()
|
|
|
|
|
{
|
|
|
|
|
CpuHandle ctx = CPU_Init(2);
|
|
|
|
|
if(!ctx)
|
|
|
|
|
{
|
|
|
|
|
std::printf("CPU_Init failed\n");
|
|
|
|
|
return 2;
|
|
|
|
|
}
|
|
|
|
|
std::printf("Mini-batch accumulation check (WarriorCPU.dll)\n");
|
|
|
|
|
|
|
|
|
|
TestDense(ctx, 4, 7, 1);
|
|
|
|
|
TestDense(ctx, 16, 64, 32);
|
|
|
|
|
TestDense(ctx, 3, 1281, 32); // the shipped SP500 H1 raw input width
|
|
|
|
|
TestDenseMatchesUnbatched(ctx);
|
|
|
|
|
TestConv(ctx, 6, 4, 2, 40, 1);
|
|
|
|
|
TestConv(ctx, 192, 32, 64, 1280, 32); // the shipped conv shape: 3 bars x 64 features, step 1 bar
|
|
|
|
|
TestBufferAdd(ctx);
|
fix: the Adam second moment was never Adam - all four tiers
Root cause of the B=32 regression, and it predates F4 entirely. Every Adam
kernel stored v already square-rooted and then fed that stored value back in
as if it were the variance:
v_new = sqrt(b2 * v_old + (1 - b2) * g^2)
That recursion has a fixed point at v ~= b2 = 0.999 for ANY gradient below
unit scale, so the denominator stops tracking the gradient and Adam degrades
into plain SGD with lr = lt. Measured against the shipped WarriorCPU.dll
(batch_accum_check.cpp, TestOptimizerScaleInvariance), 4000 steps of a
constant gradient: 3285x less displacement at |g|=1e-5 than at |g|=1, where
a scale-invariant optimizer gives the same distance for both. After the fix
all six magnitudes read 1.199 and v tracks |g| exactly.
It hit conv/LSTM specifically because they sit behind a batch-norm with
running variance ~2.6e+05, so their gradients arrive divided by ~500 - deep
in the degraded regime - while the dense stack near the loss stayed in the
working one. In situ on SP500 H1: lstm1 dW/W 2.62/10.0/7.14% -> 0.024/0.022/
0.003%, conv1 decaying to 0.000% by era 30. NeuronBatchNorm.mqh already
squared v back for gamma/beta and its comment named the kernels as wrong,
which is exactly why gamma/beta kept training while the stages behind froze.
Persisted .nnw needs no migration - v keeps its std-dev meaning.
Also, the two ways F4 exposed it, both mine:
- No LR compensation for B fewer steps per era. sqrt(B) for adaptive methods
(Krizhevsky 2014; Granziol et al. 2022), applied once in
InitialEtaForOptimizer(). Linear scaling (Goyal et al. 2017) is for SGD.
- Plateau patience denominated in eras, so raising B made the ladder 32x more
impatient in its only unit. PAI converged at era 41 on ~49k updates where
the same config had been finding new bests at era 1028.
TrainPlateauPatienceEras() stretches it by the same sqrt(B).
TRAIN_BATCH_SIZE 32 -> 8 so the patience stretch stays affordable (8 -> 23
eras per stage, not 8 -> 45). Both helpers are identities at B=1.
Deploy gate: DEPLOY_MIN_SIDE_RECALL_PCT (10%) folded into tradeableOK. The
perceptron reported Sell:0% recall in all 41 eras, cleared the floor on Buy
alone at 36.6% vs 34% chance, deployed, and sprayed buy arrows. Folded into
the ranking key rather than checked at deploy time so a one-sided era cannot
become best-so-far in the first place.
Deinit: the arrow purge now runs BEFORE ExtPanel.Destroy(), an unbounded
CAppDialog teardown that sat ahead of it - the same ordering inversion the
rule there exists to prevent. CONV was force-terminated 4.8 s into OnDeinit
(vs ~1.1 s for the three that finished) having reached none of its cleanup,
so its arrows stayed on the chart. Steps are now timed in the log.
PurgeChart's verification rescan filtered on OBJ_ARROW, the same blind spot
as the bulk delete, so "persisted 10 ... cleared 0" passed silently. It now
walks every object type and reports the object counts when both are zero.
Both build variants compile 0 errors / 0 warnings; both DLLs rebuilt.
FORCES A RETRAIN (already forced by N1) and both DLLs must ship with the .ex5.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 14:02:35 -04:00
|
|
|
TestOptimizerScaleInvariance(ctx);
|
feat: mini-batch gradient accumulation (F4), front-end-aware capacity budget (F6), split Wyckoff categoricals (N1)
Completes the 2026-08-09 training audit. FORCES A RETRAIN of every
Wyckoff-enabled config (N1 re-keys the fingerprint), and BOTH DLLs must be
redeployed alongside the .ex5 - they carry new exports.
F4 - mini-batch accumulation, TRAIN_BATCH_SIZE=32. Training was pure online
SGD (one weight update per bar), which is the mechanical source of the
era-to-era whipsaw every downstream guard was built to cope with. The O(n^2)
outer product is native - AccumulateWeightGrad / AccumulateWeightGradConv /
AccumulateBufferInto in Network.cl, WarriorCPU and WarriorDML - while the
optimizer step is host-side MQL5 shared by all tiers (ApplyAccumToBlock), so
there is one Adam/SGD implementation instead of four that can drift.
- the LSTM needs no outer-product kernel (WeightsGradient already holds the
sample's full dW) but could NOT simply be left un-zeroed between samples:
CPU_LSTMSeqBackward/DML_LSTMSeqBackward memset it on entry. Hence a
separate accumulator plus an elementwise add.
- batch-norm gamma/beta accumulate in host arrays, not new BatchOptions
slots - BN_OPT_STRIDE is baked into every persisted .nnw.
- scoped to pass 2; online learning keeps immediate updates. Every save /
checkpoint / scoring boundary flushes, scaling by the real sample count.
- degrades to per-sample updates (one log line) on a tier that cannot
accumulate, so old devices and DLL-free builds are unaffected.
- verified offline: DirectML/batch_accum_check.cpp drives the real exports
against an independent reference; at B=1 the accumulator matches the
shipped unbatched kernel's own gradient to 1.1e-16. Math only - the
in-situ check remains the per-layer dW/W report on a real era.
F6 - ComputeFirstLayerWidth budgeted against the RAW input width even where a
conv/LSTM front end had already reduced it, so an LSTM's dense stack was
charged for 1,280 inputs when it receives 64. Confirmed from the deployed
.cfg files: CONV, LSTM and HYBRID were all pinned at the 16-unit floor. Now
budgeted against the front-end output and capped at it (never fan out), with
the derivation reordered so both stages settle first.
N1 - EventCode/EventPhase/StructuralPhase are signed categoricals packing
direction and Wyckoff stage into one scalar across a sign discontinuity. Split
into direction + [0,1] magnitude, the same convention the base OHLC block uses.
Information-preserving; 13 readings now occupy 16 inputs.
Compiled clean (0 errors, 0 warnings); both DLLs rebuilt.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 11:48:03 -04:00
|
|
|
|
|
|
|
|
CPU_Shutdown(ctx);
|
|
|
|
|
if(g_failures == 0)
|
|
|
|
|
std::printf("ALL CHECKS PASSED\n");
|
|
|
|
|
else
|
|
|
|
|
std::printf("%d CHECK(S) FAILED\n", g_failures);
|
|
|
|
|
return (g_failures == 0) ? 0 : 1;
|
|
|
|
|
}
|