forked from animatedread/Warrior_EA
THE OPTIMIZER ("0.1% an hour per agent", 0 of 39 passes in 78 min,
12 agents): the tester fires OnTimer on SIMULATED time, so the live
chart's 500ms EventSetMillisecondTimer over a 2016-2026 pass is ~600
MILLION OnTimer calls - each walking 4x PollTraining, the vote
readout's string build, the overlay advance and the deployed census.
None of it serves an inference-only pass: training never runs, per-bar
inference is driven by OnTickHandler off the tick stream, the risk
budget re-checks in OnTick, and there is no chart to keep fresh.
StepSetTimer now arms EventSetTimer(3600) in tester/optimizer/forward
(~2,600 calls per pass) and keeps the 500ms timer for live charts.
Plus a TESTER PASS SELF-PROFILE: per-tick buckets (pre / Expert.OnTick
/ journal) and the timer total, printed once at the pass's OnDeinit -
so if a pass is still slow it names its own consumer instead of being
diagnosed from outside.
OFFLOAD (operator: "as much calculation as possible to DLL/OpenCL"):
batch norm was the ONE stage still host-side on the DLL tier - the
device path was OpenCL-only, so every sample crossed the bus twice per
BN layer and normalized in interpreted MQL5 (and every model runs
batchnorm ON). Four new exports mirror AI\Network.cl's BatchNorm*
kernels 1:1 in DOUBLE precision (closer to the host reference than
the float OpenCL kernels): forward with running stats + frozen flag,
hidden gradient with the clamp derivative, gamma/beta accumulate, and
the batch-mean apply (no weight decay, moments-before-skip ordering,
sqrt-stored v). BnDeviceEligible/EnsureBnDeviceBuffers/all four
Dispatch* now route by backend; the EXISTING in-situ self-checks
(host-vs-device on the first real sample, latch-off + host fallback on
mismatch) verify the DLL kernels exactly as they verified OpenCL ones.
batch_accum_check regression: ALL CHECKS PASSED on the rebuilt DLL.
Same deployment coupling as bd46374: the .ex5 imports the new exports
- copy DirectML\WarriorCPU.dll into MQL5\Libraries (terminal closed)
together with the new .ex5, and re-copy it to the tester agents (or
just run DirectML\build_cpu.bat once with everything closed - it
deploys to every discovered Libraries folder).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
155 lines
9.9 KiB
C++
155 lines
9.9 KiB
C++
//+------------------------------------------------------------------+
|
|
//| Warrior_EA |
|
|
//| Multithreaded CPU compute fallback - used when neither |
|
|
//| OpenCL nor the D3D12/DirectML tier are available (e.g. |
|
|
//| a VM with no GPU passthrough). Same buffer-handle model |
|
|
//| as WarriorDML.dll but full double precision throughout |
|
|
//| (no GPU float roundtrip) and work is spread across a |
|
|
//| configurable pool of worker threads instead of a device. |
|
|
//+------------------------------------------------------------------+
|
|
// Flat C ABI so MQL5 can #import this DLL directly. Function names and
|
|
// argument order/semantics mirror WarriorDML.h / AI\Network.cl 1:1 so the
|
|
// two backends are interchangeable behind CDirectMLMy in AI\Network.mqh.
|
|
#pragma once
|
|
|
|
#define WARRIORCPU_API extern "C" __declspec(dllexport)
|
|
|
|
// Every exported function (besides CPU_Init/CPU_GetHardwareConcurrency) takes a
|
|
// CpuHandle as its first argument - an opaque pointer to a heap-allocated,
|
|
// self-contained context (its own thread pool, buffer table and mutex) that
|
|
// CPU_Init() allocates and the caller (CDirectMLMy on the MQL5 side) is
|
|
// responsible for remembering and passing back on every subsequent call, then
|
|
// releasing via CPU_Shutdown(). No state is shared between contexts and this
|
|
// DLL keeps no global/static mutable state of its own, so it can be loaded any
|
|
// number of times and driven by any number of instances/threads in parallel -
|
|
// a fault or a wedged call against one context can never poison another
|
|
// context's calls, unlike a process-wide singleton would.
|
|
// A plain `long long` (not C++ `long`, which is only 32 bits on Windows) so it
|
|
// round-trips exactly through MQL5's 64-bit `long` on the #import side.
|
|
typedef long long CpuHandle;
|
|
|
|
// Lifecycle. CPU_Init(threads) always succeeds (returns a non-zero handle)
|
|
// since it needs no hardware - threads<=0 means "use
|
|
// std::thread::hardware_concurrency()". Returns 0 on failure.
|
|
WARRIORCPU_API CpuHandle __stdcall CPU_Init(int threads);
|
|
WARRIORCPU_API void __stdcall CPU_Shutdown(CpuHandle ctx);
|
|
WARRIORCPU_API int __stdcall CPU_GetLastError(CpuHandle ctx);
|
|
WARRIORCPU_API int __stdcall CPU_GetThreadCount(CpuHandle ctx);
|
|
// Stateless: true std::thread::hardware_concurrency(), independent of any
|
|
// context's pool size. Use this (not a CPU_Init(0)/GetThreadCount()/Shutdown()
|
|
// probe) to size a CPU_Init() request.
|
|
WARRIORCPU_API int __stdcall CPU_GetHardwareConcurrency();
|
|
|
|
// Buffer management. Handles are small non-negative integers scoped to ctx;
|
|
// -1 means failure. Buffers are plain double vectors owned by ctx; Write/Read
|
|
// just memcpy in/out, there is no upload/download step on CPU.
|
|
WARRIORCPU_API int __stdcall CPU_BufferCreate(CpuHandle ctx, int elementCount);
|
|
WARRIORCPU_API int __stdcall CPU_BufferWrite(CpuHandle ctx, int handle, const double *data, int count);
|
|
WARRIORCPU_API int __stdcall CPU_BufferRead(CpuHandle ctx, int handle, double *data, int count);
|
|
WARRIORCPU_API void __stdcall CPU_BufferFree(CpuHandle ctx, int handle);
|
|
|
|
// Compute kernels - argument order/semantics mirror AI\Network.cl and
|
|
// WarriorDML.h 1:1. activation: 0 = TANH, 1 = SIGMOID, 2 = PReLU (conv only).
|
|
WARRIORCPU_API int __stdcall CPU_FeedForward(CpuHandle ctx, int wHandle, int iHandle, int oHandle,
|
|
int inputs, int activation);
|
|
|
|
WARRIORCPU_API int __stdcall CPU_CalcOutputGradient(CpuHandle ctx, int tHandle, int oHandle, int igHandle,
|
|
int activation, int count);
|
|
|
|
WARRIORCPU_API int __stdcall CPU_CalcHiddenGradient(CpuHandle ctx, int wHandle, int gHandle, int oHandle, int igHandle,
|
|
int outputs, int activation, int count);
|
|
|
|
WARRIORCPU_API int __stdcall CPU_UpdateWeightsMomentum(CpuHandle ctx, int wHandle, int gHandle, int iHandle, int dwHandle,
|
|
int inputs, double learningRate, double momentumRate, int neurons, int optimizer);
|
|
|
|
WARRIORCPU_API int __stdcall CPU_UpdateWeightsAdam(CpuHandle ctx, int wHandle, int gHandle, int iHandle,
|
|
int mHandle, int vHandle,
|
|
int inputs, double lt, double b1, double b2, int neurons);
|
|
|
|
// Mini-batch gradient accumulation (2026-08-09 audit, F4) - mirrors AI\Network.cl's
|
|
// AccumulateWeightGrad / AccumulateWeightGradConv. These ADD one sample's per-weight gradient into an
|
|
// accumulator buffer; the optimizer step on the batch mean is CPU_ApplyAccumAdam /
|
|
// CPU_ApplyAccumMomentum below (2026-08-25 - the original design applied it host-side in MQL5, which
|
|
// profiled at ~10x the whole forward pass and dominated every training era on the CPU tier).
|
|
WARRIORCPU_API int __stdcall CPU_AccumulateWeightGrad(CpuHandle ctx, int accHandle, int gHandle, int iHandle,
|
|
int inputs, int neurons);
|
|
|
|
WARRIORCPU_API int __stdcall CPU_AccumulateWeightGradConv(CpuHandle ctx, int accHandle, int gHandle, int iHandle,
|
|
int inputs, int windowIn, int windowOut, int step);
|
|
|
|
WARRIORCPU_API int __stdcall CPU_AccumulateBufferInto(CpuHandle ctx, int dstHandle, int srcHandle, int count);
|
|
|
|
// Mini-batch APPLY (2026-08-25): one optimizer step on the batch mean, element-wise over `total`
|
|
// weights, then zero the accumulator. Same math as CPU_UpdateWeightsAdam / the momentum kernel, so
|
|
// batch size 1 reproduces the unbatched path exactly. Generic over any flat weight block (dense,
|
|
// conv, LSTM, batch norm) - the caller passes the block's buffers, no shape needed.
|
|
WARRIORCPU_API int __stdcall CPU_ApplyAccumAdam(CpuHandle ctx, int wHandle, int accHandle, int mHandle, int vHandle,
|
|
int total, double scale, double lt, double b1, double b2);
|
|
|
|
WARRIORCPU_API int __stdcall CPU_ApplyAccumMomentum(CpuHandle ctx, int wHandle, int accHandle, int dwHandle,
|
|
int total, double scale, double learningRate, double momentum);
|
|
|
|
// Batch norm (2026-08-25) - mirrors AI\Network.cl's BatchNorm* kernels in double precision. Until
|
|
// these existed the BN layers were the one stage that ran host-side on the DLL tier, with two bus
|
|
// crossings per layer per sample. The options layout is AI\Network.mqh's BN_OPT_* stride-9 record.
|
|
WARRIORCPU_API int __stdcall CPU_BatchNormForward(CpuHandle ctx, int iHandle, int oHandle, int optHandle,
|
|
double w, int frozen, int count);
|
|
|
|
WARRIORCPU_API int __stdcall CPU_BatchNormHiddenGrad(CpuHandle ctx, int gHandle, int prevOHandle, int prevGHandle,
|
|
int optHandle, int activation, int count);
|
|
|
|
WARRIORCPU_API int __stdcall CPU_BatchNormAccumGammaBeta(CpuHandle ctx, int gHandle, int optHandle, int accHandle, int count);
|
|
|
|
WARRIORCPU_API int __stdcall CPU_BatchNormApplyGammaBeta(CpuHandle ctx, int optHandle, int accHandle,
|
|
double scale, double lt, double b1, double b2, double lr, double momentum, int optimizer, int count);
|
|
|
|
WARRIORCPU_API int __stdcall CPU_FeedForwardConv(CpuHandle ctx, int wHandle, int iHandle, int oHandle,
|
|
int inputs, int step, int windowIn, int windowOut, int activation, int positions);
|
|
|
|
WARRIORCPU_API int __stdcall CPU_CalcHiddenGradientConv(CpuHandle ctx, int wHandle, int gHandle, int oHandle, int igHandle,
|
|
int outputs, int step, int windowIn, int windowOut, int activation, int inputCount);
|
|
|
|
WARRIORCPU_API int __stdcall CPU_UpdateWeightsConvMomentum(CpuHandle ctx, int wHandle, int gHandle, int iHandle, int dwHandle,
|
|
int inputs, double learningRate, double momentumRate, int windowIn, int windowOut, int step, int optimizer);
|
|
|
|
WARRIORCPU_API int __stdcall CPU_UpdateWeightsConvAdam(CpuHandle ctx, int wHandle, int gHandle, int iHandle,
|
|
int mHandle, int vHandle,
|
|
int inputs, double lt, double b1, double b2, int windowIn, int windowOut, int step);
|
|
|
|
WARRIORCPU_API int __stdcall CPU_LSTMGates(CpuHandle ctx, int wHandle, int hiddenPrevHandle, int inputsHandle,
|
|
int concatenatedHandle, int hiddenSize, int inputSize);
|
|
|
|
WARRIORCPU_API int __stdcall CPU_LSTMState(CpuHandle ctx, int concatenatedHandle, int memoryHandle, int hiddenPrevHandle,
|
|
int hiddenCacheHandle, int outputHandle, int hiddenSize);
|
|
|
|
WARRIORCPU_API int __stdcall CPU_LSTMGateGradient(CpuHandle ctx, int gradientHandle, int memoryHandle, int concatenatedHandle,
|
|
int concatenatedGradientHandle, int hiddenSize);
|
|
|
|
WARRIORCPU_API int __stdcall CPU_LSTMWeightsGradient(CpuHandle ctx, int concatenatedGradientHandle, int hiddenCacheHandle,
|
|
int inputsHandle, int weightsGradientHandle, int hiddenSize, int inputSize);
|
|
|
|
WARRIORCPU_API int __stdcall CPU_LSTMInputsGradient(CpuHandle ctx, int concatenatedGradientHandle, int wHandle,
|
|
int inputsGradientHandle, int hiddenSize, int inputSize);
|
|
|
|
// Fused unrolled sequence LSTM - see the block comment above the definitions in WarriorCPU.cpp.
|
|
// These replace the per-step entry points above for sequence models: the per-step ones treat the whole
|
|
// input as ONE timestep, and CPU_LSTMGateGradient cannot accept dc from the following step, so real
|
|
// backpropagation-through-time cannot be assembled from them.
|
|
WARRIORCPU_API int __stdcall CPU_LSTMSeqForward(CpuHandle ctx, int wHandle, int inputsHandle,
|
|
int cacheGatesHandle, int cacheCellHandle, int cacheHiddenHandle, int outputHandle,
|
|
int hiddenSize, int stepInputs, int steps);
|
|
|
|
WARRIORCPU_API int __stdcall CPU_LSTMSeqBackward(CpuHandle ctx, int wHandle, int inputsHandle,
|
|
int cacheGatesHandle, int cacheCellHandle, int cacheHiddenHandle, int outGradientHandle,
|
|
int weightsGradientHandle, int inputsGradientHandle, int hiddenSize, int stepInputs, int steps);
|
|
|
|
WARRIORCPU_API int __stdcall CPU_LSTMUpdateWeightsAdam(CpuHandle ctx, int wHandle, int weightsGradientHandle,
|
|
int mHandle, int vHandle, double l, double b1, double b2, int total);
|
|
|
|
WARRIORCPU_API int __stdcall CPU_LSTMUpdateWeightsMomentum(CpuHandle ctx, int wHandle, int weightsGradientHandle,
|
|
int dwHandle, double learningRate, double momentumRate, int total, int optimizer);
|
|
|
|
WARRIORCPU_API int __stdcall CPU_FeedForwardProof(CpuHandle ctx, int iHandle, int oHandle, int inputs, int window, int step, int outputs);
|
|
|
|
WARRIORCPU_API int __stdcall CPU_CalcInputGradientProof(CpuHandle ctx, int iHandle, int gHandle, int oHandle, int igHandle,
|
|
int outputs, int window, int step, int inputs);
|