Warrior_EA/AI/NeuronOCLConvPool.mqh
AnimateDread ea9d86b3ee fix: dense backprop read the weight matrix transposed - on every backend
CaclHiddenGradient computed this layer's gradient as matrix_w[(outputs+1)*i + k]
against a buffer whose actual layout (one row per NEXT-layer neuron, stride
inputs+1) makes the correct read matrix_w[k*(inputs+1) + i]: the transpose for
square layers, and for the non-square boundaries this EA actually builds
(tapered stacks, the 3-neuron head) a mis-strided walk that ran past the buffer
end - garbage on OpenCL, zeroed reads on the CPU DLL, so the tiers did not even
agree with each other. Every gradient crossing a dense boundary on its way down
- the entire learning signal reaching the BN/conv/LSTM front ends - passed
through a fixed wrong matrix: feedback-alignment dynamics, not backprop, which
is why nets still "learned something" and this survived. The book reference
(NeuroNet_DNG) fixed this in a later article version; our kernel descended from
the earlier one. Confounds every model-based negative verdict to date.

Also in this commit, same root cause family:
- per-sample UpdateWeightsAdam (OpenCL): input for slot group j was read at
  matrix_i[j] instead of matrix_i[j*4] (corrupted outer product past group 0),
  and dispatch dim 1 sized on ceil(inputs/4) left the bias column unreachable
  whenever inputs%4==0 - dense biases never trained on OpenCL. Rewritten as a
  lane-guarded scalar loop keeping our Adam conventions (sqrt-stored v,
  decoupled decay, both clamps, no sign gate). The batched accum path never had
  either bug; this kernel is what SetBatchSize(1) runs - including online
  continual learning on client machines, where OpenCL is the only tier.
- conv backward passed raw (int)Activation() where the kernels expect
  NativeActivationCode(): NONE took the tanh branch (clamping a BN layer's
  unbounded z-scores), TANH took sigmoid, PRELU took none. Dormant only because
  the conv sits at layer 1 today.
- hidden-gradient dispatch over Neurons()+1 dropped to Neurons(): biases get no
  backprop gradient and the extra work-item only ever read past matrix_o.

All three backends (Network.cl, WarriorCPU.cpp, WarriorDML.cpp HLSL) changed in
lockstep; DML gained an `inputs` constant to derive the row stride. New
dense_backprop_check.cpp proves the CPU kernel is central-finite-difference
consistent with the real forward kernel on 8x8, 64x3, 33x64, 5x3 (max diff
3e-9) and that all three activation branches match transcription. All 16 checks
pass. Offline math check only - the in-situ proof remains the per-layer dW/W
report on a real era.

FORCES FULL RETRAIN. Both DLLs rebuilt and redeployed to MQL5\Libraries.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-11 18:06:09 -04:00

832 lines
38 KiB
MQL5

//+------------------------------------------------------------------+
//| NeuronOCLConvPool.mqh |
//| AnimateDread |
//| https://www.mql5.com |
//+------------------------------------------------------------------+
//| CNeuronConvOCL/CNeuronPoolOCL - the GPU-accelerated (OpenCL + |
//| DirectML) convolution and max-pooling layers. Both derive from |
//| CNeuronBaseOCL (AI\Network.mqh, must already be declared) and use|
//| CBufferDouble (AI\BufferDouble.mqh). Included from Network.mqh at|
//| the exact point these classes used to sit (right after |
//| CNeuronBaseOCL's own declaration, before CNeuronLSTMOCL), so |
//| ordering matches the original file. Extracted verbatim (SOLID |
//| cleanup) - no logic changes. |
//+------------------------------------------------------------------+
#include "BufferDouble.mqh"
//| GPU-accelerated convolution layer (OpenCL + DirectML). Ported |
//| from the NeuroNet_DNG reference library's CNeuronConvOCL/ |
//| CNeuronProofOCL, adapted to this project's double-precision |
//| CBufferDouble/COpenCLMy/CDirectMLMy conventions. A single |
//| (window+1)*window_out weight block is shared across every |
//| sliding position - unlike CNeuronBaseOCL, where every output has |
//| its own private weight vector. |
//+------------------------------------------------------------------+
class CNeuronConvOCL : public CNeuronBaseOCL
{
protected:
uint iWindow;
uint iStep;
uint iWindowOut;
CBufferDouble *WeightsConv;
CBufferDouble *DeltaWeightsConv;
CBufferDouble *FirstMomentumConv;
CBufferDouble *SecondMomentumConv;
//--- Mini-batch accumulator for the CONVOLUTION KERNEL block. Distinct from the base class's
//--- GradAccum, which shadows the base Weights - i.e. the outgoing dense matrix that the layer ABOVE
//--- accumulates into. A conv neuron owns both tensors, so it needs both accumulators; sharing one
//--- slot would have the two layers writing each other's gradients.
CBufferDouble *GradAccumConv;
//---
virtual bool feedForward(CNeuronBaseOCL *NeuronOCL);
virtual bool feedForwardCPU(CNeuronBaseOCL *NeuronOCL); // pure-MQL5 mirror of FeedForwardConv
virtual bool updateInputWeights(CNeuronBaseOCL *NeuronOCL);
virtual bool accumulateInputWeightGrads(CNeuronBaseOCL *NeuronOCL);
public:
CNeuronConvOCL(void) : iWindow(1), iStep(1), iWindowOut(1)
{
WeightsConv = NULL;
DeltaWeightsConv = NULL;
FirstMomentumConv = NULL;
SecondMomentumConv = NULL;
GradAccumConv = NULL;
}
~CNeuronConvOCL(void);
virtual bool Init(uint numOutputs, uint myIndex, COpenCLMy *open_cl, uint window_in, uint step, uint window_out, uint units_count, ENUM_OPTIMIZATION optimization_type);
virtual bool Init(uint numOutputs, uint myIndex, CDirectMLMy *direct_ml, uint window_in, uint step, uint window_out, uint units_count, ENUM_OPTIMIZATION optimization_type);
virtual bool calcInputGradients(CNeuronBaseOCL *NeuronOCL);
virtual bool Save(int const file_handle);
virtual bool Load(int const file_handle);
virtual int Type(void) const { return defNeuronConvOCL; }
// Shape accessors. A .nnw persists the window/step it was BUILT with, so a model saved by an older
// build keeps that shape forever on load - these let EnforceTopologyContract() detect a stale
// receptive field instead of training on silently. See AI\Network.mqh's CNet::FirstConvWindow.
uint Window(void) const { return iWindow; }
uint Step(void) const { return iStep; }
uint WindowOut(void) const { return iWindowOut; }
// See CNeuronBaseOCL::getWeights/setWeights - same pair, targeting WeightsConv instead of the
// base class's Weights, for CNet::BlendWeightsFrom()'s EMA shadow-weight deployment.
virtual int getWeightsConv(double &values[]) { return (CheckPointer(WeightsConv) == POINTER_INVALID ? 0 : WeightsConv.GetData(values)); }
virtual bool setWeightsConv(double &values[])
{
if(CheckPointer(WeightsConv) == POINTER_INVALID)
return false;
if(!WeightsConv.AssignArray(values))
return false;
return WeightsConv.BufferWrite();
}
//--- The conv kernel block keeps its own moment/momentum buffers beside the base class's - see
//--- CNet::ResetOptimizerState.
virtual bool ResetOptimizerState(void)
{
bool ok = CNeuronBaseOCL::ResetOptimizerState();
ok = ZeroOptimizerBuffer(FirstMomentumConv) && ok;
ok = ZeroOptimizerBuffer(SecondMomentumConv) && ok;
ok = ZeroOptimizerBuffer(DeltaWeightsConv) && ok;
return ok;
}
//--- Both accumulators - the base one for the outgoing dense matrix, this class's for the conv
//--- kernel. See GradAccumConv's declaration for why they cannot share a slot.
virtual bool BeginGradAccum(void)
{
bool ok = CNeuronBaseOCL::BeginGradAccum();
if(CheckPointer(GradAccumConv) != POINTER_INVALID && GradAccumConv.Total() > 0)
ok = ZeroOptimizerBuffer(GradAccumConv) && ok;
return ok;
}
virtual bool ApplyAccumulatedGradients(double scale)
{
//--- Both blocks step on the SAME batch, so t advances once here rather than once per block.
bool ok = ApplyAccumToBlock(Weights, GradAccum, FirstMomentum, SecondMomentum, DeltaWeights, scale);
ok = ApplyAccumToBlock(WeightsConv, GradAccumConv, FirstMomentumConv, SecondMomentumConv,
DeltaWeightsConv, scale) && ok;
if(optimization == ADAM)
t++;
return ok;
}
};
//+------------------------------------------------------------------+
//| |
//+------------------------------------------------------------------+
CNeuronConvOCL::~CNeuronConvOCL(void)
{
if(CheckPointer(WeightsConv) != POINTER_INVALID)
delete WeightsConv;
if(CheckPointer(DeltaWeightsConv) != POINTER_INVALID)
delete DeltaWeightsConv;
if(CheckPointer(FirstMomentumConv) != POINTER_INVALID)
delete FirstMomentumConv;
if(CheckPointer(SecondMomentumConv) != POINTER_INVALID)
delete SecondMomentumConv;
if(CheckPointer(GradAccumConv) != POINTER_INVALID)
delete GradAccumConv;
}
//+------------------------------------------------------------------+
//| |
//+------------------------------------------------------------------+
bool CNeuronConvOCL::Init(uint numOutputs, uint myIndex, COpenCLMy *open_cl, uint window_in, uint step, uint window_out, uint units_count, ENUM_OPTIMIZATION optimization_type)
{
if(window_out <= 0)
return false;
if(!CNeuronBaseOCL::Init(numOutputs, myIndex, open_cl, units_count * window_out, optimization_type))
return false;
//---
iWindow = window_in;
iStep = step;
iWindowOut = (uint)fmax(window_out, 1);
//---
int count = (int)((iWindow + 1) * iWindowOut);
if(CheckPointer(WeightsConv) == POINTER_INVALID)
{
WeightsConv = new CBufferDouble();
if(CheckPointer(WeightsConv) == POINTER_INVALID)
return false;
}
if(!WeightsConv.Reserve(count))
return false;
// Fan-in-scaled (LeCun-uniform) init - see CNeuronBaseOCL::Init's OpenCL overload for the full
// rationale; fan-in here is the conv window size.
double weighScale = 1.0 / MathSqrt((double)iWindow + 1.0);
for(int i = 0; i < count; i++)
{
double weigh = ((MathRand() + 1) / 32768.0 - 0.5) * 2.0 * weighScale;
if(weigh == 0)
weigh = 0.001;
if(!WeightsConv.Add(weigh))
return false;
}
if(!WeightsConv.BufferCreate(OpenCL))
return false;
//---
if(optimization == SGD)
{
if(CheckPointer(DeltaWeightsConv) == POINTER_INVALID)
{
DeltaWeightsConv = new CBufferDouble();
if(CheckPointer(DeltaWeightsConv) == POINTER_INVALID)
return false;
}
if(!DeltaWeightsConv.BufferInit(count, 0))
return false;
if(!DeltaWeightsConv.BufferCreate(OpenCL))
return false;
}
else
{
if(CheckPointer(FirstMomentumConv) == POINTER_INVALID)
{
FirstMomentumConv = new CBufferDouble();
if(CheckPointer(FirstMomentumConv) == POINTER_INVALID)
return false;
}
if(!FirstMomentumConv.BufferInit(count, 0))
return false;
if(!FirstMomentumConv.BufferCreate(OpenCL))
return false;
//---
if(CheckPointer(SecondMomentumConv) == POINTER_INVALID)
{
SecondMomentumConv = new CBufferDouble();
if(CheckPointer(SecondMomentumConv) == POINTER_INVALID)
return false;
}
if(!SecondMomentumConv.BufferInit(count, 0))
return false;
if(!SecondMomentumConv.BufferCreate(OpenCL))
return false;
}
//---
return true;
}
//+------------------------------------------------------------------+
//| DirectML/D3D12 tier equivalent of Init(COpenCLMy*) above. |
//+------------------------------------------------------------------+
bool CNeuronConvOCL::Init(uint numOutputs, uint myIndex, CDirectMLMy *direct_ml, uint window_in, uint step, uint window_out, uint units_count, ENUM_OPTIMIZATION optimization_type)
{
if(window_out <= 0)
return false;
if(!CNeuronBaseOCL::Init(numOutputs, myIndex, direct_ml, units_count * window_out, optimization_type))
return false;
//---
iWindow = window_in;
iStep = step;
iWindowOut = (uint)fmax(window_out, 1);
//---
int count = (int)((iWindow + 1) * iWindowOut);
if(CheckPointer(WeightsConv) == POINTER_INVALID)
{
WeightsConv = new CBufferDouble();
if(CheckPointer(WeightsConv) == POINTER_INVALID)
return false;
}
if(!WeightsConv.Reserve(count))
return false;
// Fan-in-scaled (LeCun-uniform) init - see the matching OpenCL Init() overload above.
double weighScale = 1.0 / MathSqrt((double)iWindow + 1.0);
for(int i = 0; i < count; i++)
{
double weigh = ((MathRand() + 1) / 32768.0 - 0.5) * 2.0 * weighScale;
if(weigh == 0)
weigh = 0.001;
if(!WeightsConv.Add(weigh))
return false;
}
if(!WeightsConv.BufferCreate(DirectML))
return false;
//---
if(optimization == SGD)
{
if(CheckPointer(DeltaWeightsConv) == POINTER_INVALID)
{
DeltaWeightsConv = new CBufferDouble();
if(CheckPointer(DeltaWeightsConv) == POINTER_INVALID)
return false;
}
if(!DeltaWeightsConv.BufferInit(count, 0))
return false;
if(!DeltaWeightsConv.BufferCreate(DirectML))
return false;
}
else
{
if(CheckPointer(FirstMomentumConv) == POINTER_INVALID)
{
FirstMomentumConv = new CBufferDouble();
if(CheckPointer(FirstMomentumConv) == POINTER_INVALID)
return false;
}
if(!FirstMomentumConv.BufferInit(count, 0))
return false;
if(!FirstMomentumConv.BufferCreate(DirectML))
return false;
//---
if(CheckPointer(SecondMomentumConv) == POINTER_INVALID)
{
SecondMomentumConv = new CBufferDouble();
if(CheckPointer(SecondMomentumConv) == POINTER_INVALID)
return false;
}
if(!SecondMomentumConv.BufferInit(count, 0))
return false;
if(!SecondMomentumConv.BufferCreate(DirectML))
return false;
}
//---
return true;
}
//+------------------------------------------------------------------+
//| |
//+------------------------------------------------------------------+
bool CNeuronConvOCL::feedForward(CNeuronBaseOCL *NeuronOCL)
{
if(CheckPointer(NeuronOCL) == POINTER_INVALID)
return false;
int positions = Output.Total() / (int)iWindowOut;
if(CheckPointer(DirectML) != POINTER_INVALID)
{
if(!DirectML.FeedForwardConv(WeightsConv.GetIndex(), NeuronOCL.getOutputIndex(), Output.GetIndex(),
NeuronOCL.Neurons(), (int)iStep, (int)iWindow, (int)iWindowOut, NativeActivationCode(activation), positions))
{
Print(__FUNCTION__ + ": " + DirectML.BackendName() + " FeedForwardConv failed, error " + IntegerToString(DirectML.LastError()));
return false;
}
return Output.BufferRead();
}
if(CheckPointer(OpenCL) == POINTER_INVALID)
return false;
uint global_work_offset[1] = {0};
uint global_work_size[1];
global_work_size[0] = (uint)positions;
OpenCL.SetArgumentBuffer(def_k_FeedForwardConv, def_k_ffc_matrix_w, WeightsConv.GetIndex());
OpenCL.SetArgumentBuffer(def_k_FeedForwardConv, def_k_ffc_matrix_i, NeuronOCL.getOutputIndex());
OpenCL.SetArgumentBuffer(def_k_FeedForwardConv, def_k_ffc_matrix_o, Output.GetIndex());
OpenCL.SetArgument(def_k_FeedForwardConv, def_k_ffc_inputs, NeuronOCL.Neurons());
OpenCL.SetArgument(def_k_FeedForwardConv, def_k_ffc_step, (int)iStep);
OpenCL.SetArgument(def_k_FeedForwardConv, def_k_ffc_window_in, (int)iWindow);
OpenCL.SetArgument(def_k_FeedForwardConv, def_k_ffc_window_out, (int)iWindowOut);
OpenCL.SetArgument(def_k_FeedForwardConv, def_k_ffc_activation, NativeActivationCode(activation));
if(!OpenCL.Execute(def_k_FeedForwardConv, 1, global_work_offset, global_work_size))
{
printf("Error of execution kernel FeedForwardConv: %d", GetLastError());
return false;
}
//--- Output stays GPU-resident; see the note in CNeuronBaseOCL::feedForward().
return true;
}
//+------------------------------------------------------------------+
//| Pure-MQL5 double-precision mirror of Network.cl's FeedForwardConv |
//| kernel (host buffers only) - the CPU inference path. Filters live |
//| in this layer's own WeightsConv, [window_out][window_in+1] with |
//| the +1 bias last; inputs are the previous layer's Output. |
//+------------------------------------------------------------------+
bool CNeuronConvOCL::feedForwardCPU(CNeuronBaseOCL *NeuronOCL)
{
if(CheckPointer(NeuronOCL) == POINTER_INVALID || CheckPointer(Output) == POINTER_INVALID || CheckPointer(WeightsConv) == POINTER_INVALID)
return false;
int inputs = NeuronOCL.Neurons();
int window_in = (int)iWindow;
int step = (int)iStep;
int window_out = (int)iWindowOut;
if(window_out <= 0)
return false;
int positions = Output.Total() / window_out;
int wTotal = WeightsConv.Total();
for(int i = 0; i < positions; i++)
{
int shift_out = window_out * i;
int shift_in = step * i;
for(int out = 0; out < window_out; out++)
{
int shift = (window_in + 1) * out;
if(shift + window_in >= wTotal)
return false;
int stop = (window_in <= (inputs - shift_in)) ? window_in : (inputs - shift_in);
double sum = 0.0;
for(int k = 0; k < stop; k++)
sum += NeuronOCL.OutputHost(shift_in + k) * WeightsConv.At(shift + k);
sum += WeightsConv.At(shift + window_in); // bias
switch(activation)
{
case TANH:
sum = tanh(sum);
break;
case SIGMOID:
sum = 1.0 / (1.0 + exp(-MathMax(-50.0, MathMin(50.0, sum))));
break;
case PRELU:
if(sum < 0.0)
sum *= 0.01;
break;
}
if(!Output.Update(out + shift_out, sum))
return false;
}
}
return true;
}
//+------------------------------------------------------------------+
//| Writes the gradient into NeuronOCL (the EARLIER/input-side layer) |
//| - opposite call direction from the dense calcHiddenGradients, but |
//| matches the reference library's own CNeuronConvOCL convention. |
//+------------------------------------------------------------------+
bool CNeuronConvOCL::calcInputGradients(CNeuronBaseOCL *NeuronOCL)
{
if(CheckPointer(NeuronOCL) == POINTER_INVALID)
return false;
int outputs = Neurons();
if(CheckPointer(DirectML) != POINTER_INVALID)
{
//--- NativeActivationCode, NOT a raw enum cast: the backends number activations differently
//--- from ENUM_ACTIVATION (see NativeActivationCode's declaration comment). The raw cast sent
//--- NONE(0) into the kernels' tanh branch (which clamps and damps a BN layer's unbounded
//--- z-score outputs), TANH(1) into the sigmoid branch, and PRELU(3) past every branch (no
//--- derivative at all). Dormant in current presets only because the conv sits at layer 1
//--- with nothing trainable below it - any deeper conv placement activates it silently.
if(!DirectML.CalcHiddenGradientConv(WeightsConv.GetIndex(), getGradientIndex(), NeuronOCL.getOutputIndex(), NeuronOCL.getGradientIndex(),
outputs, (int)iStep, (int)iWindow, (int)iWindowOut, NativeActivationCode(NeuronOCL.Activation()), NeuronOCL.Neurons()))
{
Print(__FUNCTION__ + ": " + DirectML.BackendName() + " CalcHiddenGradientConv failed, error " + IntegerToString(DirectML.LastError()));
return false;
}
double temp[];
return NeuronOCL.getGradient(temp) > 0;
}
if(CheckPointer(OpenCL) == POINTER_INVALID)
return false;
uint global_work_offset[1] = {0};
uint global_work_size[1];
global_work_size[0] = NeuronOCL.Neurons();
OpenCL.SetArgumentBuffer(def_k_CalcHiddenGradientConv, def_k_chgc_matrix_w, WeightsConv.GetIndex());
OpenCL.SetArgumentBuffer(def_k_CalcHiddenGradientConv, def_k_chgc_matrix_g, getGradientIndex());
OpenCL.SetArgumentBuffer(def_k_CalcHiddenGradientConv, def_k_chgc_matrix_o, NeuronOCL.getOutputIndex());
OpenCL.SetArgumentBuffer(def_k_CalcHiddenGradientConv, def_k_chgc_matrix_ig, NeuronOCL.getGradientIndex());
OpenCL.SetArgument(def_k_CalcHiddenGradientConv, def_k_chgc_outputs, outputs);
OpenCL.SetArgument(def_k_CalcHiddenGradientConv, def_k_chgc_step, (int)iStep);
OpenCL.SetArgument(def_k_CalcHiddenGradientConv, def_k_chgc_window_in, (int)iWindow);
OpenCL.SetArgument(def_k_CalcHiddenGradientConv, def_k_chgc_window_out, (int)iWindowOut);
//--- NativeActivationCode, NOT a raw enum cast - see the DirectML branch's comment above.
OpenCL.SetArgument(def_k_CalcHiddenGradientConv, def_k_chgc_activation, NativeActivationCode(NeuronOCL.Activation()));
if(!OpenCL.Execute(def_k_CalcHiddenGradientConv, 1, global_work_offset, global_work_size))
{
printf("Error of execution kernel CalcHiddenGradientConv: %d", GetLastError());
return false;
}
//--- NeuronOCL's Gradient stays GPU-resident; its own calcHiddenGradients/calcInputGradients
//--- reads it via getGradientIndex(). The old getGradient(temp) call was a discarded-result
//--- sync (GetData() BufferRead()s internally) with no consumer of the read - pure overhead.
return true;
}
//+------------------------------------------------------------------+
//| |
//+------------------------------------------------------------------+
bool CNeuronConvOCL::updateInputWeights(CNeuronBaseOCL *NeuronOCL)
{
if(CheckPointer(NeuronOCL) == POINTER_INVALID)
return false;
int inputs = NeuronOCL.Neurons();
if(CheckPointer(DirectML) != POINTER_INVALID)
{
if(optimization == SGD)
{
if(!DirectML.UpdateWeightsConvMomentum(WeightsConv.GetIndex(), getGradientIndex(), NeuronOCL.getOutputIndex(), DeltaWeightsConv.GetIndex(),
inputs, eta, alpha, (int)iWindow, (int)iWindowOut, (int)iStep, 0))
{
Print(__FUNCTION__ + ": " + DirectML.BackendName() + " UpdateWeightsConvMomentum failed, error " + IntegerToString(DirectML.LastError()));
return false;
}
}
else
{
double lt = eta * sqrt(1 - pow(b2, t)) / (1 - pow(b1, t));
if(!DirectML.UpdateWeightsConvAdam(WeightsConv.GetIndex(), getGradientIndex(), NeuronOCL.getOutputIndex(),
FirstMomentumConv.GetIndex(), SecondMomentumConv.GetIndex(), inputs, lt, b1, b2, (int)iWindow, (int)iWindowOut, (int)iStep))
{
Print(__FUNCTION__ + ": " + DirectML.BackendName() + " UpdateWeightsConvAdam failed, error " + IntegerToString(DirectML.LastError()));
return false;
}
t++;
}
//--- WeightsConv stays DLL-resident; feedForward reads it via GetIndex() (same as the OpenCL
//--- branch below). Save()/BlendWeightsFrom() BufferRead() on demand.
return true;
}
if(CheckPointer(OpenCL) == POINTER_INVALID)
return false;
uint global_work_offset[1] = {0};
uint global_work_size[1];
if(optimization == SGD)
{
global_work_size[0] = WeightsConv.Total();
OpenCL.SetArgumentBuffer(def_k_UpdateWeightsConvMomentum, def_k_uwcm_matrix_w, WeightsConv.GetIndex());
OpenCL.SetArgumentBuffer(def_k_UpdateWeightsConvMomentum, def_k_uwcm_matrix_g, getGradientIndex());
OpenCL.SetArgumentBuffer(def_k_UpdateWeightsConvMomentum, def_k_uwcm_matrix_i, NeuronOCL.getOutputIndex());
OpenCL.SetArgumentBuffer(def_k_UpdateWeightsConvMomentum, def_k_uwcm_matrix_dw, DeltaWeightsConv.GetIndex());
OpenCL.SetArgument(def_k_UpdateWeightsConvMomentum, def_k_uwcm_inputs, inputs);
OpenCL.SetArgument(def_k_UpdateWeightsConvMomentum, def_k_uwcm_learning_rates, (float)eta);
OpenCL.SetArgument(def_k_UpdateWeightsConvMomentum, def_k_uwcm_momentum, (float)alpha);
OpenCL.SetArgument(def_k_UpdateWeightsConvMomentum, def_k_uwcm_window_in, (int)iWindow);
OpenCL.SetArgument(def_k_UpdateWeightsConvMomentum, def_k_uwcm_window_out, (int)iWindowOut);
OpenCL.SetArgument(def_k_UpdateWeightsConvMomentum, def_k_uwcm_step, (int)iStep);
OpenCL.SetArgument(def_k_UpdateWeightsConvMomentum, def_k_uwcm_optimizer, 0);
ResetLastError();
if(!OpenCL.Execute(def_k_UpdateWeightsConvMomentum, 1, global_work_offset, global_work_size))
{
printf("Error of execution kernel UpdateWeightsConvMomentum: %d", GetLastError());
return false;
}
}
else
{
global_work_size[0] = iWindow + 1;
double lt = eta * sqrt(1 - pow(b2, t)) / (1 - pow(b1, t));
OpenCL.SetArgumentBuffer(def_k_UpdateWeightsConvAdam, def_k_uwca_matrix_w, WeightsConv.GetIndex());
OpenCL.SetArgumentBuffer(def_k_UpdateWeightsConvAdam, def_k_uwca_matrix_g, getGradientIndex());
OpenCL.SetArgumentBuffer(def_k_UpdateWeightsConvAdam, def_k_uwca_matrix_i, NeuronOCL.getOutputIndex());
OpenCL.SetArgumentBuffer(def_k_UpdateWeightsConvAdam, def_k_uwca_matrix_m, FirstMomentumConv.GetIndex());
OpenCL.SetArgumentBuffer(def_k_UpdateWeightsConvAdam, def_k_uwca_matrix_v, SecondMomentumConv.GetIndex());
OpenCL.SetArgument(def_k_UpdateWeightsConvAdam, def_k_uwca_inputs, inputs);
OpenCL.SetArgument(def_k_UpdateWeightsConvAdam, def_k_uwca_l, (float)lt);
OpenCL.SetArgument(def_k_UpdateWeightsConvAdam, def_k_uwca_b1, (float)b1);
OpenCL.SetArgument(def_k_UpdateWeightsConvAdam, def_k_uwca_b2, (float)b2);
OpenCL.SetArgument(def_k_UpdateWeightsConvAdam, def_k_uwca_window_in, (int)iWindow);
OpenCL.SetArgument(def_k_UpdateWeightsConvAdam, def_k_uwca_window_out, (int)iWindowOut);
OpenCL.SetArgument(def_k_UpdateWeightsConvAdam, def_k_uwca_step, (int)iStep);
ResetLastError();
if(!OpenCL.Execute(def_k_UpdateWeightsConvAdam, 1, global_work_offset, global_work_size))
{
printf("Error of execution kernel UpdateWeightsConvAdam: %d", GetLastError());
return false;
}
t++;
}
//--- WeightsConv stays GPU-resident; see the note in CNeuronBaseOCL::updateInputWeights().
return true;
}
//+------------------------------------------------------------------+
//| MINI-BATCH ACCUMULATE (conv kernel block). Batched counterpart of |
//| updateInputWeights above - identical dispatch, identical operands,|
//| but it only ADDS into GradAccumConv. The base class's own |
//| accumulator (the outgoing dense matrix) is filled separately by |
//| the layer ABOVE this one calling its accumulate on this neuron. |
//+------------------------------------------------------------------+
bool CNeuronConvOCL::accumulateInputWeightGrads(CNeuronBaseOCL *NeuronOCL)
{
if(CheckPointer(NeuronOCL) == POINTER_INVALID || CheckPointer(WeightsConv) == POINTER_INVALID)
return false;
if(!EnsureGradAccumFor(GradAccumConv, WeightsConv))
return false;
int inputs = NeuronOCL.Neurons();
if(CheckPointer(DirectML) != POINTER_INVALID)
{
if(!DirectML.AccumulateWeightGradConv(GradAccumConv.GetIndex(), getGradientIndex(), NeuronOCL.getOutputIndex(),
inputs, (int)iWindow, (int)iWindowOut, (int)iStep))
{
Print(__FUNCTION__ + ": " + DirectML.BackendName() + " AccumulateWeightGradConv failed, error " + IntegerToString(DirectML.LastError()));
return false;
}
return true;
}
if(CheckPointer(OpenCL) == POINTER_INVALID)
return false;
uint global_work_offset[1] = {0};
uint global_work_size[1];
//--- Flat over the whole kernel block, matching AccumulateWeightGradConv's indexing (and
//--- UpdateWeightsConvMomentum's), NOT the Adam kernel's (window_in+1) shape.
global_work_size[0] = WeightsConv.Total();
if(!OpenCL.SetArgumentBuffer(def_k_AccumulateWeightGradConv, def_k_awgc_matrix_acc, GradAccumConv.GetIndex()))
return false;
if(!OpenCL.SetArgumentBuffer(def_k_AccumulateWeightGradConv, def_k_awgc_matrix_g, getGradientIndex()))
return false;
if(!OpenCL.SetArgumentBuffer(def_k_AccumulateWeightGradConv, def_k_awgc_matrix_i, NeuronOCL.getOutputIndex()))
return false;
if(!OpenCL.SetArgument(def_k_AccumulateWeightGradConv, def_k_awgc_inputs, inputs))
return false;
if(!OpenCL.SetArgument(def_k_AccumulateWeightGradConv, def_k_awgc_window_in, (int)iWindow))
return false;
if(!OpenCL.SetArgument(def_k_AccumulateWeightGradConv, def_k_awgc_window_out, (int)iWindowOut))
return false;
if(!OpenCL.SetArgument(def_k_AccumulateWeightGradConv, def_k_awgc_step, (int)iStep))
return false;
ResetLastError();
if(!OpenCL.Execute(def_k_AccumulateWeightGradConv, 1, global_work_offset, global_work_size))
{
printf("Error of execution kernel AccumulateWeightGradConv: %d", GetLastError());
return false;
}
return true;
}
//+------------------------------------------------------------------+
//| |
//+------------------------------------------------------------------+
bool CNeuronConvOCL::Save(const int file_handle)
{
if(!CNeuronBaseOCL::Save(file_handle))
return false;
if(FileWriteInteger(file_handle, (int)iWindow, INT_VALUE) < INT_VALUE)
return false;
if(FileWriteInteger(file_handle, (int)iStep, INT_VALUE) < INT_VALUE)
return false;
if(FileWriteInteger(file_handle, (int)iWindowOut, INT_VALUE) < INT_VALUE)
return false;
if(CheckPointer(WeightsConv) == POINTER_INVALID || !WeightsConv.BufferRead() || !WeightsConv.Save(file_handle))
return false;
if(optimization == SGD)
{
if(CheckPointer(DeltaWeightsConv) == POINTER_INVALID || !DeltaWeightsConv.BufferRead() || !DeltaWeightsConv.Save(file_handle))
return false;
}
else
{
if(CheckPointer(FirstMomentumConv) == POINTER_INVALID || !FirstMomentumConv.BufferRead() || !FirstMomentumConv.Save(file_handle))
return false;
if(CheckPointer(SecondMomentumConv) == POINTER_INVALID || !SecondMomentumConv.BufferRead() || !SecondMomentumConv.Save(file_handle))
return false;
}
//---
return true;
}
//+------------------------------------------------------------------+
//| |
//+------------------------------------------------------------------+
bool CNeuronConvOCL::Load(const int file_handle)
{
if(!CNeuronBaseOCL::Load(file_handle))
return false;
iWindow = (uint)FileReadInteger(file_handle, INT_VALUE);
iStep = (uint)FileReadInteger(file_handle, INT_VALUE);
iWindowOut = (uint)FileReadInteger(file_handle, INT_VALUE);
//---
if(CheckPointer(WeightsConv) == POINTER_INVALID)
{
WeightsConv = new CBufferDouble();
if(CheckPointer(WeightsConv) == POINTER_INVALID)
return false;
}
if(WeightsConv.GetIndex() >= 0)
WeightsConv.BufferFree();
if(!WeightsConv.Load(file_handle))
return false;
if(!BackendBufferCreate(WeightsConv))
return false;
//---
if(optimization == SGD)
{
if(CheckPointer(DeltaWeightsConv) == POINTER_INVALID)
{
DeltaWeightsConv = new CBufferDouble();
if(CheckPointer(DeltaWeightsConv) == POINTER_INVALID)
return false;
}
if(DeltaWeightsConv.GetIndex() >= 0)
DeltaWeightsConv.BufferFree();
if(!DeltaWeightsConv.Load(file_handle))
return false;
if(!BackendBufferCreate(DeltaWeightsConv))
return false;
}
else
{
if(CheckPointer(FirstMomentumConv) == POINTER_INVALID)
{
FirstMomentumConv = new CBufferDouble();
if(CheckPointer(FirstMomentumConv) == POINTER_INVALID)
return false;
}
if(FirstMomentumConv.GetIndex() >= 0)
FirstMomentumConv.BufferFree();
if(!FirstMomentumConv.Load(file_handle))
return false;
if(!BackendBufferCreate(FirstMomentumConv))
return false;
//---
if(CheckPointer(SecondMomentumConv) == POINTER_INVALID)
{
SecondMomentumConv = new CBufferDouble();
if(CheckPointer(SecondMomentumConv) == POINTER_INVALID)
return false;
}
if(SecondMomentumConv.GetIndex() >= 0)
SecondMomentumConv.BufferFree();
if(!SecondMomentumConv.Load(file_handle))
return false;
if(!BackendBufferCreate(SecondMomentumConv))
return false;
}
//---
return true;
}
//+------------------------------------------------------------------+
//| GPU-accelerated max-pooling layer (OpenCL + DirectML). No weights,|
//| so no updateInputWeights work - just a sliding max. Ported from |
//| the NeuroNet_DNG reference's CNeuronProofOCL kernels. |
//+------------------------------------------------------------------+
class CNeuronPoolOCL : public CNeuronBaseOCL
{
protected:
uint iWindow;
uint iStep;
//---
virtual bool feedForward(CNeuronBaseOCL *NeuronOCL);
virtual bool feedForwardCPU(CNeuronBaseOCL *NeuronOCL); // pure-MQL5 mirror of FeedForwardProof
virtual bool updateInputWeights(CNeuronBaseOCL *NeuronOCL) { return true; }
public:
CNeuronPoolOCL(void) : iWindow(2), iStep(1) {}
~CNeuronPoolOCL(void) {}
virtual bool Init(uint numOutputs, uint myIndex, COpenCLMy *open_cl, uint window, uint step, uint units_count, ENUM_OPTIMIZATION optimization_type);
virtual bool Init(uint numOutputs, uint myIndex, CDirectMLMy *direct_ml, uint window, uint step, uint units_count, ENUM_OPTIMIZATION optimization_type);
virtual bool calcInputGradients(CNeuronBaseOCL *NeuronOCL);
virtual bool Save(int const file_handle);
virtual bool Load(int const file_handle);
virtual int Type(void) const { return defNeuronPoolOCL; }
};
//+------------------------------------------------------------------+
//| |
//+------------------------------------------------------------------+
bool CNeuronPoolOCL::Init(uint numOutputs, uint myIndex, COpenCLMy *open_cl, uint window, uint step, uint units_count, ENUM_OPTIMIZATION optimization_type)
{
if(!CNeuronBaseOCL::Init(numOutputs, myIndex, open_cl, units_count, optimization_type))
return false;
iWindow = window;
iStep = step;
return true;
}
//+------------------------------------------------------------------+
//| |
//+------------------------------------------------------------------+
bool CNeuronPoolOCL::Init(uint numOutputs, uint myIndex, CDirectMLMy *direct_ml, uint window, uint step, uint units_count, ENUM_OPTIMIZATION optimization_type)
{
if(!CNeuronBaseOCL::Init(numOutputs, myIndex, direct_ml, units_count, optimization_type))
return false;
iWindow = window;
iStep = step;
return true;
}
//+------------------------------------------------------------------+
//| |
//+------------------------------------------------------------------+
bool CNeuronPoolOCL::feedForward(CNeuronBaseOCL *NeuronOCL)
{
if(CheckPointer(NeuronOCL) == POINTER_INVALID)
return false;
int outputs = Output.Total();
if(CheckPointer(DirectML) != POINTER_INVALID)
{
if(!DirectML.FeedForwardProof(NeuronOCL.getOutputIndex(), Output.GetIndex(), NeuronOCL.Neurons(), (int)iWindow, (int)iStep, outputs))
{
Print(__FUNCTION__ + ": " + DirectML.BackendName() + " FeedForwardProof failed, error " + IntegerToString(DirectML.LastError()));
return false;
}
return Output.BufferRead();
}
if(CheckPointer(OpenCL) == POINTER_INVALID)
return false;
uint offset1[1] = {0};
uint size1[1] = {(uint)outputs};
OpenCL.SetArgumentBuffer(def_k_FeedForwardProof, def_k_ffp_matrix_i, NeuronOCL.getOutputIndex());
OpenCL.SetArgumentBuffer(def_k_FeedForwardProof, def_k_ffp_matrix_o, Output.GetIndex());
OpenCL.SetArgument(def_k_FeedForwardProof, def_k_ffp_inputs, NeuronOCL.Neurons());
OpenCL.SetArgument(def_k_FeedForwardProof, def_k_ffp_window, (int)iWindow);
OpenCL.SetArgument(def_k_FeedForwardProof, def_k_ffp_step, (int)iStep);
if(!OpenCL.Execute(def_k_FeedForwardProof, 1, offset1, size1))
{
printf("Error of execution kernel FeedForwardProof: %d", GetLastError());
return false;
}
//--- Output stays GPU-resident; see the note in CNeuronBaseOCL::feedForward().
return true;
}
//+------------------------------------------------------------------+
//| Pure-MQL5 mirror of Network.cl's FeedForwardProof kernel (sliding |
//| max-pool over the previous layer's Output). Host buffers only. |
//+------------------------------------------------------------------+
bool CNeuronPoolOCL::feedForwardCPU(CNeuronBaseOCL *NeuronOCL)
{
if(CheckPointer(NeuronOCL) == POINTER_INVALID || CheckPointer(Output) == POINTER_INVALID)
return false;
int outputs = Output.Total();
int inputs = NeuronOCL.Neurons();
int window = (int)iWindow;
int step = (int)iStep;
for(int i = 0; i < outputs; i++)
{
int pos = i * step;
if(pos >= inputs)
return false;
double result = NeuronOCL.OutputHost(pos);
for(int k = 1; k < window; k++)
{
int shift = k + pos;
if(shift >= inputs)
break;
result = MathMax(result, NeuronOCL.OutputHost(shift));
}
if(!Output.Update(i, result))
return false;
}
return true;
}
//+------------------------------------------------------------------+
//| Inverted-call convention, same as Conv/LSTM's calcInputGradients. |
//+------------------------------------------------------------------+
bool CNeuronPoolOCL::calcInputGradients(CNeuronBaseOCL *NeuronOCL)
{
if(CheckPointer(NeuronOCL) == POINTER_INVALID)
return false;
int outputs = Neurons();
int inputs = NeuronOCL.Neurons();
if(CheckPointer(DirectML) != POINTER_INVALID)
{
if(!DirectML.CalcInputGradientProof(NeuronOCL.getOutputIndex(), getGradientIndex(), getOutputIndex(), NeuronOCL.getGradientIndex(),
outputs, (int)iWindow, (int)iStep, inputs))
{
Print(__FUNCTION__ + ": " + DirectML.BackendName() + " CalcInputGradientProof failed, error " + IntegerToString(DirectML.LastError()));
return false;
}
double temp[];
return NeuronOCL.getGradient(temp) > 0;
}
if(CheckPointer(OpenCL) == POINTER_INVALID)
return false;
uint offset1[1] = {0};
uint size1[1] = {(uint)inputs};
OpenCL.SetArgumentBuffer(def_k_CalcInputGradientProof, def_k_cigp_matrix_i, NeuronOCL.getOutputIndex());
OpenCL.SetArgumentBuffer(def_k_CalcInputGradientProof, def_k_cigp_matrix_g, getGradientIndex());
OpenCL.SetArgumentBuffer(def_k_CalcInputGradientProof, def_k_cigp_matrix_o, getOutputIndex());
OpenCL.SetArgumentBuffer(def_k_CalcInputGradientProof, def_k_cigp_matrix_ig, NeuronOCL.getGradientIndex());
OpenCL.SetArgument(def_k_CalcInputGradientProof, def_k_cigp_outputs, outputs);
OpenCL.SetArgument(def_k_CalcInputGradientProof, def_k_cigp_window, (int)iWindow);
OpenCL.SetArgument(def_k_CalcInputGradientProof, def_k_cigp_step, (int)iStep);
if(!OpenCL.Execute(def_k_CalcInputGradientProof, 1, offset1, size1))
{
printf("Error of execution kernel CalcInputGradientProof: %d", GetLastError());
return false;
}
//--- NeuronOCL's Gradient stays GPU-resident; see the note in
//--- CNeuronConvOCL::calcInputGradients().
return true;
}
//+------------------------------------------------------------------+
//| |
//+------------------------------------------------------------------+
bool CNeuronPoolOCL::Save(const int file_handle)
{
if(!CNeuronBaseOCL::Save(file_handle))
return false;
if(FileWriteInteger(file_handle, (int)iWindow, INT_VALUE) < INT_VALUE)
return false;
if(FileWriteInteger(file_handle, (int)iStep, INT_VALUE) < INT_VALUE)
return false;
return true;
}
//+------------------------------------------------------------------+
//| |
//+------------------------------------------------------------------+
bool CNeuronPoolOCL::Load(const int file_handle)
{
if(!CNeuronBaseOCL::Load(file_handle))
return false;
iWindow = (uint)FileReadInteger(file_handle, INT_VALUE);
iStep = (uint)FileReadInteger(file_handle, INT_VALUE);
return true;
}