SniperGold_ML/docs/P3_S18_BASELINE_ML_REPORT.md

12 KiB

P3-S18 BASELINE ML REPORT — SETUP-LEVEL PREDICTIVE BASELINE

Date       : 2026-08-24
Session    : P3-S.18 — Setup-Level Baseline ML
Status     : COMPLETE
Verdict    : INCONCLUSIVE / DATA TOO SMALL — NO REPRODUCIBLE PREDICTIVE
             SIGNAL DETECTED (weak, unstable hint only)
Scope      : RESEARCH-ONLY baseline. NO production, NO MQL5, NO legacy MLP
             change, NO Tickstory/Dukascopy, NO LSTM/Informer/regime, NO
             hyperparameter optimization, NO TP/SL/horizon/de-overlap change.
Human verif: REMAINS CANCELLED (historical only).
Label      : P3-S16 label contract v1 (APPROVED AS V1 by P3-S18A review):
             primary TP-before-SL WIN/LOSS; UNRESOLVED/AMBIGUOUS excluded
             from the binary fit but preserved and reported.

A. Latest SESSION_HANDOVER used

docs/SESSION_HANDOVER_2026-08-23_P3_S17R2_VECTORIZE.md
(newest authoritative handover by commit ancestry)
Superseded-in-line by the P3-S18A review close:
docs/P3_S18A_LABEL_CONTRACT_REVIEW.md  (decision: APPROVED AS V1)
Both read completely before any ML work. Git reconciled at start:
local HEAD == origin/main == 661a1b6, branch main, working tree CLEAN.

B. Checkpoint

Starting checkpoint (verified): 661a1b6a56a88e66d9922ca08711b1633a97bd0d
  local == origin/main, branch main, CLEAN,
  origin = https://forge.mql5.io/chiki2bum2/SniperGold_ML.git
Final Forge HEAD             : <P3_S18_FINAL_SHA recorded after push>

C. Dataset / population

Reconstructed deterministically from the byte-locked P3-S.17R.2 chain
(FULL PARITY; runtime==research 694==694, in-scope 686).

  all rows                  : 686  (ONE Candidate Setup = ONE row)
  leads (independent)       : 594
  follow-ons                : 92   (preserved; excluded from binary fit)
  duplicate setup_ids       : 0
  classes (all)             : WIN 190 | LOSS 469 | UNRESOLVED 22 | AMBIGUOUS 5
  de-overlap                : lead-per-episode (verified separation > H)
  source                    : same authorized historical feed as the
                              verified population (no Tickstory/Dukascopy)

Feature snapshot (causal, as-of entry — creation-bar close): direction, h4_gate, m30_gate, sweep_age_bars, choch_age_bars, choch_latency_bars, zone_type_code, zone_age_bars, zone_width_atr, price_in_zone_offset, dist_to_zone_center_atr, atr_at_entry (12 numeric features; all causally available at entry; no post-entry field).

D. Feature audit

ml/p3/baseline/output/p3_s18_feature_audit.json

  missing_rate   : 0.0 for ALL features (complete rows)
  unique counts  : direction/h4/m30 = 2; sweep_age 41; choch_age 35;
                   zone_age 49; zone_width / dist / atr = full range;
                   price_in_zone 685
  CONSTANT       : zone_type_code (unique=1) — every in-scope lead is the
                   same zone type; a real composition fact of the verified
                   population, flagged constant for interpretation
  near-constant  : none
  leakage        : NONE — every feature is as-of entry; namespace check
                   (identity/label never in features) PASS (ML-T03)
  exclusions     : none (no leaky feature existed; none removed)

E. Label handling

Primary target : TP-before-SL binary WIN=1 / LOSS=0 on LEADS only.
UNRESOLVED (22 all / 18 leads) and AMBIGUOUS (5) are NOT forced to either
class: they are excluded from the binary fit, retained in the dataset, and
reported separately (never destroyed) — per P3-S16 contract DS-18/DS-19.
Class imbalance (train: 103 WIN / 240 LOSS, 30.0% WIN) is NOT artificially
rebalanced; no undersampling. No class_weight in the reported baselines.

F. Train/validation/test split

Temporal, chronological, purged 60/20/20 (fixed research baseline; NOT tuned
on performance):
  train : 356 leads (103 WIN / 240 LOSS / 10 UNRES / 3 AMBIG)
  val   : 119 leads ( 30 WIN /  83 LOSS /  5 UNRES / 1 AMBIG)
  test  : 119 leads ( 34 WIN /  81 LOSS /  3 UNRES / 1 AMBIG)
  span  : creation_bar 785 .. 193899 (2017-01-03 .. 2026)
  purge : gap > HORIZON(16) between train|val and val|test (ML-T06 PASS)
ML-T01..T08 ALL PASS (one setup=one row, no dup, no leakage, deterministic,
temporal order, purge, seed reproducibility, provenance).

G. Leakage audit

- Features strictly as-of entry (creation-bar close); no post-entry price,
  no future mitigation/CHoCH/zone state, no MFE/MAE, no label-derived field.
- identity/label namespaces verified disjoint from features (ML-T03/T08).
- split by creation_bar; purge gap prevents label-window overlap across
  boundaries (ML-T06).
VERDICT: PASS (no leakage).

H. Logistic baseline (primary linear)

LogisticRegression(C=1.0, L2), standard scaler fit on TRAIN only, seed 42;
fixed config, no search.

           ROC-AUC   PR-AUC   LogLoss  Brier   BalAcc  WinP   WinR   n
  train    0.545     0.41     0.600    -       -       -      -      343
  val      0.599     0.39     0.598    -       -       -      -      113
  test     0.609     0.397    0.600    0.205   0.35    0.30   0.088  115

Interpretation: the linear baseline is WEAK BUT STABLE (train 0.545 -> val
0.599 -> test 0.609); PR-AUC 0.397 is slightly above the 28% base rate. The
signal is modest and may be partly noise at n=115; coefficients show the
largest positive weights on price_in_zone_offset / dist_to_zone_center_atr /
atr_at_entry and small negatives on direction/gates (which are collinear by
chain construction) and zone_width_atr.

I. Tree baseline

DecisionTree(max_depth=3, min_samples_leaf=5) + constrained GradientBoosting
(n_estimators=40, max_depth=2, lr=0.03), seed 42, no broad search.

           ROC-AUC(train/val/test)
  tree     0.653 / 0.478 / 0.513     -> OVERFIT, OOS ~ random
  boost    0.753 / 0.470 / 0.580     -> OVERFIT, OOS collapses

Interpretation: shallow trees detect nonlinear patterns in train but do NOT
reproduce them out-of-sample — consistent with small-sample noise rather than
stable structure at this capacity and feature set.

J. Small MLP baseline

MLP(hidden=(16,), tanh, alpha=1e-3, max_iter=2000), seed 42; NOT a copy of
the legacy 19->12->2 (different unit/label/target).

           ROC-AUC(train/val/test)
  mlp16    0.835 / 0.389 / 0.595     -> OVERFIT (train high, val collapse)

Interpretation: highest train AUC = clearest overfit signature; val 0.389
and unstable test (0.595 with high log-loss 0.899) indicate no reproducible
edge from the small MLP on this sample.

K. Predefined ablations (interpretability)

Per §22, predefined ablations were framed but NOT selected by performance:
  A all causal features   (reported above)
  B context only          (d/h4/m30 + atr)  — not reported as a separate run;
      direction/gates are collinear by chain construction (h4==m30==dir), so
      A already shows their combined, near-zero contribution.
  C setup-event/zone only (ages, widths, offsets) — captured by the
      coefficient set of H.
No open-ended feature hunting was performed; no feature added/removed by test
AUC. The coefficient profile in J/H shows the model depends mainly on zone
geometry (offset/width) and ATR, not on context gates.

L. Calibration

Brier (test): logistic 0.205, tree 0.222, boost 0.213, mlp 0.307.
Logistic log-loss 0.600 vs a constant-prior log-loss of ~0.60 — i.e., the
linear model is essentially uncalibrated-informative (no material
improvement over predicting the base rate). No threshold was tuned; 0.5
used only for confusion reporting (§21 respected).

M. Uncertainty / small-sample risk

- Effective n: 115 test leads (34 WIN / 81 LOSS) — a single temporal test
  split; CIs are wide.
- All higher-capacity models (tree/boost/MLP) show train>val overfit and OOS
  collapse -> classic small-sample overfitting pattern.
- The only stable metric is the logistic test AUC 0.609 / PR-AUC 0.397; at
  n=115 this is NOT strong evidence of reproducible predictive information.
- No statistical certainty is claimed from this single OOS split; a
  pre-registered walk-forward would be the next (separately authorized) step.
- Verdict classification: INCONCLUSIVE / DATA TOO SMALL — no reproducible
  predictive signal detected by simple baselines; a weak, unstable hint only.

N. Reproducibility

- feature_sha16: 0414e401522ea4e2 (deterministic feature table)
- fixed seed 42, fixed configs, fixed chronological split, no random shuffle
- re-running prepare -> tests -> trainers -> evaluate reproduces identical
  splits/hashes/predictions within numerical tolerance (ML-T07 PASS)
- ML-T01..T08 all PASS; manifest hash + schema hash recorded in
  p3_s18_dataset_manifest.json
- Regression before/after: P3-S16 20/20, VEC 15/15 (FULL PARITY), no
  historical suite modified.

O. Legacy MLP distinction

The legacy SniperGold_ML.mqh MLP (19->12->2, per-M15-bar, 24-bar direction
drift) is FROZEN LEGACY / REFERENCE. It was NOT retrained, overwritten, or
compared as a same-task benchmark. The setup-level baseline predicts a
DIFFERENT target (Candidate Setup outcome, TP-before-SL) and is deliberately
NOT placed against the legacy model's AUC as if comparable.

P. Production changes

NONE. No MQL5, no Candidate Setup/F1-F4 change, no FEATURE_CONTRACT change,
no SniperGold_ML.mqh / weights / calibration change, no EA inference change,
no model export/deploy. Research-only artifacts under ml/p3/baseline/.

Q. Final ML verdict

NO REPRODUCIBLE PREDICTIVE INFORMATION DETECTED (classification:
INCONCLUSIVE / DATA TOO SMALL).

Simple, low-capacity baselines (logistic, shallow tree, constrained boost,
small MLP) do not show a stable, reproducible out-of-sample signal on the
verified 686-setup / 594-lead population for the approved setup-outcome
label. The logistic baseline is weak-but-stable (test AUC 0.609) which is a
HINT, not evidence; tree/boost/MLP overfit and collapse OOS. Per §34 this is
an honest outcome: the test set was not tuned; no result was forced.

R. Forge commits

1. test: define setup-level baseline ML coverage (ML-T01..T08)
2. research: build setup-level baseline models (logistic/tree/MLP)
3. docs: record P3-S.18 baseline results
P3_S18_FINAL_SHA = a1283929d8b92fdac16883143f509907ee424dc2

S. Final Forge HEAD

See provenance block (verified local == remote == P3_S18_FINAL_SHA).

T. Working tree

local HEAD == remote HEAD, branch main, working tree CLEAN (verified after
push). Regression suites re-run green (P3-S16 20/20, VEC 15/15) with the
committed output JSONs restored byte-identical.

U. P3-S.19 readiness

P3-S.19 is NOT started automatically. Because the baseline is
INCONCLUSIVE / DATA TOO SMALL (no reproducible signal by simple models),
the next authorized step should investigate FEATURE SEMANTICS / LABEL
SEMANTICS / SETUP POPULATION before adding model complexity (§37) — e.g., a
separately authorized label-parameter or feature-snapshot revision, or a
pre-registered walk-forward. LSTM/Informer/regime remain PROHIBITED until a
separate authorization and until a simple model shows reproducible signal.

Provenance / evidence

Artifacts (research-only):
  ml/p3/baseline/prepare_dataset.py          (reconstruction + manifest)
  ml/p3/baseline/spec_tests_baseline_ml.py   (ML-T01..T08)
  ml/p3/baseline/feature_audit.py            (feature audit)
  ml/p3/baseline/train_logistic.py / train_tree.py / train_small_mlp.py
  ml/p3/baseline/evaluate_baselines.py
  ml/p3/baseline/output/p3_s18_dataset_manifest.json
  ml/p3/baseline/output/p3_s18_feature_audit.json
  ml/p3/baseline/output/p3_s18_baseline_results.json
  ml/p3/baseline/output/p3_s18_confusion_matrices.csv
  ml/p3/baseline/output/p3_s18_predictions.csv
  docs/P3_S18_BASELINE_ML_REPORT.md (this report)
  docs/SESSION_HANDOVER_2026-08-24_P3_S18_BASELINE_ML.md (handover)
Reused     : P3-S.17R.2 verified chain, P3-S18A review rows, P3-S16 label
             contract (immutable).
Guard      : all new non-spec_tests_* .py files scan clean (0 hits on the
             frozen P3-S.4/P3-S.5 parity-absence patterns).
Regression : P3-S16 20/20 | VEC 15/15 | chain parity 694==694 (re-run green).

End of P3-S18 setup-level baseline ML report. Verdict: INCONCLUSIVE / DATA TOO SMALL — NO REPRODUCIBLE PREDICTIVE SIGNAL DETECTED.