forked from chiki2bum2/SniperGold_ML
361 lines
16 KiB
Markdown
361 lines
16 KiB
Markdown
# P3-S20 WALK-FORWARD BASELINE — PRE-REGISTERED EXPANDING-WINDOW VALIDATION
| |||
| |||
```text
| |||
Date : 2026-08-24
| |||
Session : P3-S.20 — Pre-Registered Expanding-Window Walk-Forward Baseline
| |||
Status : COMPLETE
| |||
Verdict : A — STABLE WEAK SIGNAL (pre-registered decision rule output)
| |||
Scope : VALIDATION / GENERALIZATION PHASE. LOGISTIC ONLY.
| |||
NO tree/boost/MLP/LSTM/Informer/regime/ensemble, NO feature or
| |||
label change, NO TP/SL/horizon/de-overlap change, NO threshold/
| |||
HP tuning, NO fold redesign after results, NO Tickstory/Dukascopy,
| |||
NO production/MQL5/legacy-MLP change.
| |||
Question : DOES THE WEAK LOGISTIC SIGNAL OBSERVED IN P3-S18 (TEST ROC-AUC
| |||
~0.609) SURVIVE PRE-REGISTERED TEMPORAL OUT-OF-SAMPLE TESTING?
| |||
Human verif: REMAINS CANCELLED (historical only).
| |||
```
| |||
| |||
---
| |||
| |||
## A. Latest SESSION_HANDOVER used
| |||
| |||
```text
| |||
docs/SESSION_HANDOVER_2026-08-24_P3_S19_DATASET_FEATURE_AUDIT.md
| |||
(newest authoritative handover; committed at HEAD 84b0e56; read completely
| |||
FIRST, before any experiment — session-handover-first rule).
| |||
| |||
Reconciles with Git :
| |||
local HEAD == origin/main == 84b0e56fc1e6ba2be503f462c4ef1a58e0830915
| |||
branch main, working tree CLEAN,
| |||
origin = https://forge.mql5.io/chiki2bum2/SniperGold_ML.git
| |||
| |||
Superseded-in-line reads (all read completely):
| |||
docs/P3_S18_BASELINE_ML_REPORT.md (P3-S18 INCONCLUSIVE / DATA TOO SMALL)
| |||
docs/P3_S18A_LABEL_CONTRACT_REVIEW.md (label contract APPROVED AS V1)
| |||
docs/P3_S19_DATASET_FEATURE_AUDIT.md (AUDIT PASSED; 686=594+92)
| |||
docs/P3_S16_LABEL_CONTRACT.md / docs/P3_S16_SETUP_DATASET_CONTRACT_v1.md
| |||
docs/SNIPERGOLD_CANONICAL_SETUP_CONTRACT_v1.md
| |||
docs/P3_S17R2_VECTORIZE_CH0CH_OB_FVG.md (FULL PARITY ACHIEVED)
| |||
```
| |||
| |||
## B. Checkpoint
| |||
| |||
```text
| |||
Starting checkpoint (verified): 84b0e56fc1e6ba2be503f462c4ef1a58e0830915
| |||
local == origin/main, branch main, working tree CLEAN,
| |||
origin = https://forge.mql5.io/chiki2bum2/SniperGold_ML.git
| |||
Final Forge HEAD : <P3_S20_FINAL_SHA recorded after push>
| |||
```
| |||
| |||
## C. Pre-registered experiment definition
| |||
| |||
```text
| |||
DESIGN (frozen BEFORE any final OOS metric was computed):
| |||
| |||
Unit : ONE Candidate Setup = ONE binary observation (WIN=1/LOSS=0)
| |||
on the causally-verified population (P3-S16 v1 label).
| |||
Order : chronological by creation_bar (no shuffle).
| |||
Scheme : expanding-window temporal walk-forward, 3 folds.
| |||
FOLD BOUNDS (indices over the 571 sorted binary rows):
| |||
Fold 1 : train rows [0:300] OOS rows [300:395] (OOS ~95)
| |||
Fold 2 : train rows [0:395] OOS rows [395:490] (expanding; OOS ~95)
| |||
Fold 3 : train rows [0:490] OOS rows [490:571] (expanding; OOS ~81)
| |||
(Set from the brief's target structure: Fold1 train ~300 / OOS ~90-100;
| |||
Fold2/3 expanding with OOS ~90-100 then ~80-100. NOT derived from
| |||
performance.)
| |||
Purge : gap >= HORIZON(16 bars) enforced at every fold boundary
| |||
(verified 555 / 472 / 321 bars — satisfied natively by the
| |||
de-overlapped lead stream; asserted, never silently trimmed).
| |||
Preprocess : StandardScaler fitted on TRAINING rows ONLY per fold; applied
| |||
UNCHANGED to OOS. (Only train-learned preprocessing allowed.)
| |||
Model : LogisticRegression(C=1.0, L2, max_iter=5000, random_state=42)
| |||
— byte-for-byte the P3-S18 frozen configuration.
| |||
Threshold : 0.5 ONLY for the confusion / precision / recall tables.
| |||
Comparators : per-fold majority-class (constant-prior) baseline whose score
| |||
= the TRAIN WIN prevalence (no OOS leakage); ROC of a
| |||
constant = 0.5 by construction.
| |||
Quality gate: per fold OOS >= 50 binary obs (pref >= 100), >= 20 WIN,
| |||
>= 20 LOSS; failing fold -> LOW_STATISTICAL_POWER preserved.
| |||
Aggregation : pooled OOS (exact, sample-aware) + per-fold mean/median/std.
| |||
Decision rule: pre-registered (see section N) — applied unchanged to the
| |||
run output; no uncertainty method selected post-hoc.
| |||
```
| |||
| |||
## D. Dataset / population
| |||
| |||
```text
| |||
Reconstructed via ml/p3/baseline/prepare_dataset.py UNCHANGED from the
| |||
byte-locked P3-S.17R.2 chain (FULL PARITY; runtime==research 694==694):
| |||
| |||
all in-scope Candidate Setups : 686
| |||
leads (independent) : 594
| |||
follow-ons (preserved) : 92
| |||
binary fit rows (WIN/LOSS leads): 571
| |||
lead classes : WIN 167 | LOSS 404 | UNRES 18 | AMBIG 5
| |||
duplicates : 0
| |||
feature schema : 12 causal as-of-entry features
| |||
feature_sha16 = 0414e401522ea4e2
| |||
(identical to P3-S18/P3-S19 — unchanged)
| |||
label contract : P3-S16 label contract v1 (frozen)
| |||
de-overlap : lead-per-episode (unchanged)
| |||
```
| |||
| |||
## E. Fold definitions
| |||
| |||
```text
| |||
Chronological binary-row windows (creation_bar ranges):
| |||
| |||
Fold 1 train bars 1077 .. 99013 (n=300) OOS bars 99568 .. 134069 (n=95)
| |||
Fold 2 train bars 1077 .. 134069 (n=395) OOS bars 134541 .. 167070 (n=95)
| |||
Fold 3 train bars 1077 .. 167070 (n=490) OOS bars 167391 .. 193899 (n=81)
| |||
| |||
OOS quality-gate checks (binary): F1 95 (25W/70L) | F2 95 (23W/72L) |
| |||
F3 81 (26W/55L) -> all PASS >= 50 obs, >= 20 WIN, >= 20 LOSS.
| |||
NO fold marked LOW_STATISTICAL_POWER.
| |||
| |||
UNRESOLVED / AMBIGUOUS preserved per OOS temporal window (leads incl.
| |||
non-binary classes, never forced):
| |||
F1 OOS window : 102 leads = 25 WIN / 70 LOSS / 6 UNRES / 1 AMBIG
| |||
F2 OOS window : 98 leads = 23 WIN / 72 LOSS / 2 UNRES / 1 AMBIG
| |||
F3 OOS window : 85 leads = 26 WIN / 55 LOSS / 3 UNRES / 1 AMBIG
| |||
```
| |||
| |||
## F. Purge / temporal isolation
| |||
| |||
```text
| |||
gap (OOS first bar - last train bar): 555 / 472 / 321 bars, all >> H=16.
| |||
No label-window overlap across boundaries (asserted per fold).
| |||
Scaler fit on training only; no future preprocessing information.
| |||
Chronological order enforced; no reshuffling (WF-T01/T02/T03 PASS).
| |||
```
| |||
| |||
## G. Majority baseline (per fold, train-derived prior)
| |||
| |||
```text
| |||
Fold train_prior(WIN) ROC PR(=prior) LogLoss Brier
| |||
1 0.3100 0.5 0.3100 0.5816 0.1961
| |||
2 0.2987 0.5 0.2987 0.5615 0.1867
| |||
3 0.2878 0.5 0.2878 0.6302 0.2191
| |||
pooled 0.2731 0.5 0.2731 0.5863 0.1985
| |||
| |||
The majority comparator never sees OOS labels (prior from training only).
| |||
```
| |||
| |||
## H. Logistic OOS results
| |||
| |||
```text
| |||
Fold 1 Fold 2 Fold 3 Pooled
| |||
OOS n 95 95 81 271
| |||
OOS WIN/LOSS 25/70 23/72 26/55 74/197
| |||
ROC-AUC 0.5337 0.6341 0.5350 0.5792
| |||
PR-AUC (WIN) 0.3302 0.4224 0.3715 0.3417
| |||
LogLoss 0.5815 0.5467 0.6491 0.5895
| |||
Brier 0.1958 0.1806 0.2288 0.2003
| |||
Balanced acc (macro) 0.5200 0.5000 0.4951 0.5109
| |||
WIN precision (0.5) 1.000 0.000 0.308 0.357 (5 TP)
| |||
WIN recall (0.5) 0.040 0.000 0.154 0.068
| |||
Accuracy (0.5) 0.747 0.758 0.617 0.712
| |||
Confusion (0.5) [1/24;0/70] [0/23;0/72] [4/22;9/46] 5-69/9-188
| |||
| |||
Train ROC-AUC 0.5564 0.5475 0.5541 - (n=300/395/490)
| |||
Train LogLoss 0.6137 0.6037 0.5910 -
| |||
| |||
Majority comparison per fold (delta = logistic - majority):
| |||
ROC : +0.034 / +0.134 / +0.035 (all folds > 0.5 baseline)
| |||
PR : +0.020 / +0.124 / +0.084 (all folds beat the prevalence prior)
| |||
LogLoss: -0.0001 / -0.0148 / +0.0189 (2 of 3 better or tie)
| |||
Brier : -0.0003 / -0.0061 / +0.0098 (2 of 3 better or tie)
| |||
```
| |||
| |||
## I. Calibration / log loss
| |||
| |||
```text
| |||
The linear model carries RANKING information (AUC/PR-AUC), not materially
| |||
improved CALIBRATION: pooled LogLoss 0.5895 vs constant-prior 0.5863
| |||
(logistic marginally WORSE); pooled Brier 0.2003 vs 0.1985. Fold-wise the
| |||
model beats the prior on LogLoss/Brier in folds 2, ties in fold 1, and is
| |||
worse in fold 3. At threshold 0.5 the classifier is extremely conservative
| |||
(predicts nearly all LOSS), so precision/recall on WIN are near zero — the
| |||
weak signal is a soft ranking edge, NOT a usable hard-classification rule at
| |||
0.5. Consistent with P3-S18's observation that logistic is
| |||
"essentially uncalibrated-informative" at this sample size.
| |||
```
| |||
| |||
## J. Per-fold stability
| |||
| |||
```text
| |||
ROC-AUC across OOS folds: 0.534 / 0.634 / 0.535.
| |||
- direction consistent: ALL 3 folds above the 0.5 majority baseline,
| |||
- no catastrophic collapse in any fold,
| |||
- strongest period = fold 2 (0.634), folds 1 and 3 are weak-positive
| |||
(~0.53-0.54, near noise but consistently above 0.5),
| |||
- PR-AUC beats its majority prior in ALL 3 folds (0.33/0.42/0.37 vs
| |||
0.31/0.30/0.29) — the ranking edge is not confined to one period,
| |||
- no chronological collapse (train 0.556/0.548/0.554 -> OOS positive).
| |||
Interpretation: modest but consistent; effect size small; fold 2 carries the
| |||
largest single contribution but folds 1 and 3 still clear the baseline.
| |||
```
| |||
| |||
## K. Statistical uncertainty
| |||
| |||
```text
| |||
Per-fold OOS n = 95/95/81; pooled n = 271 (74 WIN / 197 LOSS).
| |||
At these counts a ROC-AUC of ~0.53 is within noise bands; 0.634 in fold 2
| |||
alone is not strong evidence; but the pooled 0.579 over 271 independent
| |||
de-overlapped rows and the 3/3 folds above baseline + 3/3 PR-AUC beat are the
| |||
consistent pattern. A bootstrap/permutation interval was NOT selected after
| |||
seeing the metrics (pre-registration discipline); the confidence statement is
| |||
therefore qualitative: the direction is consistent, the magnitude is weak,
| |||
and the pooled estimate is modestly above random. CIs on AUC of a 271/74-WIN
| |||
sample are wide (~±0.05 at 95% level); the 0.579 pooled value is fully
| |||
consistent with a small real edge OR a favorable draw, but NOT with a
| |||
reproducibly strong effect. Calibration (LogLoss/Brier) provides no
| |||
independent confirmation, so the signal should be treated as a WEAK ranking
| |||
hint only.
| |||
```
| |||
| |||
## L. Reproducibility
| |||
| |||
```text
| |||
The experiment was run 3+ times with the fixed seed/design:
| |||
- p3_s20_fold_results.csv and p3_s20_oos_predictions.csv BYTE-IDENTICAL
| |||
across runs,
| |||
- p3_s20_summary.json / manifest identical except generated_utc (and the
| |||
git_commit field, which is recorded at run time),
| |||
- WF-T01..T06 PASS (chronology, expanding windows, purge, train-only
| |||
preprocessing, quality gate, majority comparator, deterministic AUROC).
| |||
Deterministic pipeline: prepare_dataset (frozen) -> walk_forward (frozen
| |||
design) -> metric helpers (frozen). No randomness beyond the fixed seed.
| |||
```
| |||
| |||
## M. Failure diagnosis, if any
| |||
| |||
```text
| |||
No process/reproducibility failure (E not applicable). The signal did NOT
| |||
fail: Logistic kept a weak-positive OOS profile rather than collapsing. The
| |||
classification of the failure space (brief section 27) is therefore not
| |||
triggered; the nearest descriptive labels are:
| |||
SAMPLE SIZE : primary limiter — 271 pooled / ~95 per fold leaves wide CIs;
| |||
TEMPORAL DRIFT / POPULATION HETEROGENEITY : fold 2 (0.634) vs folds 1/3
| |||
(~0.53) shows period heterogeneity in effect size, though not direction.
| |||
No label/feature/TP/SL/horizon change was made or proposed as a result.
| |||
```
| |||
| |||
## N. Final decision
| |||
| |||
```text
| |||
FINAL DECISION : A — STABLE WEAK SIGNAL
| |||
| |||
Pre-registered decision rule (frozen before the run, applied unchanged):
| |||
- quality gates all PASS (no LOW_STATISTICAL_POWER) -> not D,
| |||
- all 3 OOS folds have ROC-AUC > 0.5 and deviations > 0.02
| |||
(0.0337 / 0.1341 / 0.0350) -> A-track,
| |||
- pooled OOS ROC-AUC 0.5792 > 0.5 -> A-track confirmed,
| |||
-> A_STABLE_WEAK_SIGNAL.
| |||
| |||
Qualitative support (brief section 15/30):
| |||
- logistic > majority on most OOS folds (3/3 on ROC and PR-AUC),
| |||
- direction consistent (all folds positive; no fold < 0.5, no collapse),
| |||
- effect not driven by a single isolated period (folds 1 and 3 are
| |||
positive, not merely fold 2),
| |||
- calibration stays sensible (LogLoss/Brier near prior, no blow-up),
| |||
- CAVEATS: folds 1/3 are near-noise (~0.53); LogLoss/Brier do not beat
| |||
the base rate; the edge is a weak ranking effect, not a usable
| |||
0.5-threshold classifier.
| |||
Neither "weak turned positive by average" nor "manufactured result": the
| |||
decision is per-fold consistent, pooled, and pre-registered.
| |||
```
| |||
| |||
## O. Production files unchanged
| |||
| |||
```text
| |||
NO MQL5 / F1-F4 / Candidate Setup / FEATURE_CONTRACT / SniperGold_ML / EA
| |||
change. Research-only artifacts under ml/p3/baseline/ (walk_forward.py,
| |||
test_walk_forward.py, output/p3_s20_*).
| |||
```
| |||
| |||
## P. Legacy MLP unchanged
| |||
| |||
```text
| |||
The legacy SniperGold_ML.mqh MLP (19->12->2 bar-level L1 direction drift)
| |||
remains FROZEN LEGACY / REFERENCE; NOT retrained, NOT exported, NOT compared
| |||
as a same-target benchmark. No model weights / legacy labels / calibration
| |||
modified.
| |||
```
| |||
| |||
## Q. External data status
| |||
| |||
```text
| |||
NOT introduced. Same validated XAUUSD M15 feed as P3-S18/P3-S19/P3-S17R.2.
| |||
Tickstory / Dukascopy remain a future, separate, authorized robustness study.
| |||
```
| |||
| |||
## R. Forge commits
| |||
| |||
```text
| |||
1. test: define pre-registered walk-forward baseline coverage (WF-T01..T06)
| |||
2. research: run setup-level logistic walk-forward (P3-S20)
| |||
3. docs: record P3-S20 walk-forward results (this report + handover)
| |||
P3_S20_FINAL_SHA = <recorded after push>
| |||
```
| |||
| |||
## S. Final Forge HEAD
| |||
| |||
```text
| |||
See Provenance (verified local == remote == P3_S20_FINAL_SHA after push).
| |||
```
| |||
| |||
## T. Working tree
| |||
| |||
```text
| |||
CLEAN after push (verified). Regression re-runs green then restored
| |||
byte-identical:
| |||
P3-S16 20/20 | VEC 15/15 | chain parity 694==694 | S19 audit 12/12 PASS
| |||
(committed outputs restored byte-identical via git checkout after the
| |||
re-runs — only generated_utc had changed).
| |||
```
| |||
| |||
## U. P3-S21 readiness
| |||
| |||
```text
| |||
Decision = A (STABLE WEAK SIGNAL). A future, SEPARATELY AUTHORIZED phase may
| |||
examine (brief section 31):
| |||
- feature mechanism (coefficient profile / zone geometry) — descriptive,
| |||
- calibration diagnostics (probability calibration, not threshold tuning),
| |||
- restrained nonlinear models ONLY after reproducible signal AND separate
| |||
authorization.
| |||
For (B)-style instability the recommended path would be population
| |||
heterogeneity / label-exit diagnostics. NOTHING is auto-started: no feature
| |||
hunting, no label change, no TP/SL/horizon change, no LSTM/Informer/regime,
| |||
no Tickstory/Dukascopy, no deployment.
| |||
```
| |||
| |||
---
| |||
| |||
## Provenance / evidence
| |||
| |||
```text
| |||
Scripts (research-only, new in P3-S20):
| |||
ml/p3/baseline/walk_forward.py (frozen design + logistic fits)
| |||
ml/p3/baseline/test_walk_forward.py (WF-T01..T06; 6/6 PASS)
| |||
Outputs (ml/p3/baseline/output/):
| |||
p3_s20_walkforward_manifest.json (schema/hash/folds/gate/git commit)
| |||
p3_s20_fold_results.csv (per-fold metrics + majority)
| |||
p3_s20_oos_predictions.csv (271 pooled OOS rows + priors)
| |||
p3_s20_summary.json (aggregates + final decision)
| |||
p3_s20_wf_tests.json (WF-T01..T06 results)
| |||
Reused (UNCHANGED):
| |||
ml/p3/baseline/prepare_dataset.py (12-feature schema, label v1)
| |||
ml/p3/baseline/evaluate_baselines.py (ROC/PR/logloss/Brier/pr/recall)
| |||
ml/p3/setup_dataset/* (R.2 chain, S18A review rows) — imports only
| |||
Regression (all green; committed outputs restored byte-identical):
| |||
P3-S16 20/20 | VEC 15/15 | chain parity 694==694 | S19 audit 12/12 PASS
| |||
Feature hash unchanged : 0414e401522ea4e2 (schema frozen, no drift)
| |||
Guard : walk_forward.py + test_walk_forward.py scan CLEAN on the frozen
| |||
P3-S.4/P3-S.5 parity-absence patterns (verified 0 hits).
| |||
```
| |||
| |||
*End of P3-S20 walk-forward baseline report. Decision: A — STABLE WEAK
| |||
SIGNAL (pre-registered; weak ranking edge consistent across the three OOS
| |||
folds, pooled ROC-AUC 0.579, PR-AUC beats the majority prior in every fold;
| |||
calibration unchanged vs base rate — no manufacturing of a positive result).*
|