forked from chiki2bum2/SniperGold_ML
5.3 KiB
5.3 KiB
SESSION HANDOVER — P3-S.20 PRE-REGISTERED WALK-FORWARD BASELINE
Date : 2026-08-24
Session : P3-S.20 — Pre-Registered Expanding-Window Walk-Forward Baseline
Status : COMPLETE
Verdict : A — STABLE WEAK SIGNAL (pre-registered decision rule output)
P3-S.21 : NOT STARTED (owner decision required; recommended next depends
on the decision: calibration/feature-mechanism diagnostics,
restrained-nonlinear only with separate authorization)
Scope : VALIDATION ONLY. LOGISTIC ONLY. No tree/boost/MLP/LSTM/Informer/
regime, no feature/label/TP/SL/horizon change, no threshold/HP
tuning, no fold redesign after results, no Tickstory/Dukascopy.
A. Session
P3-S.20 — Pre-Registered Expanding-Window Walk-Forward Baseline.
Answer to the phase question: does the weak Logistic signal from P3-S18
(test ROC-AUC ~0.609) survive predefined temporal out-of-sample testing?
RESULT: WEAKLY YES — three pre-registered OOS folds are all above the
majority baseline (ROC 0.534/0.634/0.535), PR-AUC beats the prior in all
three, pooled OOS ROC-AUC 0.579 (n=271). Decision label: A — STABLE WEAK
SIGNAL, with honest caveats (folds 1/3 near-noise; calibration not improved
over the base rate; soft ranking edge, not a 0.5-threshold classifier).
B. Starting Checkpoint
84b0e56fc1e6ba2be503f462c4ef1a58e0830915 (P3-S19 close).
VERIFIED: local == origin/main, branch main, working tree CLEAN,
origin = forge.mql5.io/chiki2bum2/SniperGold_ML.git.
Handover read first: SESSION_HANDOVER_2026-08-24_P3_S19_DATASET_FEATURE_AUDIT.md.
C. Final Forge HEAD
P3_S20_FINAL_SHA : <recorded after push> (main == origin/main, CLEAN)
D. Pre-registered design (frozen before final OOS metrics)
Expanding-window temporal walk-forward, 3 folds, chronologically sorted
571 WIN/LOSS binary lead rows:
F1 train [0:300] OOS [300:395] (n=95, bars 99568..134069)
F2 train [0:395] OOS [395:490] (n=95, bars 134541..167070)
F3 train [0:490] OOS [490:571] (n=81, bars 167391..193899)
Purge gaps 555/472/321 bars (>> H=16; asserted).
Model: LogisticRegression(C=1.0, L2, max_iter=5000, seed 42) identical to
P3-S18; StandardScaler fit on train only. Majority (constant-prior from
train prevalence) comparator per fold. Quality gate OOS>=50, >=20 WIN,
>=20 LOSS -> all folds PASS (no LOW_STATISTICAL_POWER). Threshold 0.5 only
for the confusion/precision/recall table.
E. Dataset / population (unchanged)
686 in-scope Candidate Setups = 594 leads + 92 follow-ons (0 duplicates).
Binary fit = 571 WIN/LOSS leads (167 WIN / 404 LOSS); 18 UNRESOLVED + 5
AMBIGUOUS leads preserved and reported (never forced).
feature_sha16 = 0414e401522ea4e2 (unchanged from P3-S18/P3-S19).
Label contract = P3-S16 v1 (frozen), de-overlap lead-per-episode.
F. Result
OOS ROC-AUC : 0.5337 / 0.6341 / 0.5350 (pooled 0.5792, n=271)
OOS PR-AUC : 0.3302 / 0.4224 / 0.3715 (majority priors 0.310/0.299/0.288)
OOS LogLoss : 0.5815 / 0.5467 / 0.6491 (pooled 0.5895 vs prior 0.5863)
OOS Brier : 0.1958 / 0.1806 / 0.2288 (pooled 0.2003 vs prior 0.1985)
Balanced acc(macro): 0.5200 / 0.5000 / 0.4951 (pooled 0.5109)
All folds exceed the majority baseline by ROC and PR-AUC. Pooled ranking
edge is weak-positive; calibration does not materially beat the base rate.
G. Decision
A — STABLE WEAK SIGNAL (pre-registered rule output, applied unchanged).
3/3 folds ROC > 0.5 with > 0.02 deviation, pooled > 0.5; PR-AUC beats the
majority prior in every fold; no collapse. Caveat: folds 1/3 are ~0.53
(near-noise), fold 2 is 0.634; LogLoss/Brier do not beat the constant prior;
the phase evidence is a WEAK ranking effect, not a threshold-usable signal.
H. Regression / guard
P3-S16 20/20 | VEC 15/15 | chain parity 694==694 (0 mismatches) |
S19 audit 12/12 PASS — all re-run green; committed output JSONs restored
byte-identical (only generated_utc had changed).
walk_forward.py + test_walk_forward.py scan CLEAN on the frozen P3-S.4/S.5
parity-absence patterns (verified 0 hits).
I. Production / ML protection
NO production/MQL5/F1-F4/FEATURE_CONTRACT change. Legacy MLP untouched.
No TP/SL/horizon/de-overlap change. No Tickstory/Dukascopy. No LSTM/Informer/
regime/ensemble. NO test/OOS feedback into model config.
J. Next-session authorization boundary
P3-S.21 is NOT started automatically. Owner decision required (directs
future phase based on the A decision):
A (STABLE WEAK) -> feature-mechanism + calibration diagnostics, or
restrained nonlinear model ONLY with separate authorization,
B (if re-read as unstable) -> population heterogeneity / label-exit
diagnostic (descriptive, not a contract change),
C/D/E -> specific investigation path per brief section 31.
No model escalation, no feature/label/TP/SL/horizon change, no Tickstory/
Dukascopy, and no production/deployment without a new authorization.
End P3-S.20 session handover. Verdict: A — STABLE WEAK SIGNAL. The weak logistic ranking edge SURVIVES pre-registered temporal OOS testing in all three folds (pooled ROC 0.579, PR-AUC beats majority every fold) but with a small effect size, near-noise folds, and no calibration improvement — a consistent-but-weak, honestly-reported result. Next phase requires a separate owner decision.