# SESSION HANDOVER — P3-S.20 PRE-REGISTERED WALK-FORWARD BASELINE ```text Date : 2026-08-24 Session : P3-S.20 — Pre-Registered Expanding-Window Walk-Forward Baseline Status : COMPLETE Verdict : A — STABLE WEAK SIGNAL (pre-registered decision rule output) P3-S.21 : NOT STARTED (owner decision required; recommended next depends on the decision: calibration/feature-mechanism diagnostics, restrained-nonlinear only with separate authorization) Scope : VALIDATION ONLY. LOGISTIC ONLY. No tree/boost/MLP/LSTM/Informer/ regime, no feature/label/TP/SL/horizon change, no threshold/HP tuning, no fold redesign after results, no Tickstory/Dukascopy. ``` --- ## A. Session ```text P3-S.20 — Pre-Registered Expanding-Window Walk-Forward Baseline. Answer to the phase question: does the weak Logistic signal from P3-S18 (test ROC-AUC ~0.609) survive predefined temporal out-of-sample testing? RESULT: WEAKLY YES — three pre-registered OOS folds are all above the majority baseline (ROC 0.534/0.634/0.535), PR-AUC beats the prior in all three, pooled OOS ROC-AUC 0.579 (n=271). Decision label: A — STABLE WEAK SIGNAL, with honest caveats (folds 1/3 near-noise; calibration not improved over the base rate; soft ranking edge, not a 0.5-threshold classifier). ``` ## B. Starting Checkpoint ```text 84b0e56fc1e6ba2be503f462c4ef1a58e0830915 (P3-S19 close). VERIFIED: local == origin/main, branch main, working tree CLEAN, origin = forge.mql5.io/chiki2bum2/SniperGold_ML.git. Handover read first: SESSION_HANDOVER_2026-08-24_P3_S19_DATASET_FEATURE_AUDIT.md. ``` ## C. Final Forge HEAD ```text P3_S20_FINAL_SHA : (main == origin/main, CLEAN) ``` ## D. Pre-registered design (frozen before final OOS metrics) ```text Expanding-window temporal walk-forward, 3 folds, chronologically sorted 571 WIN/LOSS binary lead rows: F1 train [0:300] OOS [300:395] (n=95, bars 99568..134069) F2 train [0:395] OOS [395:490] (n=95, bars 134541..167070) F3 train [0:490] OOS [490:571] (n=81, bars 167391..193899) Purge gaps 555/472/321 bars (>> H=16; asserted). Model: LogisticRegression(C=1.0, L2, max_iter=5000, seed 42) identical to P3-S18; StandardScaler fit on train only. Majority (constant-prior from train prevalence) comparator per fold. Quality gate OOS>=50, >=20 WIN, >=20 LOSS -> all folds PASS (no LOW_STATISTICAL_POWER). Threshold 0.5 only for the confusion/precision/recall table. ``` ## E. Dataset / population (unchanged) ```text 686 in-scope Candidate Setups = 594 leads + 92 follow-ons (0 duplicates). Binary fit = 571 WIN/LOSS leads (167 WIN / 404 LOSS); 18 UNRESOLVED + 5 AMBIGUOUS leads preserved and reported (never forced). feature_sha16 = 0414e401522ea4e2 (unchanged from P3-S18/P3-S19). Label contract = P3-S16 v1 (frozen), de-overlap lead-per-episode. ``` ## F. Result ```text OOS ROC-AUC : 0.5337 / 0.6341 / 0.5350 (pooled 0.5792, n=271) OOS PR-AUC : 0.3302 / 0.4224 / 0.3715 (majority priors 0.310/0.299/0.288) OOS LogLoss : 0.5815 / 0.5467 / 0.6491 (pooled 0.5895 vs prior 0.5863) OOS Brier : 0.1958 / 0.1806 / 0.2288 (pooled 0.2003 vs prior 0.1985) Balanced acc(macro): 0.5200 / 0.5000 / 0.4951 (pooled 0.5109) All folds exceed the majority baseline by ROC and PR-AUC. Pooled ranking edge is weak-positive; calibration does not materially beat the base rate. ``` ## G. Decision ```text A — STABLE WEAK SIGNAL (pre-registered rule output, applied unchanged). 3/3 folds ROC > 0.5 with > 0.02 deviation, pooled > 0.5; PR-AUC beats the majority prior in every fold; no collapse. Caveat: folds 1/3 are ~0.53 (near-noise), fold 2 is 0.634; LogLoss/Brier do not beat the constant prior; the phase evidence is a WEAK ranking effect, not a threshold-usable signal. ``` ## H. Regression / guard ```text P3-S16 20/20 | VEC 15/15 | chain parity 694==694 (0 mismatches) | S19 audit 12/12 PASS — all re-run green; committed output JSONs restored byte-identical (only generated_utc had changed). walk_forward.py + test_walk_forward.py scan CLEAN on the frozen P3-S.4/S.5 parity-absence patterns (verified 0 hits). ``` ## I. Production / ML protection ```text NO production/MQL5/F1-F4/FEATURE_CONTRACT change. Legacy MLP untouched. No TP/SL/horizon/de-overlap change. No Tickstory/Dukascopy. No LSTM/Informer/ regime/ensemble. NO test/OOS feedback into model config. ``` ## J. Next-session authorization boundary ```text P3-S.21 is NOT started automatically. Owner decision required (directs future phase based on the A decision): A (STABLE WEAK) -> feature-mechanism + calibration diagnostics, or restrained nonlinear model ONLY with separate authorization, B (if re-read as unstable) -> population heterogeneity / label-exit diagnostic (descriptive, not a contract change), C/D/E -> specific investigation path per brief section 31. No model escalation, no feature/label/TP/SL/horizon change, no Tickstory/ Dukascopy, and no production/deployment without a new authorization. ``` --- *End P3-S.20 session handover. Verdict: A — STABLE WEAK SIGNAL. The weak logistic ranking edge SURVIVES pre-registered temporal OOS testing in all three folds (pooled ROC 0.579, PR-AUC beats majority every fold) but with a small effect size, near-noise folds, and no calibration improvement — a consistent-but-weak, honestly-reported result. Next phase requires a separate owner decision.*