SniperGold_ML/docs/SESSION_HANDOVER_2026-08-24_P3_S20_WALK_FORWARD.md

5.3 KiB

SESSION HANDOVER — P3-S.20 PRE-REGISTERED WALK-FORWARD BASELINE

Date       : 2026-08-24
Session    : P3-S.20 — Pre-Registered Expanding-Window Walk-Forward Baseline
Status     : COMPLETE
Verdict    : A — STABLE WEAK SIGNAL (pre-registered decision rule output)
P3-S.21    : NOT STARTED (owner decision required; recommended next depends
             on the decision: calibration/feature-mechanism diagnostics,
             restrained-nonlinear only with separate authorization)
Scope      : VALIDATION ONLY. LOGISTIC ONLY. No tree/boost/MLP/LSTM/Informer/
             regime, no feature/label/TP/SL/horizon change, no threshold/HP
             tuning, no fold redesign after results, no Tickstory/Dukascopy.

A. Session

P3-S.20 — Pre-Registered Expanding-Window Walk-Forward Baseline.
Answer to the phase question: does the weak Logistic signal from P3-S18
(test ROC-AUC ~0.609) survive predefined temporal out-of-sample testing?
RESULT: WEAKLY YES — three pre-registered OOS folds are all above the
majority baseline (ROC 0.534/0.634/0.535), PR-AUC beats the prior in all
three, pooled OOS ROC-AUC 0.579 (n=271). Decision label: A — STABLE WEAK
SIGNAL, with honest caveats (folds 1/3 near-noise; calibration not improved
over the base rate; soft ranking edge, not a 0.5-threshold classifier).

B. Starting Checkpoint

84b0e56fc1e6ba2be503f462c4ef1a58e0830915 (P3-S19 close).
VERIFIED: local == origin/main, branch main, working tree CLEAN,
origin = forge.mql5.io/chiki2bum2/SniperGold_ML.git.
Handover read first: SESSION_HANDOVER_2026-08-24_P3_S19_DATASET_FEATURE_AUDIT.md.

C. Final Forge HEAD

P3_S20_FINAL_SHA : <recorded after push>  (main == origin/main, CLEAN)

D. Pre-registered design (frozen before final OOS metrics)

Expanding-window temporal walk-forward, 3 folds, chronologically sorted
571 WIN/LOSS binary lead rows:
  F1 train [0:300]  OOS [300:395] (n=95, bars 99568..134069)
  F2 train [0:395]  OOS [395:490] (n=95, bars 134541..167070)
  F3 train [0:490]  OOS [490:571] (n=81, bars 167391..193899)
Purge gaps 555/472/321 bars (>> H=16; asserted).
Model: LogisticRegression(C=1.0, L2, max_iter=5000, seed 42) identical to
P3-S18; StandardScaler fit on train only. Majority (constant-prior from
train prevalence) comparator per fold. Quality gate OOS>=50, >=20 WIN,
>=20 LOSS -> all folds PASS (no LOW_STATISTICAL_POWER). Threshold 0.5 only
for the confusion/precision/recall table.

E. Dataset / population (unchanged)

686 in-scope Candidate Setups = 594 leads + 92 follow-ons (0 duplicates).
Binary fit = 571 WIN/LOSS leads (167 WIN / 404 LOSS); 18 UNRESOLVED + 5
AMBIGUOUS leads preserved and reported (never forced).
feature_sha16 = 0414e401522ea4e2 (unchanged from P3-S18/P3-S19).
Label contract = P3-S16 v1 (frozen), de-overlap lead-per-episode.

F. Result

OOS ROC-AUC        : 0.5337 / 0.6341 / 0.5350   (pooled 0.5792, n=271)
OOS PR-AUC         : 0.3302 / 0.4224 / 0.3715   (majority priors 0.310/0.299/0.288)
OOS LogLoss        : 0.5815 / 0.5467 / 0.6491   (pooled 0.5895 vs prior 0.5863)
OOS Brier          : 0.1958 / 0.1806 / 0.2288   (pooled 0.2003 vs prior 0.1985)
Balanced acc(macro): 0.5200 / 0.5000 / 0.4951 (pooled 0.5109)
All folds exceed the majority baseline by ROC and PR-AUC. Pooled ranking
edge is weak-positive; calibration does not materially beat the base rate.

G. Decision

A — STABLE WEAK SIGNAL (pre-registered rule output, applied unchanged).
3/3 folds ROC > 0.5 with > 0.02 deviation, pooled > 0.5; PR-AUC beats the
majority prior in every fold; no collapse. Caveat: folds 1/3 are ~0.53
(near-noise), fold 2 is 0.634; LogLoss/Brier do not beat the constant prior;
the phase evidence is a WEAK ranking effect, not a threshold-usable signal.

H. Regression / guard

P3-S16 20/20 | VEC 15/15 | chain parity 694==694 (0 mismatches) |
S19 audit 12/12 PASS — all re-run green; committed output JSONs restored
byte-identical (only generated_utc had changed).
walk_forward.py + test_walk_forward.py scan CLEAN on the frozen P3-S.4/S.5
parity-absence patterns (verified 0 hits).

I. Production / ML protection

NO production/MQL5/F1-F4/FEATURE_CONTRACT change. Legacy MLP untouched.
No TP/SL/horizon/de-overlap change. No Tickstory/Dukascopy. No LSTM/Informer/
regime/ensemble. NO test/OOS feedback into model config.

J. Next-session authorization boundary

P3-S.21 is NOT started automatically. Owner decision required (directs
future phase based on the A decision):
  A (STABLE WEAK) -> feature-mechanism + calibration diagnostics, or
     restrained nonlinear model ONLY with separate authorization,
  B (if re-read as unstable) -> population heterogeneity / label-exit
     diagnostic (descriptive, not a contract change),
  C/D/E -> specific investigation path per brief section 31.
No model escalation, no feature/label/TP/SL/horizon change, no Tickstory/
Dukascopy, and no production/deployment without a new authorization.

End P3-S.20 session handover. Verdict: A — STABLE WEAK SIGNAL. The weak logistic ranking edge SURVIVES pre-registered temporal OOS testing in all three folds (pooled ROC 0.579, PR-AUC beats majority every fold) but with a small effect size, near-noise folds, and no calibration improvement — a consistent-but-weak, honestly-reported result. Next phase requires a separate owner decision.