# P3-S24 — RESEARCH VALUE ASSESSMENT & SIGNAL RECOVERY STRATEGY ```text Date : 2026-08-26 Session : P3-S24 — RESEARCH VALUE ASSESSMENT (ANALYSIS ONLY) Status : COMPLETE Policy : docs/CODE_VERIFICATION_AND_SILENT_BUG_POLICY_v1.md Classification : D — CURRENT EDGE NOT DEMONSTRATED Next : P3-S25 = NOT STARTED (owner authorization required) Predecessor: docs/SESSION_HANDOVER_2026-08-26_P3_S23_CORRECTED_POPULATION_ML_BASELINE.md Namespace : ml/p3/p3_s24_research_value_assessment/ ``` ## Objective Determine the most scientifically valuable next research direction AFTER the observation that the corrected (UTC-clock M30) Candidate Setup population does NOT reproduce the weak logistic ranking signal seen on the frozen population (P3-S20 pooled ROC-AUC 0.5792 → P3-S23 corrected pooled ROC-AUC 0.5235). This is NOT a model-optimization phase. The research principle is **increase information quality, not model complexity**. Output = the highest-value next research investment. --- ## 0. Preflight (mandatory — all PASS) ```text Newest handover used (Git ancestry): docs/SESSION_HANDOVER_2026-08-26_P3_S23_... Read completely (policy + required evidence): CODE_VERIFICATION_AND_SILENT_BUG_POLICY_v1.md P3_S23_CORRECTED_POPULATION_ML_BASELINE.md P3_S22_4_CORRECTED_POPULATION_VALIDATION.md P3_S22_3_RESEARCH_M30_REPAIR.md P3_S21_R_RETROSPECTIVE_SILENT_BUG_VERIFICATION.md P3_S21_2_CALIBRATION_REPORT.md P3_S21_1_FEATURE_CONTRIBUTION_AUDIT.md P3_S20_WALK_FORWARD_BASELINE.md P3_S19_DATASET_FEATURE_AUDIT.md (+ P3_S18, P3_S18A, P3_S16 label contract, P3_S23 pre-registration, S23 machine evidence JSON) Git reconciled at start (before any work): local HEAD == origin/main == 05ba7762fe4e17c1cbf9bf8bebdbbb1347209d24 branch main ; working tree CLEAN ; no stash ; no untracked files ``` --- ## Phase 1 — Research Diagnosis Tree The four candidate causes are evaluated on the CORRECTED 579 binary-fit population (175 WIN / 404 LOSS) and the P3-S23 pooled OOS (279; 87 WIN). ### Hypothesis A — Sample Size Limitation **Question:** is 579 binary-fit observations insufficient to detect a small but real edge? **Quantified (Hanley-McNeil variance + normal-approximation power, α=0.05 two-sided; see §Phase 3):** | Detect AUC | total N, balanced (80%/90% power) | total N @ current ratio 1:2.31 (80%/90%) | |---|---|---| | 0.55 | 1,036 / 1,386 | 1,261 / 1,687 | | 0.58 | 398 / 532 | 493 / 658 | | 0.60 | 250 / 336 | 314 / 420 | At the current sample, power to reject AUC=0.5 is: | True AUC | power @ binary fit (175/404) | power @ pooled OOS (87/192) | |---|---|---| | 0.55 | 0.476 | 0.264 | | 0.58 | 0.861 | 0.568 | | 0.60 | 0.969 | 0.763 | Smallest AUC detectable at 80% power: **≈0.574 at the full fit**, **≈0.605 at the pooled-OOS evaluation**. The observed corrected point estimate (0.5235) is **below even the detectable boundary at the fit**, and its 95% CI [0.450, 0.592] is ~±0.074 wide and includes 0.5. **Verdict:** Sample size is a **real second-order limiter** (a 0.55-level, real edge is far from resolvable; power targets ≈ 1,050–1,700 observations). But sample size alone does NOT explain why the corrected point estimate (0.5235) is at/under noise *and* why the frozen→corrected change moved it −0.056. The observed residual is indistinguishable from noise. ### Hypothesis B — Population Heterogeneity **Question:** is the setup population mixing different mechanisms? **Descriptive diagnostic (existing frozen dimensions only, corrected 579 rows; no new features, no model):** | Dimension (bin) | WIN-rate spread | chi-square p | |---|---|---| | direction (short/long) | 0.314 vs 0.290 → 0.024 | 0.522 | | zone position (tertiles) | 0.254–0.342 | 0.249 | | CHoCH age (fresh→old) | 0.333→0.206 | 0.265 | | CHoCH latency (tertiles) | 0.291–0.317 | 0.941 | | liquidity / sweep age (tertiles) | 0.281–0.325 | 0.838 | | volatility / ATR (tertiles) | 0.254–0.344 | 0.240 | | **calendar year** | **0.151–0.419** | **0.024*** | - **No setup family** (direction, zone position, CHoCH timing, liquidity behaviour, volatility scale) shows a statistically resolvable WIN-rate difference. The population is **NOT** mixing meaningfully different, separable mechanisms along the frozen structural dimensions. - The **only** significant heterogeneity is **temporal**: the outcome WIN base rate itself drifts strongly by year (2018 ≈ 0.15 → 2026 ≈ 0.42, p=0.024). This is regime/population non-stationarity, not a "different setup families" story, and it is consistent with P3-S23's fold-3 collapse and period dependence. **Verdict:** Population heterogeneity is **real but temporal/regime-based**, not structural family-mixing. Redesigning the setup mechanism to "separate families" is not supported by this diagnostic. ### Hypothesis C — Label / Outcome Information Limit **Question:** is the current label contract too noisy? Findings from P3-S16 v1 + P3-S18A (unchanged, read-only): - Censoring is low: UNRESOLVED leads ≈ 3% (25/610 corrected); AMBIGUOUS (same-bar TP+SL) ≈ 1% (6/610). All corrected UNRESOLVED are timeouts. - First-hit resolution is deterministic; survival says SL resolves ~2x faster than TP (median 2 vs 3 bars), but this is a frequency fact, not a leak. - Horizon H=16 captures >98% of outcomes; it is a label window (lifetime-based), distinct from runtime validity, and was approved as v1. **Verdict:** Label **noise/censoring is NOT a first-order limiter.** The contract was semantically validated and the ambiguous/uncensored share is small. What IS mutable is the *base rate* of WIN/LOSS over time (see Hypothesis B), which is a regime effect, not label noise. No label v2 is required to answer the strategic question. ### Hypothesis D — Genuine Absence of Predictive Edge **Question:** does the corrected evidence indicate the current setup definition carries little predictive information? Evidence on the corrected population (P3-S23, independently verified): - pooled OOS ROC-AUC **0.5235**, 95% CI **[0.450, 0.592] ⊇ 0.5**; - fold 3 collapses below random (**0.464**) while folds 1–2 are weak-positive (0.539/0.538) ⇒ **period-instanced, not stable**; - LogLoss/Brier are **worse than the constant prior** (0.647 vs 0.621; 0.226 vs 0.215) ⇒ not even a usable probability output; - the frozen→corrected change moved pooled AUC **−0.056 → 0.5235**, i.e. the earlier "signal" was substantially a construction artifact. **Verdict:** On the validated corrected population, **a stable predictive edge is NOT demonstrated** for the current setup definition. This is the honest, evidence-led conclusion; it is NOT forced negative — it is what the corrected numbers say. --- ## Phase 2 — Value Ranking of Future Research Options | # | Option | Expected value | Rationale | |---|---|---|---| | 1 | Increase dataset size (period/instruments) | **MEDIUM** | Raises power (needed to see 0.55–0.58), but alone cannot recreate an edge the corrected pop does not show; multi-instrument changes the traded mechanism. | | 2 | External data validation (Tickstory/Dukascopy/second feed) | **HIGH** | Directly answers broker-feed bias, M30 UTC-clock basis-parity, and population robustness — the largest residual unknown behind S22.3–S23. Lowest cost, highest information. | | 3 | Setup mechanism redesign | **MEDIUM (conditional/deferred)** | Hypothesis B shows no structural family-mix; redesign is premature until a replicable basis signal exists. | | 4 | Advanced ML | **LOW now** | Nonlinear models already overfit on this sample (P3-S18 tree/boost/MLP OOS collapse), calibration didn't help, corrected baseline adds no signal. Defer. | --- ## Phase 3 — Required Quantitative Analysis ### Sample Power Analysis (Hypothesis A numerals) Current: n=579 binary-fit (175 WIN / 404 LOSS); pooled OOS n=279 (87/192). Method: Hanley-McNeil AUC variance + normal-approx power vs A=0.5 (two-sided α=0.05). Full table in `output/p3_s24_power.json`. **Minimum sample recommendation (reliable detection, 80%/90% power):** | Target AUC | balanced total | @ current LOSS:WIN ratio | |---|---|---| | 0.55 | ~1,036 / ~1,386 | ~1,261 / ~1,687 | | 0.58 | ~398 / ~532 | ~493 / ~658 | | 0.60 | ~250 / ~336 | ~314 / ~420 | Practical recommendation for this framework (current 1:2.3 class ratio): **≈1,300–1,700 independent WIN/LOSS leads to resolve a 0.55-level edge;** **≈500–660 for 0.58; ≈320–420 for 0.60.** For a tight ±0.03 CI near noise: **≈1,700 leads** (current ratio) or **≈1,420** (balanced). The current 579-row population is roughly **half to one-third** of what a 0.55 detection needs. ### Evidence Matrix (Finding | Evidence | Confidence | Implication) | Phase | Finding | Evidence | Confidence | Implication | |---|---|---|---|---| | P3-S18 | Baseline INCONCLUSIVE / too small; nonlinear overfits OOS | logistic t=0.609; tree 0.513, boost 0.580, mlp 0.595; 115 test | Datasets/features verified (S19); single split | Capacity above logistic adds nothing on this n | | P3-S20 | Frozen-pop walk-forward pooled 0.5792, 3/3 folds>0.5 | pooled 0.5792/PR 0.342 vs prior 0.273 | **INDEPENDENTLY VERIFIED** (S21.R) | Weak ranking signal existed on the OLD (faulty) population | | P3-S21.1 | ZONE position + short context are stable leading coefs; ~8 effective dims | std-coef zone 0.138, direction 0.082; heavy collinearity | Reproduced; descriptive | If real, signal rides on zone geometry + context, low dims | | P3-S21.2 | Ranking exists; calibration does not improve | raw 0.5895/0.2003; isotonic worse; prior best | **VERIFIED** | No usable/actionable probability output | | P3-S22.3 | M30 construction defect; 686→707, 202/223 drifted | oracle 100%; 6 mutations; 0-mismatch | **VERIFIED** | A hidden data-pipeline factor, not the model, drove much of the signal | | P3-S22.4 | Corrected pop + contracts validated | 707=610+97; leakage 8/8; schema const | **VERIFIED** | Clean deterministic basis for S23 | | P3-S23 | Corrected pooled 0.5235, fold3 0.464, cal worse | CI [0.450,0.592]; 8/8 mutations | **VERIFIED** | Edge does NOT reproduce as stable on the corrected population | ### Research Decision Matrix (Rank) | Research direction | Expected info gain | Cost | Priority | |---|---|---|---| | **Option 2 — External data validation** | **HIGH** | LOW–MEDIUM | **1** | | Option 1 — Dataset expansion | MEDIUM | LOW–MEDIUM | 2 | | Option 3 — Setup mechanism redesign | MEDIUM (deferred) | HIGH | 3 | | Option 4 — Advanced ML | LOW | MEDIUM–HIGH | 4 | --- ## Final Decision Classification ```text D — CURRENT EDGE NOT DEMONSTRATED ``` Basis (all independently verified, corrected population): - pooled OOS ROC-AUC 0.5235 with 95% CI **[0.450, 0.592]** containing 0.5; - fold-3 collapse below random (0.464) ⇒ period-instanced; - LogLoss/Brier worse than the constant prior ⇒ no usable probability; - corrected construction (S22.3) removed the bulk of the earlier apparent edge. Rejected options and reasons: - **A (dataset expansion) — rejected as PRIMARY.** Sample size is a genuine second-order limit (power to detect 0.55 is ~48%), but the corrected point estimate (0.5235) is *below* the detectable threshold and the frozen→corrected change (−0.056) shows the earlier edge was largely artifact — more samples on the same definition/feed will not conjure an undemonstrated edge. - **B (mechanism research) — rejected.** No structural family heterogeneity exists (all frozen-dimension splits p>0.2); only temporal base-rate drift is significant. Separating "set-up families" is not supported. - **C (label/contract review) — rejected.** Label v1 was semantically validated; censoring/ambiguity are small (≈3%/1%). Outcome base-rate drift is a regime effect, not label noise. - **E (data/pipeline defect) — resolved, not current.** The M30 construction defect was corrected and verified in S22.3; however the **basis-parity of the corrected build against the broker feed was never independently re-established** — this becomes the top validation target (Option 2), not an assumption. Highest-value next action (Option 2): **independently validate the broker-feed / UTC-clock M30 basis-parity and the corrected population** with external data — because it (a) resolves the single largest residual unknown, and (b) is the only step that can separate "genuine absence of edge" (D) from a "residual feed/construction artifact" (E). It is low-cost, HIGH-information, and NOT auto-imported in this phase. Only after a validated basis should dataset expansion to the S24 power targets proceed. --- ## Hard constraints (this phase honoured) ```text ML training / tuning : NONE New features : NONE (descriptive bins of existing features only) Label / TP / SL / horizon : NONE changed MQL5 / production : NONE changed External data : NONE imported (MQL5/Files assessed only) Calibration / nonlinear : NONE Deployment / trading : NONE P3-S25 : NOT STARTED (owner authorization required) ``` ## Artifacts ```text Namespace : ml/p3/p3_s24_research_value_assessment/ p3_s24_config.py / p3_s24_power.py / p3_s24_heterogeneity.py / p3_s24_evidence.py / p3_s24_run_main.py / README.md output : p3_s24_power.json, p3_s24_heterogeneity.json, p3_s24_evidence_matrix.json, p3_s24_decision_matrix.json, p3_s24_classification.json, p3_s24_manifest.json Docs : this file, docs/P3_S24_NEXT_RESEARCH_PRIORITY.md, docs/SESSION_HANDOVER_2026-08-26_P3_S24_RESEARCH_VALUE_ASSESSMENT.md Reproducibility: two runs byte-identical except generated_utc (aggregate SHA256 9912f7ab887bd0ec… identical across runs) ``` *End of P3-S24 research value assessment. Classification D — CURRENT EDGE NOT DEMONSTRATED; highest-value next action = independent/external data validation of the corrected basis, then dataset expansion to the computed power targets.*