SniperGold_ML/docs/P3_S24_RESEARCH_VALUE_ASSESSMENT.md

14 KiB

P3-S24 — RESEARCH VALUE ASSESSMENT & SIGNAL RECOVERY STRATEGY

Date       : 2026-08-26
Session    : P3-S24 — RESEARCH VALUE ASSESSMENT (ANALYSIS ONLY)
Status     : COMPLETE
Policy     : docs/CODE_VERIFICATION_AND_SILENT_BUG_POLICY_v1.md
Classification : D — CURRENT EDGE NOT DEMONSTRATED
Next       : P3-S25 = NOT STARTED (owner authorization required)
Predecessor: docs/SESSION_HANDOVER_2026-08-26_P3_S23_CORRECTED_POPULATION_ML_BASELINE.md
Namespace  : ml/p3/p3_s24_research_value_assessment/

Objective

Determine the most scientifically valuable next research direction AFTER the observation that the corrected (UTC-clock M30) Candidate Setup population does NOT reproduce the weak logistic ranking signal seen on the frozen population (P3-S20 pooled ROC-AUC 0.5792 → P3-S23 corrected pooled ROC-AUC 0.5235).

This is NOT a model-optimization phase. The research principle is increase information quality, not model complexity. Output = the highest-value next research investment.


0. Preflight (mandatory — all PASS)

Newest handover used (Git ancestry): docs/SESSION_HANDOVER_2026-08-26_P3_S23_...
Read completely (policy + required evidence):
  CODE_VERIFICATION_AND_SILENT_BUG_POLICY_v1.md
  P3_S23_CORRECTED_POPULATION_ML_BASELINE.md
  P3_S22_4_CORRECTED_POPULATION_VALIDATION.md
  P3_S22_3_RESEARCH_M30_REPAIR.md
  P3_S21_R_RETROSPECTIVE_SILENT_BUG_VERIFICATION.md
  P3_S21_2_CALIBRATION_REPORT.md
  P3_S21_1_FEATURE_CONTRIBUTION_AUDIT.md
  P3_S20_WALK_FORWARD_BASELINE.md
  P3_S19_DATASET_FEATURE_AUDIT.md  (+ P3_S18, P3_S18A, P3_S16 label contract,
     P3_S23 pre-registration, S23 machine evidence JSON)

Git reconciled at start (before any work):
  local HEAD == origin/main == 05ba7762fe4e17c1cbf9bf8bebdbbb1347209d24
  branch main ; working tree CLEAN ; no stash ; no untracked files

Phase 1 — Research Diagnosis Tree

The four candidate causes are evaluated on the CORRECTED 579 binary-fit population (175 WIN / 404 LOSS) and the P3-S23 pooled OOS (279; 87 WIN).

Hypothesis A — Sample Size Limitation

Question: is 579 binary-fit observations insufficient to detect a small but real edge?

Quantified (Hanley-McNeil variance + normal-approximation power, α=0.05 two-sided; see §Phase 3):

Detect AUC total N, balanced (80%/90% power) total N @ current ratio 1:2.31 (80%/90%)
0.55 1,036 / 1,386 1,261 / 1,687
0.58 398 / 532 493 / 658
0.60 250 / 336 314 / 420

At the current sample, power to reject AUC=0.5 is:

True AUC power @ binary fit (175/404) power @ pooled OOS (87/192)
0.55 0.476 0.264
0.58 0.861 0.568
0.60 0.969 0.763

Smallest AUC detectable at 80% power: ≈0.574 at the full fit, ≈0.605 at the pooled-OOS evaluation. The observed corrected point estimate (0.5235) is below even the detectable boundary at the fit, and its 95% CI [0.450, 0.592] is ~±0.074 wide and includes 0.5.

Verdict: Sample size is a real second-order limiter (a 0.55-level, real edge is far from resolvable; power targets ≈ 1,050–1,700 observations). But sample size alone does NOT explain why the corrected point estimate (0.5235) is at/under noise and why the frozen→corrected change moved it −0.056. The observed residual is indistinguishable from noise.

Hypothesis B — Population Heterogeneity

Question: is the setup population mixing different mechanisms?

Descriptive diagnostic (existing frozen dimensions only, corrected 579 rows; no new features, no model):

Dimension (bin) WIN-rate spread chi-square p
direction (short/long) 0.314 vs 0.290 → 0.024 0.522
zone position (tertiles) 0.254–0.342 0.249
CHoCH age (fresh→old) 0.333→0.206 0.265
CHoCH latency (tertiles) 0.291–0.317 0.941
liquidity / sweep age (tertiles) 0.281–0.325 0.838
volatility / ATR (tertiles) 0.254–0.344 0.240
calendar year 0.151–0.419 0.024*
  • No setup family (direction, zone position, CHoCH timing, liquidity behaviour, volatility scale) shows a statistically resolvable WIN-rate difference. The population is NOT mixing meaningfully different, separable mechanisms along the frozen structural dimensions.
  • The only significant heterogeneity is temporal: the outcome WIN base rate itself drifts strongly by year (2018 ≈ 0.15 → 2026 ≈ 0.42, p=0.024). This is regime/population non-stationarity, not a "different setup families" story, and it is consistent with P3-S23's fold-3 collapse and period dependence.

Verdict: Population heterogeneity is real but temporal/regime-based, not structural family-mixing. Redesigning the setup mechanism to "separate families" is not supported by this diagnostic.

Hypothesis C — Label / Outcome Information Limit

Question: is the current label contract too noisy?

Findings from P3-S16 v1 + P3-S18A (unchanged, read-only):

  • Censoring is low: UNRESOLVED leads ≈ 3% (25/610 corrected); AMBIGUOUS (same-bar TP+SL) ≈ 1% (6/610). All corrected UNRESOLVED are timeouts.
  • First-hit resolution is deterministic; survival says SL resolves ~2x faster than TP (median 2 vs 3 bars), but this is a frequency fact, not a leak.
  • Horizon H=16 captures >98% of outcomes; it is a label window (lifetime-based), distinct from runtime validity, and was approved as v1.

Verdict: Label noise/censoring is NOT a first-order limiter. The contract was semantically validated and the ambiguous/uncensored share is small. What IS mutable is the base rate of WIN/LOSS over time (see Hypothesis B), which is a regime effect, not label noise. No label v2 is required to answer the strategic question.

Hypothesis D — Genuine Absence of Predictive Edge

Question: does the corrected evidence indicate the current setup definition carries little predictive information?

Evidence on the corrected population (P3-S23, independently verified):

  • pooled OOS ROC-AUC 0.5235, 95% CI [0.450, 0.592] ⊇ 0.5;
  • fold 3 collapses below random (0.464) while folds 1–2 are weak-positive (0.539/0.538) ⇒ period-instanced, not stable;
  • LogLoss/Brier are worse than the constant prior (0.647 vs 0.621; 0.226 vs 0.215) ⇒ not even a usable probability output;
  • the frozen→corrected change moved pooled AUC −0.056 → 0.5235, i.e. the earlier "signal" was substantially a construction artifact.

Verdict: On the validated corrected population, a stable predictive edge is NOT demonstrated for the current setup definition. This is the honest, evidence-led conclusion; it is NOT forced negative — it is what the corrected numbers say.


Phase 2 — Value Ranking of Future Research Options

# Option Expected value Rationale
1 Increase dataset size (period/instruments) MEDIUM Raises power (needed to see 0.55–0.58), but alone cannot recreate an edge the corrected pop does not show; multi-instrument changes the traded mechanism.
2 External data validation (Tickstory/Dukascopy/second feed) HIGH Directly answers broker-feed bias, M30 UTC-clock basis-parity, and population robustness — the largest residual unknown behind S22.3–S23. Lowest cost, highest information.
3 Setup mechanism redesign MEDIUM (conditional/deferred) Hypothesis B shows no structural family-mix; redesign is premature until a replicable basis signal exists.
4 Advanced ML LOW now Nonlinear models already overfit on this sample (P3-S18 tree/boost/MLP OOS collapse), calibration didn't help, corrected baseline adds no signal. Defer.

Phase 3 — Required Quantitative Analysis

Sample Power Analysis (Hypothesis A numerals)

Current: n=579 binary-fit (175 WIN / 404 LOSS); pooled OOS n=279 (87/192). Method: Hanley-McNeil AUC variance + normal-approx power vs A=0.5 (two-sided α=0.05). Full table in output/p3_s24_power.json.

Minimum sample recommendation (reliable detection, 80%/90% power):

Target AUC balanced total @ current LOSS:WIN ratio
0.55 ~1,036 / ~1,386 ~1,261 / ~1,687
0.58 ~398 / ~532 ~493 / ~658
0.60 ~250 / ~336 ~314 / ~420

Practical recommendation for this framework (current 1:2.3 class ratio): ≈1,300–1,700 independent WIN/LOSS leads to resolve a 0.55-level edge; ≈500–660 for 0.58; ≈320–420 for 0.60. For a tight ±0.03 CI near noise: ≈1,700 leads (current ratio) or ≈1,420 (balanced). The current 579-row population is roughly half to one-third of what a 0.55 detection needs.

Evidence Matrix (Finding | Evidence | Confidence | Implication)

Phase Finding Evidence Confidence Implication
P3-S18 Baseline INCONCLUSIVE / too small; nonlinear overfits OOS logistic t=0.609; tree 0.513, boost 0.580, mlp 0.595; 115 test Datasets/features verified (S19); single split Capacity above logistic adds nothing on this n
P3-S20 Frozen-pop walk-forward pooled 0.5792, 3/3 folds>0.5 pooled 0.5792/PR 0.342 vs prior 0.273 INDEPENDENTLY VERIFIED (S21.R) Weak ranking signal existed on the OLD (faulty) population
P3-S21.1 ZONE position + short context are stable leading coefs; ~8 effective dims std-coef zone 0.138, direction 0.082; heavy collinearity Reproduced; descriptive If real, signal rides on zone geometry + context, low dims
P3-S21.2 Ranking exists; calibration does not improve raw 0.5895/0.2003; isotonic worse; prior best VERIFIED No usable/actionable probability output
P3-S22.3 M30 construction defect; 686→707, 202/223 drifted oracle 100%; 6 mutations; 0-mismatch VERIFIED A hidden data-pipeline factor, not the model, drove much of the signal
P3-S22.4 Corrected pop + contracts validated 707=610+97; leakage 8/8; schema const VERIFIED Clean deterministic basis for S23
P3-S23 Corrected pooled 0.5235, fold3 0.464, cal worse CI [0.450,0.592]; 8/8 mutations VERIFIED Edge does NOT reproduce as stable on the corrected population

Research Decision Matrix (Rank)

Research direction Expected info gain Cost Priority
Option 2 — External data validation HIGH LOW–MEDIUM 1
Option 1 — Dataset expansion MEDIUM LOW–MEDIUM 2
Option 3 — Setup mechanism redesign MEDIUM (deferred) HIGH 3
Option 4 — Advanced ML LOW MEDIUM–HIGH 4

Final Decision Classification

D — CURRENT EDGE NOT DEMONSTRATED

Basis (all independently verified, corrected population):

  • pooled OOS ROC-AUC 0.5235 with 95% CI [0.450, 0.592] containing 0.5;
  • fold-3 collapse below random (0.464) ⇒ period-instanced;
  • LogLoss/Brier worse than the constant prior ⇒ no usable probability;
  • corrected construction (S22.3) removed the bulk of the earlier apparent edge.

Rejected options and reasons:

  • A (dataset expansion) — rejected as PRIMARY. Sample size is a genuine second-order limit (power to detect 0.55 is ~48%), but the corrected point estimate (0.5235) is below the detectable threshold and the frozen→corrected change (−0.056) shows the earlier edge was largely artifact — more samples on the same definition/feed will not conjure an undemonstrated edge.
  • B (mechanism research) — rejected. No structural family heterogeneity exists (all frozen-dimension splits p>0.2); only temporal base-rate drift is significant. Separating "set-up families" is not supported.
  • C (label/contract review) — rejected. Label v1 was semantically validated; censoring/ambiguity are small (≈3%/1%). Outcome base-rate drift is a regime effect, not label noise.
  • E (data/pipeline defect) — resolved, not current. The M30 construction defect was corrected and verified in S22.3; however the basis-parity of the corrected build against the broker feed was never independently re-established — this becomes the top validation target (Option 2), not an assumption.

Highest-value next action (Option 2): independently validate the broker-feed / UTC-clock M30 basis-parity and the corrected population with external data — because it (a) resolves the single largest residual unknown, and (b) is the only step that can separate "genuine absence of edge" (D) from a "residual feed/construction artifact" (E). It is low-cost, HIGH-information, and NOT auto-imported in this phase. Only after a validated basis should dataset expansion to the S24 power targets proceed.


Hard constraints (this phase honoured)

ML training / tuning        : NONE
New features                : NONE (descriptive bins of existing features only)
Label / TP / SL / horizon   : NONE changed
MQL5 / production           : NONE changed
External data               : NONE imported (MQL5/Files assessed only)
Calibration / nonlinear     : NONE
Deployment / trading        : NONE
P3-S25                      : NOT STARTED (owner authorization required)

Artifacts

Namespace : ml/p3/p3_s24_research_value_assessment/
  p3_s24_config.py / p3_s24_power.py / p3_s24_heterogeneity.py /
  p3_s24_evidence.py / p3_s24_run_main.py / README.md
output    : p3_s24_power.json, p3_s24_heterogeneity.json,
            p3_s24_evidence_matrix.json, p3_s24_decision_matrix.json,
            p3_s24_classification.json, p3_s24_manifest.json
Docs      : this file, docs/P3_S24_NEXT_RESEARCH_PRIORITY.md,
            docs/SESSION_HANDOVER_2026-08-26_P3_S24_RESEARCH_VALUE_ASSESSMENT.md
Reproducibility: two runs byte-identical except generated_utc (aggregate
            SHA256 9912f7ab887bd0ec… identical across runs)

End of P3-S24 research value assessment. Classification D — CURRENT EDGE NOT DEMONSTRATED; highest-value next action = independent/external data validation of the corrected basis, then dataset expansion to the computed power targets.