14 KiB
P3-S24 — RESEARCH VALUE ASSESSMENT & SIGNAL RECOVERY STRATEGY
Date : 2026-08-26
Session : P3-S24 — RESEARCH VALUE ASSESSMENT (ANALYSIS ONLY)
Status : COMPLETE
Policy : docs/CODE_VERIFICATION_AND_SILENT_BUG_POLICY_v1.md
Classification : D — CURRENT EDGE NOT DEMONSTRATED
Next : P3-S25 = NOT STARTED (owner authorization required)
Predecessor: docs/SESSION_HANDOVER_2026-08-26_P3_S23_CORRECTED_POPULATION_ML_BASELINE.md
Namespace : ml/p3/p3_s24_research_value_assessment/
Objective
Determine the most scientifically valuable next research direction AFTER the observation that the corrected (UTC-clock M30) Candidate Setup population does NOT reproduce the weak logistic ranking signal seen on the frozen population (P3-S20 pooled ROC-AUC 0.5792 → P3-S23 corrected pooled ROC-AUC 0.5235).
This is NOT a model-optimization phase. The research principle is increase information quality, not model complexity. Output = the highest-value next research investment.
0. Preflight (mandatory — all PASS)
Newest handover used (Git ancestry): docs/SESSION_HANDOVER_2026-08-26_P3_S23_...
Read completely (policy + required evidence):
CODE_VERIFICATION_AND_SILENT_BUG_POLICY_v1.md
P3_S23_CORRECTED_POPULATION_ML_BASELINE.md
P3_S22_4_CORRECTED_POPULATION_VALIDATION.md
P3_S22_3_RESEARCH_M30_REPAIR.md
P3_S21_R_RETROSPECTIVE_SILENT_BUG_VERIFICATION.md
P3_S21_2_CALIBRATION_REPORT.md
P3_S21_1_FEATURE_CONTRIBUTION_AUDIT.md
P3_S20_WALK_FORWARD_BASELINE.md
P3_S19_DATASET_FEATURE_AUDIT.md (+ P3_S18, P3_S18A, P3_S16 label contract,
P3_S23 pre-registration, S23 machine evidence JSON)
Git reconciled at start (before any work):
local HEAD == origin/main == 05ba7762fe4e17c1cbf9bf8bebdbbb1347209d24
branch main ; working tree CLEAN ; no stash ; no untracked files
Phase 1 — Research Diagnosis Tree
The four candidate causes are evaluated on the CORRECTED 579 binary-fit population (175 WIN / 404 LOSS) and the P3-S23 pooled OOS (279; 87 WIN).
Hypothesis A — Sample Size Limitation
Question: is 579 binary-fit observations insufficient to detect a small but real edge?
Quantified (Hanley-McNeil variance + normal-approximation power, α=0.05 two-sided; see §Phase 3):
| Detect AUC | total N, balanced (80%/90% power) | total N @ current ratio 1:2.31 (80%/90%) |
|---|---|---|
| 0.55 | 1,036 / 1,386 | 1,261 / 1,687 |
| 0.58 | 398 / 532 | 493 / 658 |
| 0.60 | 250 / 336 | 314 / 420 |
At the current sample, power to reject AUC=0.5 is:
| True AUC | power @ binary fit (175/404) | power @ pooled OOS (87/192) |
|---|---|---|
| 0.55 | 0.476 | 0.264 |
| 0.58 | 0.861 | 0.568 |
| 0.60 | 0.969 | 0.763 |
Smallest AUC detectable at 80% power: ≈0.574 at the full fit, ≈0.605 at the pooled-OOS evaluation. The observed corrected point estimate (0.5235) is below even the detectable boundary at the fit, and its 95% CI [0.450, 0.592] is ~±0.074 wide and includes 0.5.
Verdict: Sample size is a real second-order limiter (a 0.55-level, real edge is far from resolvable; power targets ≈ 1,050–1,700 observations). But sample size alone does NOT explain why the corrected point estimate (0.5235) is at/under noise and why the frozen→corrected change moved it −0.056. The observed residual is indistinguishable from noise.
Hypothesis B — Population Heterogeneity
Question: is the setup population mixing different mechanisms?
Descriptive diagnostic (existing frozen dimensions only, corrected 579 rows; no new features, no model):
| Dimension (bin) | WIN-rate spread | chi-square p |
|---|---|---|
| direction (short/long) | 0.314 vs 0.290 → 0.024 | 0.522 |
| zone position (tertiles) | 0.254–0.342 | 0.249 |
| CHoCH age (fresh→old) | 0.333→0.206 | 0.265 |
| CHoCH latency (tertiles) | 0.291–0.317 | 0.941 |
| liquidity / sweep age (tertiles) | 0.281–0.325 | 0.838 |
| volatility / ATR (tertiles) | 0.254–0.344 | 0.240 |
| calendar year | 0.151–0.419 | 0.024* |
- No setup family (direction, zone position, CHoCH timing, liquidity behaviour, volatility scale) shows a statistically resolvable WIN-rate difference. The population is NOT mixing meaningfully different, separable mechanisms along the frozen structural dimensions.
- The only significant heterogeneity is temporal: the outcome WIN base rate itself drifts strongly by year (2018 ≈ 0.15 → 2026 ≈ 0.42, p=0.024). This is regime/population non-stationarity, not a "different setup families" story, and it is consistent with P3-S23's fold-3 collapse and period dependence.
Verdict: Population heterogeneity is real but temporal/regime-based, not structural family-mixing. Redesigning the setup mechanism to "separate families" is not supported by this diagnostic.
Hypothesis C — Label / Outcome Information Limit
Question: is the current label contract too noisy?
Findings from P3-S16 v1 + P3-S18A (unchanged, read-only):
- Censoring is low: UNRESOLVED leads ≈ 3% (25/610 corrected); AMBIGUOUS (same-bar TP+SL) ≈ 1% (6/610). All corrected UNRESOLVED are timeouts.
- First-hit resolution is deterministic; survival says SL resolves ~2x faster than TP (median 2 vs 3 bars), but this is a frequency fact, not a leak.
- Horizon H=16 captures >98% of outcomes; it is a label window (lifetime-based), distinct from runtime validity, and was approved as v1.
Verdict: Label noise/censoring is NOT a first-order limiter. The contract was semantically validated and the ambiguous/uncensored share is small. What IS mutable is the base rate of WIN/LOSS over time (see Hypothesis B), which is a regime effect, not label noise. No label v2 is required to answer the strategic question.
Hypothesis D — Genuine Absence of Predictive Edge
Question: does the corrected evidence indicate the current setup definition carries little predictive information?
Evidence on the corrected population (P3-S23, independently verified):
- pooled OOS ROC-AUC 0.5235, 95% CI [0.450, 0.592] ⊇ 0.5;
- fold 3 collapses below random (0.464) while folds 1–2 are weak-positive (0.539/0.538) ⇒ period-instanced, not stable;
- LogLoss/Brier are worse than the constant prior (0.647 vs 0.621; 0.226 vs 0.215) ⇒ not even a usable probability output;
- the frozen→corrected change moved pooled AUC −0.056 → 0.5235, i.e. the earlier "signal" was substantially a construction artifact.
Verdict: On the validated corrected population, a stable predictive edge is NOT demonstrated for the current setup definition. This is the honest, evidence-led conclusion; it is NOT forced negative — it is what the corrected numbers say.
Phase 2 — Value Ranking of Future Research Options
| # | Option | Expected value | Rationale |
|---|---|---|---|
| 1 | Increase dataset size (period/instruments) | MEDIUM | Raises power (needed to see 0.55–0.58), but alone cannot recreate an edge the corrected pop does not show; multi-instrument changes the traded mechanism. |
| 2 | External data validation (Tickstory/Dukascopy/second feed) | HIGH | Directly answers broker-feed bias, M30 UTC-clock basis-parity, and population robustness — the largest residual unknown behind S22.3–S23. Lowest cost, highest information. |
| 3 | Setup mechanism redesign | MEDIUM (conditional/deferred) | Hypothesis B shows no structural family-mix; redesign is premature until a replicable basis signal exists. |
| 4 | Advanced ML | LOW now | Nonlinear models already overfit on this sample (P3-S18 tree/boost/MLP OOS collapse), calibration didn't help, corrected baseline adds no signal. Defer. |
Phase 3 — Required Quantitative Analysis
Sample Power Analysis (Hypothesis A numerals)
Current: n=579 binary-fit (175 WIN / 404 LOSS); pooled OOS n=279 (87/192).
Method: Hanley-McNeil AUC variance + normal-approx power vs A=0.5 (two-sided
α=0.05). Full table in output/p3_s24_power.json.
Minimum sample recommendation (reliable detection, 80%/90% power):
| Target AUC | balanced total | @ current LOSS:WIN ratio |
|---|---|---|
| 0.55 | ~1,036 / ~1,386 | ~1,261 / ~1,687 |
| 0.58 | ~398 / ~532 | ~493 / ~658 |
| 0.60 | ~250 / ~336 | ~314 / ~420 |
Practical recommendation for this framework (current 1:2.3 class ratio): ≈1,300–1,700 independent WIN/LOSS leads to resolve a 0.55-level edge; ≈500–660 for 0.58; ≈320–420 for 0.60. For a tight ±0.03 CI near noise: ≈1,700 leads (current ratio) or ≈1,420 (balanced). The current 579-row population is roughly half to one-third of what a 0.55 detection needs.
Evidence Matrix (Finding | Evidence | Confidence | Implication)
| Phase | Finding | Evidence | Confidence | Implication |
|---|---|---|---|---|
| P3-S18 | Baseline INCONCLUSIVE / too small; nonlinear overfits OOS | logistic t=0.609; tree 0.513, boost 0.580, mlp 0.595; 115 test | Datasets/features verified (S19); single split | Capacity above logistic adds nothing on this n |
| P3-S20 | Frozen-pop walk-forward pooled 0.5792, 3/3 folds>0.5 | pooled 0.5792/PR 0.342 vs prior 0.273 | INDEPENDENTLY VERIFIED (S21.R) | Weak ranking signal existed on the OLD (faulty) population |
| P3-S21.1 | ZONE position + short context are stable leading coefs; ~8 effective dims | std-coef zone 0.138, direction 0.082; heavy collinearity | Reproduced; descriptive | If real, signal rides on zone geometry + context, low dims |
| P3-S21.2 | Ranking exists; calibration does not improve | raw 0.5895/0.2003; isotonic worse; prior best | VERIFIED | No usable/actionable probability output |
| P3-S22.3 | M30 construction defect; 686→707, 202/223 drifted | oracle 100%; 6 mutations; 0-mismatch | VERIFIED | A hidden data-pipeline factor, not the model, drove much of the signal |
| P3-S22.4 | Corrected pop + contracts validated | 707=610+97; leakage 8/8; schema const | VERIFIED | Clean deterministic basis for S23 |
| P3-S23 | Corrected pooled 0.5235, fold3 0.464, cal worse | CI [0.450,0.592]; 8/8 mutations | VERIFIED | Edge does NOT reproduce as stable on the corrected population |
Research Decision Matrix (Rank)
| Research direction | Expected info gain | Cost | Priority |
|---|---|---|---|
| Option 2 — External data validation | HIGH | LOW–MEDIUM | 1 |
| Option 1 — Dataset expansion | MEDIUM | LOW–MEDIUM | 2 |
| Option 3 — Setup mechanism redesign | MEDIUM (deferred) | HIGH | 3 |
| Option 4 — Advanced ML | LOW | MEDIUM–HIGH | 4 |
Final Decision Classification
D — CURRENT EDGE NOT DEMONSTRATED
Basis (all independently verified, corrected population):
- pooled OOS ROC-AUC 0.5235 with 95% CI [0.450, 0.592] containing 0.5;
- fold-3 collapse below random (0.464) ⇒ period-instanced;
- LogLoss/Brier worse than the constant prior ⇒ no usable probability;
- corrected construction (S22.3) removed the bulk of the earlier apparent edge.
Rejected options and reasons:
- A (dataset expansion) — rejected as PRIMARY. Sample size is a genuine second-order limit (power to detect 0.55 is ~48%), but the corrected point estimate (0.5235) is below the detectable threshold and the frozen→corrected change (−0.056) shows the earlier edge was largely artifact — more samples on the same definition/feed will not conjure an undemonstrated edge.
- B (mechanism research) — rejected. No structural family heterogeneity exists (all frozen-dimension splits p>0.2); only temporal base-rate drift is significant. Separating "set-up families" is not supported.
- C (label/contract review) — rejected. Label v1 was semantically validated; censoring/ambiguity are small (≈3%/1%). Outcome base-rate drift is a regime effect, not label noise.
- E (data/pipeline defect) — resolved, not current. The M30 construction defect was corrected and verified in S22.3; however the basis-parity of the corrected build against the broker feed was never independently re-established — this becomes the top validation target (Option 2), not an assumption.
Highest-value next action (Option 2): independently validate the broker-feed / UTC-clock M30 basis-parity and the corrected population with external data — because it (a) resolves the single largest residual unknown, and (b) is the only step that can separate "genuine absence of edge" (D) from a "residual feed/construction artifact" (E). It is low-cost, HIGH-information, and NOT auto-imported in this phase. Only after a validated basis should dataset expansion to the S24 power targets proceed.
Hard constraints (this phase honoured)
ML training / tuning : NONE
New features : NONE (descriptive bins of existing features only)
Label / TP / SL / horizon : NONE changed
MQL5 / production : NONE changed
External data : NONE imported (MQL5/Files assessed only)
Calibration / nonlinear : NONE
Deployment / trading : NONE
P3-S25 : NOT STARTED (owner authorization required)
Artifacts
Namespace : ml/p3/p3_s24_research_value_assessment/
p3_s24_config.py / p3_s24_power.py / p3_s24_heterogeneity.py /
p3_s24_evidence.py / p3_s24_run_main.py / README.md
output : p3_s24_power.json, p3_s24_heterogeneity.json,
p3_s24_evidence_matrix.json, p3_s24_decision_matrix.json,
p3_s24_classification.json, p3_s24_manifest.json
Docs : this file, docs/P3_S24_NEXT_RESEARCH_PRIORITY.md,
docs/SESSION_HANDOVER_2026-08-26_P3_S24_RESEARCH_VALUE_ASSESSMENT.md
Reproducibility: two runs byte-identical except generated_utc (aggregate
SHA256 9912f7ab887bd0ec… identical across runs)
End of P3-S24 research value assessment. Classification D — CURRENT EDGE NOT DEMONSTRATED; highest-value next action = independent/external data validation of the corrected basis, then dataset expansion to the computed power targets.