SniperGold_ML/docs/P3_S24_RESEARCH_VALUE_ASSESSMENT.md

283 lines
14 KiB
Markdown

# P3-S24 — RESEARCH VALUE ASSESSMENT & SIGNAL RECOVERY STRATEGY
```text
Date : 2026-08-26
Session : P3-S24 — RESEARCH VALUE ASSESSMENT (ANALYSIS ONLY)
Status : COMPLETE
Policy : docs/CODE_VERIFICATION_AND_SILENT_BUG_POLICY_v1.md
Classification : D — CURRENT EDGE NOT DEMONSTRATED
Next : P3-S25 = NOT STARTED (owner authorization required)
Predecessor: docs/SESSION_HANDOVER_2026-08-26_P3_S23_CORRECTED_POPULATION_ML_BASELINE.md
Namespace : ml/p3/p3_s24_research_value_assessment/
```
## Objective
Determine the most scientifically valuable next research direction AFTER the
observation that the corrected (UTC-clock M30) Candidate Setup population does
NOT reproduce the weak logistic ranking signal seen on the frozen population
(P3-S20 pooled ROC-AUC 0.5792 → P3-S23 corrected pooled ROC-AUC 0.5235).
This is NOT a model-optimization phase. The research principle is **increase
information quality, not model complexity**. Output = the highest-value next
research investment.
---
## 0. Preflight (mandatory — all PASS)
```text
Newest handover used (Git ancestry): docs/SESSION_HANDOVER_2026-08-26_P3_S23_...
Read completely (policy + required evidence):
CODE_VERIFICATION_AND_SILENT_BUG_POLICY_v1.md
P3_S23_CORRECTED_POPULATION_ML_BASELINE.md
P3_S22_4_CORRECTED_POPULATION_VALIDATION.md
P3_S22_3_RESEARCH_M30_REPAIR.md
P3_S21_R_RETROSPECTIVE_SILENT_BUG_VERIFICATION.md
P3_S21_2_CALIBRATION_REPORT.md
P3_S21_1_FEATURE_CONTRIBUTION_AUDIT.md
P3_S20_WALK_FORWARD_BASELINE.md
P3_S19_DATASET_FEATURE_AUDIT.md (+ P3_S18, P3_S18A, P3_S16 label contract,
P3_S23 pre-registration, S23 machine evidence JSON)
Git reconciled at start (before any work):
local HEAD == origin/main == 05ba7762fe4e17c1cbf9bf8bebdbbb1347209d24
branch main ; working tree CLEAN ; no stash ; no untracked files
```
---
## Phase 1 — Research Diagnosis Tree
The four candidate causes are evaluated on the CORRECTED 579 binary-fit
population (175 WIN / 404 LOSS) and the P3-S23 pooled OOS (279; 87 WIN).
### Hypothesis A — Sample Size Limitation
**Question:** is 579 binary-fit observations insufficient to detect a small
but real edge?
**Quantified (Hanley-McNeil variance + normal-approximation power, α=0.05
two-sided; see §Phase 3):**
| Detect AUC | total N, balanced (80%/90% power) | total N @ current ratio 1:2.31 (80%/90%) |
|---|---|---|
| 0.55 | 1,036 / 1,386 | 1,261 / 1,687 |
| 0.58 | 398 / 532 | 493 / 658 |
| 0.60 | 250 / 336 | 314 / 420 |
At the current sample, power to reject AUC=0.5 is:
| True AUC | power @ binary fit (175/404) | power @ pooled OOS (87/192) |
|---|---|---|
| 0.55 | 0.476 | 0.264 |
| 0.58 | 0.861 | 0.568 |
| 0.60 | 0.969 | 0.763 |
Smallest AUC detectable at 80% power: **≈0.574 at the full fit**, **≈0.605 at
the pooled-OOS evaluation**. The observed corrected point estimate (0.5235) is
**below even the detectable boundary at the fit**, and its 95% CI
[0.450, 0.592] is ~±0.074 wide and includes 0.5.
**Verdict:** Sample size is a **real second-order limiter** (a 0.55-level,
real edge is far from resolvable; power targets ≈ 1,050–1,700 observations).
But sample size alone does NOT explain why the corrected point estimate (0.5235)
is at/under noise *and* why the frozen→corrected change moved it −0.056. The
observed residual is indistinguishable from noise.
### Hypothesis B — Population Heterogeneity
**Question:** is the setup population mixing different mechanisms?
**Descriptive diagnostic (existing frozen dimensions only, corrected 579 rows;
no new features, no model):**
| Dimension (bin) | WIN-rate spread | chi-square p |
|---|---|---|
| direction (short/long) | 0.314 vs 0.290 → 0.024 | 0.522 |
| zone position (tertiles) | 0.254–0.342 | 0.249 |
| CHoCH age (fresh→old) | 0.333→0.206 | 0.265 |
| CHoCH latency (tertiles) | 0.291–0.317 | 0.941 |
| liquidity / sweep age (tertiles) | 0.281–0.325 | 0.838 |
| volatility / ATR (tertiles) | 0.254–0.344 | 0.240 |
| **calendar year** | **0.151–0.419** | **0.024*** |
- **No setup family** (direction, zone position, CHoCH timing, liquidity
behaviour, volatility scale) shows a statistically resolvable WIN-rate
difference. The population is **NOT** mixing meaningfully different,
separable mechanisms along the frozen structural dimensions.
- The **only** significant heterogeneity is **temporal**: the outcome WIN
base rate itself drifts strongly by year (2018 ≈ 0.15 → 2026 ≈ 0.42,
p=0.024). This is regime/population non-stationarity, not a
"different setup families" story, and it is consistent with P3-S23's fold-3
collapse and period dependence.
**Verdict:** Population heterogeneity is **real but temporal/regime-based**,
not structural family-mixing. Redesigning the setup mechanism to "separate
families" is not supported by this diagnostic.
### Hypothesis C — Label / Outcome Information Limit
**Question:** is the current label contract too noisy?
Findings from P3-S16 v1 + P3-S18A (unchanged, read-only):
- Censoring is low: UNRESOLVED leads ≈ 3% (25/610 corrected); AMBIGUOUS
(same-bar TP+SL) ≈ 1% (6/610). All corrected UNRESOLVED are timeouts.
- First-hit resolution is deterministic; survival says SL resolves ~2x faster
than TP (median 2 vs 3 bars), but this is a frequency fact, not a leak.
- Horizon H=16 captures >98% of outcomes; it is a label window (lifetime-based),
distinct from runtime validity, and was approved as v1.
**Verdict:** Label **noise/censoring is NOT a first-order limiter.** The
contract was semantically validated and the ambiguous/uncensored share is small.
What IS mutable is the *base rate* of WIN/LOSS over time (see Hypothesis B),
which is a regime effect, not label noise. No label v2 is required to answer
the strategic question.
### Hypothesis D — Genuine Absence of Predictive Edge
**Question:** does the corrected evidence indicate the current setup definition
carries little predictive information?
Evidence on the corrected population (P3-S23, independently verified):
- pooled OOS ROC-AUC **0.5235**, 95% CI **[0.450, 0.592] ⊇ 0.5**;
- fold 3 collapses below random (**0.464**) while folds 1–2 are weak-positive
(0.539/0.538) ⇒ **period-instanced, not stable**;
- LogLoss/Brier are **worse than the constant prior** (0.647 vs 0.621;
0.226 vs 0.215) ⇒ not even a usable probability output;
- the frozen→corrected change moved pooled AUC **−0.056 → 0.5235**, i.e. the
earlier "signal" was substantially a construction artifact.
**Verdict:** On the validated corrected population, **a stable predictive edge
is NOT demonstrated** for the current setup definition. This is the honest,
evidence-led conclusion; it is NOT forced negative — it is what the corrected
numbers say.
---
## Phase 2 — Value Ranking of Future Research Options
| # | Option | Expected value | Rationale |
|---|---|---|---|
| 1 | Increase dataset size (period/instruments) | **MEDIUM** | Raises power (needed to see 0.55–0.58), but alone cannot recreate an edge the corrected pop does not show; multi-instrument changes the traded mechanism. |
| 2 | External data validation (Tickstory/Dukascopy/second feed) | **HIGH** | Directly answers broker-feed bias, M30 UTC-clock basis-parity, and population robustness — the largest residual unknown behind S22.3–S23. Lowest cost, highest information. |
| 3 | Setup mechanism redesign | **MEDIUM (conditional/deferred)** | Hypothesis B shows no structural family-mix; redesign is premature until a replicable basis signal exists. |
| 4 | Advanced ML | **LOW now** | Nonlinear models already overfit on this sample (P3-S18 tree/boost/MLP OOS collapse), calibration didn't help, corrected baseline adds no signal. Defer. |
---
## Phase 3 — Required Quantitative Analysis
### Sample Power Analysis (Hypothesis A numerals)
Current: n=579 binary-fit (175 WIN / 404 LOSS); pooled OOS n=279 (87/192).
Method: Hanley-McNeil AUC variance + normal-approx power vs A=0.5 (two-sided
α=0.05). Full table in `output/p3_s24_power.json`.
**Minimum sample recommendation (reliable detection, 80%/90% power):**
| Target AUC | balanced total | @ current LOSS:WIN ratio |
|---|---|---|
| 0.55 | ~1,036 / ~1,386 | ~1,261 / ~1,687 |
| 0.58 | ~398 / ~532 | ~493 / ~658 |
| 0.60 | ~250 / ~336 | ~314 / ~420 |
Practical recommendation for this framework (current 1:2.3 class ratio):
**≈1,300–1,700 independent WIN/LOSS leads to resolve a 0.55-level edge;**
**≈500–660 for 0.58; ≈320–420 for 0.60.** For a tight ±0.03 CI near noise:
**≈1,700 leads** (current ratio) or **≈1,420** (balanced). The current 579-row
population is roughly **half to one-third** of what a 0.55 detection needs.
### Evidence Matrix (Finding | Evidence | Confidence | Implication)
| Phase | Finding | Evidence | Confidence | Implication |
|---|---|---|---|---|
| P3-S18 | Baseline INCONCLUSIVE / too small; nonlinear overfits OOS | logistic t=0.609; tree 0.513, boost 0.580, mlp 0.595; 115 test | Datasets/features verified (S19); single split | Capacity above logistic adds nothing on this n |
| P3-S20 | Frozen-pop walk-forward pooled 0.5792, 3/3 folds>0.5 | pooled 0.5792/PR 0.342 vs prior 0.273 | **INDEPENDENTLY VERIFIED** (S21.R) | Weak ranking signal existed on the OLD (faulty) population |
| P3-S21.1 | ZONE position + short context are stable leading coefs; ~8 effective dims | std-coef zone 0.138, direction 0.082; heavy collinearity | Reproduced; descriptive | If real, signal rides on zone geometry + context, low dims |
| P3-S21.2 | Ranking exists; calibration does not improve | raw 0.5895/0.2003; isotonic worse; prior best | **VERIFIED** | No usable/actionable probability output |
| P3-S22.3 | M30 construction defect; 686→707, 202/223 drifted | oracle 100%; 6 mutations; 0-mismatch | **VERIFIED** | A hidden data-pipeline factor, not the model, drove much of the signal |
| P3-S22.4 | Corrected pop + contracts validated | 707=610+97; leakage 8/8; schema const | **VERIFIED** | Clean deterministic basis for S23 |
| P3-S23 | Corrected pooled 0.5235, fold3 0.464, cal worse | CI [0.450,0.592]; 8/8 mutations | **VERIFIED** | Edge does NOT reproduce as stable on the corrected population |
### Research Decision Matrix (Rank)
| Research direction | Expected info gain | Cost | Priority |
|---|---|---|---|
| **Option 2 — External data validation** | **HIGH** | LOW–MEDIUM | **1** |
| Option 1 — Dataset expansion | MEDIUM | LOW–MEDIUM | 2 |
| Option 3 — Setup mechanism redesign | MEDIUM (deferred) | HIGH | 3 |
| Option 4 — Advanced ML | LOW | MEDIUM–HIGH | 4 |
---
## Final Decision Classification
```text
D — CURRENT EDGE NOT DEMONSTRATED
```
Basis (all independently verified, corrected population):
- pooled OOS ROC-AUC 0.5235 with 95% CI **[0.450, 0.592]** containing 0.5;
- fold-3 collapse below random (0.464) ⇒ period-instanced;
- LogLoss/Brier worse than the constant prior ⇒ no usable probability;
- corrected construction (S22.3) removed the bulk of the earlier apparent edge.
Rejected options and reasons:
- **A (dataset expansion) — rejected as PRIMARY.** Sample size is a genuine
second-order limit (power to detect 0.55 is ~48%), but the corrected point
estimate (0.5235) is *below* the detectable threshold and the frozen→corrected
change (−0.056) shows the earlier edge was largely artifact — more samples on
the same definition/feed will not conjure an undemonstrated edge.
- **B (mechanism research) — rejected.** No structural family heterogeneity
exists (all frozen-dimension splits p>0.2); only temporal base-rate drift is
significant. Separating "set-up families" is not supported.
- **C (label/contract review) — rejected.** Label v1 was semantically validated;
censoring/ambiguity are small (≈3%/1%). Outcome base-rate drift is a regime
effect, not label noise.
- **E (data/pipeline defect) — resolved, not current.** The M30 construction
defect was corrected and verified in S22.3; however the **basis-parity of the
corrected build against the broker feed was never independently re-established**
— this becomes the top validation target (Option 2), not an assumption.
Highest-value next action (Option 2): **independently validate the broker-feed /
UTC-clock M30 basis-parity and the corrected population** with external data —
because it (a) resolves the single largest residual unknown, and (b) is the only
step that can separate "genuine absence of edge" (D) from a "residual
feed/construction artifact" (E). It is low-cost, HIGH-information, and NOT
auto-imported in this phase. Only after a validated basis should dataset
expansion to the S24 power targets proceed.
---
## Hard constraints (this phase honoured)
```text
ML training / tuning : NONE
New features : NONE (descriptive bins of existing features only)
Label / TP / SL / horizon : NONE changed
MQL5 / production : NONE changed
External data : NONE imported (MQL5/Files assessed only)
Calibration / nonlinear : NONE
Deployment / trading : NONE
P3-S25 : NOT STARTED (owner authorization required)
```
## Artifacts
```text
Namespace : ml/p3/p3_s24_research_value_assessment/
p3_s24_config.py / p3_s24_power.py / p3_s24_heterogeneity.py /
p3_s24_evidence.py / p3_s24_run_main.py / README.md
output : p3_s24_power.json, p3_s24_heterogeneity.json,
p3_s24_evidence_matrix.json, p3_s24_decision_matrix.json,
p3_s24_classification.json, p3_s24_manifest.json
Docs : this file, docs/P3_S24_NEXT_RESEARCH_PRIORITY.md,
docs/SESSION_HANDOVER_2026-08-26_P3_S24_RESEARCH_VALUE_ASSESSMENT.md
Reproducibility: two runs byte-identical except generated_utc (aggregate
SHA256 9912f7ab887bd0ec… identical across runs)
```
*End of P3-S24 research value assessment. Classification D — CURRENT EDGE NOT
DEMONSTRATED; highest-value next action = independent/external data validation
of the corrected basis, then dataset expansion to the computed power targets.*