forked from chiki2bum2/SniperGold_ML
283 lines
14 KiB
Markdown
283 lines
14 KiB
Markdown
# P3-S24 — RESEARCH VALUE ASSESSMENT & SIGNAL RECOVERY STRATEGY
| |||
| |||
```text
| |||
Date : 2026-08-26
| |||
Session : P3-S24 — RESEARCH VALUE ASSESSMENT (ANALYSIS ONLY)
| |||
Status : COMPLETE
| |||
Policy : docs/CODE_VERIFICATION_AND_SILENT_BUG_POLICY_v1.md
| |||
Classification : D — CURRENT EDGE NOT DEMONSTRATED
| |||
Next : P3-S25 = NOT STARTED (owner authorization required)
| |||
Predecessor: docs/SESSION_HANDOVER_2026-08-26_P3_S23_CORRECTED_POPULATION_ML_BASELINE.md
| |||
Namespace : ml/p3/p3_s24_research_value_assessment/
| |||
```
| |||
| |||
## Objective
| |||
| |||
Determine the most scientifically valuable next research direction AFTER the
| |||
observation that the corrected (UTC-clock M30) Candidate Setup population does
| |||
NOT reproduce the weak logistic ranking signal seen on the frozen population
| |||
(P3-S20 pooled ROC-AUC 0.5792 → P3-S23 corrected pooled ROC-AUC 0.5235).
| |||
| |||
This is NOT a model-optimization phase. The research principle is **increase
| |||
information quality, not model complexity**. Output = the highest-value next
| |||
research investment.
| |||
| |||
---
| |||
| |||
## 0. Preflight (mandatory — all PASS)
| |||
| |||
```text
| |||
Newest handover used (Git ancestry): docs/SESSION_HANDOVER_2026-08-26_P3_S23_...
| |||
Read completely (policy + required evidence):
| |||
CODE_VERIFICATION_AND_SILENT_BUG_POLICY_v1.md
| |||
P3_S23_CORRECTED_POPULATION_ML_BASELINE.md
| |||
P3_S22_4_CORRECTED_POPULATION_VALIDATION.md
| |||
P3_S22_3_RESEARCH_M30_REPAIR.md
| |||
P3_S21_R_RETROSPECTIVE_SILENT_BUG_VERIFICATION.md
| |||
P3_S21_2_CALIBRATION_REPORT.md
| |||
P3_S21_1_FEATURE_CONTRIBUTION_AUDIT.md
| |||
P3_S20_WALK_FORWARD_BASELINE.md
| |||
P3_S19_DATASET_FEATURE_AUDIT.md (+ P3_S18, P3_S18A, P3_S16 label contract,
| |||
P3_S23 pre-registration, S23 machine evidence JSON)
| |||
| |||
Git reconciled at start (before any work):
| |||
local HEAD == origin/main == 05ba7762fe4e17c1cbf9bf8bebdbbb1347209d24
| |||
branch main ; working tree CLEAN ; no stash ; no untracked files
| |||
```
| |||
| |||
---
| |||
| |||
## Phase 1 — Research Diagnosis Tree
| |||
| |||
The four candidate causes are evaluated on the CORRECTED 579 binary-fit
| |||
population (175 WIN / 404 LOSS) and the P3-S23 pooled OOS (279; 87 WIN).
| |||
| |||
### Hypothesis A — Sample Size Limitation
| |||
| |||
**Question:** is 579 binary-fit observations insufficient to detect a small
| |||
but real edge?
| |||
| |||
**Quantified (Hanley-McNeil variance + normal-approximation power, α=0.05
| |||
two-sided; see §Phase 3):**
| |||
| |||
| Detect AUC | total N, balanced (80%/90% power) | total N @ current ratio 1:2.31 (80%/90%) |
| |||
|---|---|---|
| |||
| 0.55 | 1,036 / 1,386 | 1,261 / 1,687 |
| |||
| 0.58 | 398 / 532 | 493 / 658 |
| |||
| 0.60 | 250 / 336 | 314 / 420 |
| |||
| |||
At the current sample, power to reject AUC=0.5 is:
| |||
| |||
| True AUC | power @ binary fit (175/404) | power @ pooled OOS (87/192) |
| |||
|---|---|---|
| |||
| 0.55 | 0.476 | 0.264 |
| |||
| 0.58 | 0.861 | 0.568 |
| |||
| 0.60 | 0.969 | 0.763 |
| |||
| |||
Smallest AUC detectable at 80% power: **≈0.574 at the full fit**, **≈0.605 at
| |||
the pooled-OOS evaluation**. The observed corrected point estimate (0.5235) is
| |||
**below even the detectable boundary at the fit**, and its 95% CI
| |||
[0.450, 0.592] is ~±0.074 wide and includes 0.5.
| |||
| |||
**Verdict:** Sample size is a **real second-order limiter** (a 0.55-level,
| |||
real edge is far from resolvable; power targets ≈ 1,050–1,700 observations).
| |||
But sample size alone does NOT explain why the corrected point estimate (0.5235)
| |||
is at/under noise *and* why the frozen→corrected change moved it −0.056. The
| |||
observed residual is indistinguishable from noise.
| |||
| |||
### Hypothesis B — Population Heterogeneity
| |||
| |||
**Question:** is the setup population mixing different mechanisms?
| |||
| |||
**Descriptive diagnostic (existing frozen dimensions only, corrected 579 rows;
| |||
no new features, no model):**
| |||
| |||
| Dimension (bin) | WIN-rate spread | chi-square p |
| |||
|---|---|---|
| |||
| direction (short/long) | 0.314 vs 0.290 → 0.024 | 0.522 |
| |||
| zone position (tertiles) | 0.254–0.342 | 0.249 |
| |||
| CHoCH age (fresh→old) | 0.333→0.206 | 0.265 |
| |||
| CHoCH latency (tertiles) | 0.291–0.317 | 0.941 |
| |||
| liquidity / sweep age (tertiles) | 0.281–0.325 | 0.838 |
| |||
| volatility / ATR (tertiles) | 0.254–0.344 | 0.240 |
| |||
| **calendar year** | **0.151–0.419** | **0.024*** |
| |||
| |||
- **No setup family** (direction, zone position, CHoCH timing, liquidity
| |||
behaviour, volatility scale) shows a statistically resolvable WIN-rate
| |||
difference. The population is **NOT** mixing meaningfully different,
| |||
separable mechanisms along the frozen structural dimensions.
| |||
- The **only** significant heterogeneity is **temporal**: the outcome WIN
| |||
base rate itself drifts strongly by year (2018 ≈ 0.15 → 2026 ≈ 0.42,
| |||
p=0.024). This is regime/population non-stationarity, not a
| |||
"different setup families" story, and it is consistent with P3-S23's fold-3
| |||
collapse and period dependence.
| |||
| |||
**Verdict:** Population heterogeneity is **real but temporal/regime-based**,
| |||
not structural family-mixing. Redesigning the setup mechanism to "separate
| |||
families" is not supported by this diagnostic.
| |||
| |||
### Hypothesis C — Label / Outcome Information Limit
| |||
| |||
**Question:** is the current label contract too noisy?
| |||
| |||
Findings from P3-S16 v1 + P3-S18A (unchanged, read-only):
| |||
- Censoring is low: UNRESOLVED leads ≈ 3% (25/610 corrected); AMBIGUOUS
| |||
(same-bar TP+SL) ≈ 1% (6/610). All corrected UNRESOLVED are timeouts.
| |||
- First-hit resolution is deterministic; survival says SL resolves ~2x faster
| |||
than TP (median 2 vs 3 bars), but this is a frequency fact, not a leak.
| |||
- Horizon H=16 captures >98% of outcomes; it is a label window (lifetime-based),
| |||
distinct from runtime validity, and was approved as v1.
| |||
| |||
**Verdict:** Label **noise/censoring is NOT a first-order limiter.** The
| |||
contract was semantically validated and the ambiguous/uncensored share is small.
| |||
What IS mutable is the *base rate* of WIN/LOSS over time (see Hypothesis B),
| |||
which is a regime effect, not label noise. No label v2 is required to answer
| |||
the strategic question.
| |||
| |||
### Hypothesis D — Genuine Absence of Predictive Edge
| |||
| |||
**Question:** does the corrected evidence indicate the current setup definition
| |||
carries little predictive information?
| |||
| |||
Evidence on the corrected population (P3-S23, independently verified):
| |||
- pooled OOS ROC-AUC **0.5235**, 95% CI **[0.450, 0.592] ⊇ 0.5**;
| |||
- fold 3 collapses below random (**0.464**) while folds 1–2 are weak-positive
| |||
(0.539/0.538) ⇒ **period-instanced, not stable**;
| |||
- LogLoss/Brier are **worse than the constant prior** (0.647 vs 0.621;
| |||
0.226 vs 0.215) ⇒ not even a usable probability output;
| |||
- the frozen→corrected change moved pooled AUC **−0.056 → 0.5235**, i.e. the
| |||
earlier "signal" was substantially a construction artifact.
| |||
| |||
**Verdict:** On the validated corrected population, **a stable predictive edge
| |||
is NOT demonstrated** for the current setup definition. This is the honest,
| |||
evidence-led conclusion; it is NOT forced negative — it is what the corrected
| |||
numbers say.
| |||
| |||
---
| |||
| |||
## Phase 2 — Value Ranking of Future Research Options
| |||
| |||
| # | Option | Expected value | Rationale |
| |||
|---|---|---|---|
| |||
| 1 | Increase dataset size (period/instruments) | **MEDIUM** | Raises power (needed to see 0.55–0.58), but alone cannot recreate an edge the corrected pop does not show; multi-instrument changes the traded mechanism. |
| |||
| 2 | External data validation (Tickstory/Dukascopy/second feed) | **HIGH** | Directly answers broker-feed bias, M30 UTC-clock basis-parity, and population robustness — the largest residual unknown behind S22.3–S23. Lowest cost, highest information. |
| |||
| 3 | Setup mechanism redesign | **MEDIUM (conditional/deferred)** | Hypothesis B shows no structural family-mix; redesign is premature until a replicable basis signal exists. |
| |||
| 4 | Advanced ML | **LOW now** | Nonlinear models already overfit on this sample (P3-S18 tree/boost/MLP OOS collapse), calibration didn't help, corrected baseline adds no signal. Defer. |
| |||
| |||
---
| |||
| |||
## Phase 3 — Required Quantitative Analysis
| |||
| |||
### Sample Power Analysis (Hypothesis A numerals)
| |||
| |||
Current: n=579 binary-fit (175 WIN / 404 LOSS); pooled OOS n=279 (87/192).
| |||
Method: Hanley-McNeil AUC variance + normal-approx power vs A=0.5 (two-sided
| |||
α=0.05). Full table in `output/p3_s24_power.json`.
| |||
| |||
**Minimum sample recommendation (reliable detection, 80%/90% power):**
| |||
| |||
| Target AUC | balanced total | @ current LOSS:WIN ratio |
| |||
|---|---|---|
| |||
| 0.55 | ~1,036 / ~1,386 | ~1,261 / ~1,687 |
| |||
| 0.58 | ~398 / ~532 | ~493 / ~658 |
| |||
| 0.60 | ~250 / ~336 | ~314 / ~420 |
| |||
| |||
Practical recommendation for this framework (current 1:2.3 class ratio):
| |||
**≈1,300–1,700 independent WIN/LOSS leads to resolve a 0.55-level edge;**
| |||
**≈500–660 for 0.58; ≈320–420 for 0.60.** For a tight ±0.03 CI near noise:
| |||
**≈1,700 leads** (current ratio) or **≈1,420** (balanced). The current 579-row
| |||
population is roughly **half to one-third** of what a 0.55 detection needs.
| |||
| |||
### Evidence Matrix (Finding | Evidence | Confidence | Implication)
| |||
| |||
| Phase | Finding | Evidence | Confidence | Implication |
| |||
|---|---|---|---|---|
| |||
| P3-S18 | Baseline INCONCLUSIVE / too small; nonlinear overfits OOS | logistic t=0.609; tree 0.513, boost 0.580, mlp 0.595; 115 test | Datasets/features verified (S19); single split | Capacity above logistic adds nothing on this n |
| |||
| P3-S20 | Frozen-pop walk-forward pooled 0.5792, 3/3 folds>0.5 | pooled 0.5792/PR 0.342 vs prior 0.273 | **INDEPENDENTLY VERIFIED** (S21.R) | Weak ranking signal existed on the OLD (faulty) population |
| |||
| P3-S21.1 | ZONE position + short context are stable leading coefs; ~8 effective dims | std-coef zone 0.138, direction 0.082; heavy collinearity | Reproduced; descriptive | If real, signal rides on zone geometry + context, low dims |
| |||
| P3-S21.2 | Ranking exists; calibration does not improve | raw 0.5895/0.2003; isotonic worse; prior best | **VERIFIED** | No usable/actionable probability output |
| |||
| P3-S22.3 | M30 construction defect; 686→707, 202/223 drifted | oracle 100%; 6 mutations; 0-mismatch | **VERIFIED** | A hidden data-pipeline factor, not the model, drove much of the signal |
| |||
| P3-S22.4 | Corrected pop + contracts validated | 707=610+97; leakage 8/8; schema const | **VERIFIED** | Clean deterministic basis for S23 |
| |||
| P3-S23 | Corrected pooled 0.5235, fold3 0.464, cal worse | CI [0.450,0.592]; 8/8 mutations | **VERIFIED** | Edge does NOT reproduce as stable on the corrected population |
| |||
| |||
### Research Decision Matrix (Rank)
| |||
| |||
| Research direction | Expected info gain | Cost | Priority |
| |||
|---|---|---|---|
| |||
| **Option 2 — External data validation** | **HIGH** | LOW–MEDIUM | **1** |
| |||
| Option 1 — Dataset expansion | MEDIUM | LOW–MEDIUM | 2 |
| |||
| Option 3 — Setup mechanism redesign | MEDIUM (deferred) | HIGH | 3 |
| |||
| Option 4 — Advanced ML | LOW | MEDIUM–HIGH | 4 |
| |||
| |||
---
| |||
| |||
## Final Decision Classification
| |||
| |||
```text
| |||
D — CURRENT EDGE NOT DEMONSTRATED
| |||
```
| |||
| |||
Basis (all independently verified, corrected population):
| |||
- pooled OOS ROC-AUC 0.5235 with 95% CI **[0.450, 0.592]** containing 0.5;
| |||
- fold-3 collapse below random (0.464) ⇒ period-instanced;
| |||
- LogLoss/Brier worse than the constant prior ⇒ no usable probability;
| |||
- corrected construction (S22.3) removed the bulk of the earlier apparent edge.
| |||
| |||
Rejected options and reasons:
| |||
- **A (dataset expansion) — rejected as PRIMARY.** Sample size is a genuine
| |||
second-order limit (power to detect 0.55 is ~48%), but the corrected point
| |||
estimate (0.5235) is *below* the detectable threshold and the frozen→corrected
| |||
change (−0.056) shows the earlier edge was largely artifact — more samples on
| |||
the same definition/feed will not conjure an undemonstrated edge.
| |||
- **B (mechanism research) — rejected.** No structural family heterogeneity
| |||
exists (all frozen-dimension splits p>0.2); only temporal base-rate drift is
| |||
significant. Separating "set-up families" is not supported.
| |||
- **C (label/contract review) — rejected.** Label v1 was semantically validated;
| |||
censoring/ambiguity are small (≈3%/1%). Outcome base-rate drift is a regime
| |||
effect, not label noise.
| |||
- **E (data/pipeline defect) — resolved, not current.** The M30 construction
| |||
defect was corrected and verified in S22.3; however the **basis-parity of the
| |||
corrected build against the broker feed was never independently re-established**
| |||
— this becomes the top validation target (Option 2), not an assumption.
| |||
| |||
Highest-value next action (Option 2): **independently validate the broker-feed /
| |||
UTC-clock M30 basis-parity and the corrected population** with external data —
| |||
because it (a) resolves the single largest residual unknown, and (b) is the only
| |||
step that can separate "genuine absence of edge" (D) from a "residual
| |||
feed/construction artifact" (E). It is low-cost, HIGH-information, and NOT
| |||
auto-imported in this phase. Only after a validated basis should dataset
| |||
expansion to the S24 power targets proceed.
| |||
| |||
---
| |||
| |||
## Hard constraints (this phase honoured)
| |||
| |||
```text
| |||
ML training / tuning : NONE
| |||
New features : NONE (descriptive bins of existing features only)
| |||
Label / TP / SL / horizon : NONE changed
| |||
MQL5 / production : NONE changed
| |||
External data : NONE imported (MQL5/Files assessed only)
| |||
Calibration / nonlinear : NONE
| |||
Deployment / trading : NONE
| |||
P3-S25 : NOT STARTED (owner authorization required)
| |||
```
| |||
| |||
## Artifacts
| |||
| |||
```text
| |||
Namespace : ml/p3/p3_s24_research_value_assessment/
| |||
p3_s24_config.py / p3_s24_power.py / p3_s24_heterogeneity.py /
| |||
p3_s24_evidence.py / p3_s24_run_main.py / README.md
| |||
output : p3_s24_power.json, p3_s24_heterogeneity.json,
| |||
p3_s24_evidence_matrix.json, p3_s24_decision_matrix.json,
| |||
p3_s24_classification.json, p3_s24_manifest.json
| |||
Docs : this file, docs/P3_S24_NEXT_RESEARCH_PRIORITY.md,
| |||
docs/SESSION_HANDOVER_2026-08-26_P3_S24_RESEARCH_VALUE_ASSESSMENT.md
| |||
Reproducibility: two runs byte-identical except generated_utc (aggregate
| |||
SHA256 9912f7ab887bd0ec… identical across runs)
| |||
```
| |||
| |||
*End of P3-S24 research value assessment. Classification D — CURRENT EDGE NOT
| |||
DEMONSTRATED; highest-value next action = independent/external data validation
| |||
of the corrected basis, then dataset expansion to the computed power targets.*
|