SniperGold_ML/docs/P3_FEATURE_LABEL_FORENSIC.md

473 lines
23 KiB
Markdown

# P3 FEATURE/LABEL FORENSIC REPORT — SNIPERGOLD_ML
```text
Date : 2026-08-22
Session : P3 (phase P3.1 — Feature Information Audit + Label Forensic Audit)
Status : RESEARCH / DIAGNOSTIC ONLY — no production changes
Source : P2.6 corrected baseline (immutable checkpoint)
Artifacts : ml/p3/*.py + ml/p3/output/*.json (provenance recorded)
```
---
## 0. P2 CHECKPOINT — VERIFICATION (this session)
```text
Forge HEAD : 8d330343af688e2de2b2a1d12ce19a582714709a (branch main; ls-remote origin == local) OK
Working tree : 1 untracked file (docs/P2_6_CORRECTED_BASELINE_REPORT.md) OK (P2.6 artifact not yet committed)
P2_SOURCE_SHA : 8d330343af688e2de2b2a1d12ce19a582714709a
P2.6_SOURCE_SHA : 8d330343... (commit) ; train_p26.py=e510ef3b... ; build_features_p2.py=4680f905... (per contract §5)
FEATURE_CONTRACT_SHA : C44CC6F2B740C32D06F776BD7C3E669DC5A8A6DE0484230544EBFFCF517D38DD (canonical P2.5 hash)
current file = 7b908b12... -> the file was appended with the §5 hash table AFTER the canonical hash was computed.
NOT hot-fixed (P3 rules); recorded as a provenance discrepancy.
MODEL_SHA : 06df8452a112290ecb9bae1a3dbff6df492e084f8b872885d808bb0303ec6a70
(Include\SniperGold_ML_p26_corrected.mqh, tag v20260821_p26)
DATASET_SHA : e85a08616191fb7ce4bbe79c1754228ddbcf54562034a322747501c7aabe124e
(recomputed this session from F+label+A+t 2017+; prefix e85a0861 matches P2.6) OK
Dataset : Files\AlgoForge\Data\XAUUSD_*.npz ; M15 bars 2017+ n=197032
Symbol/TF : XAUUSD / M15 ; window 2017-01-01 .. 2026-08-20
Label (baseline) : Y = sign(c[i+24]-c[i]) vs ±0.75×ATR(i) ; purge gap 24 ; seed 42 ; MLP 19->12->2
```
Frozen-model reproduction (sanity check): the P2.6 model is parsed from the .mqh and evaluated
on the same purged TEST split → **AUC LONG 0.5105 / SHORT 0.5063 — identical to the P2.6
offline result**. Model parsing + split + label are guaranteed consistent.
---
## A. FEATURE INFORMATION (P3.1a — feature_audit.py)
### A.1 Descriptive summary + per-feature information (vs 24-bar / 0.75 ATR label)
| id | feature | mean | std | uniq | zero% | AUC | r_pb | d | MI | class |
|----|---------|------|-----|------|-------|-----|------|-----|----|-------|
| 0 | htf1_bias (D1) | 0.166 | 0.986 | 2 | 0.0% | 0.5108 | +0.006 | +0.011 | 0.0000 | WEAK |
| 1 | htf2_bias (H4) | 0.080 | 0.997 | 2 | 0.0% | 0.5081 | +0.011 | +0.021 | 0.0001 | WEAK |
| 2 | htf3_bias (H1) | 0.093 | 0.996 | 2 | 0.0% | 0.4956 | -0.009 | -0.019 | 0.0000 | WEAK |
| 3 | swing_trend | 0.079 | 0.996 | 3 | 0.2% | 0.5029 | -0.002 | -0.005 | 0.0000 | NEAR-DEGENERATE |
| 4 | internal_trend | 0.056 | 0.998 | 3 | 0.1% | 0.4880 | -0.025 | -0.051 | 0.0003 | INFORMATIVE* |
| 5 | chart_bias | 0.057 | 0.998 | 3 | 0.1% | 0.4879 | -0.026 | -0.052 | 0.0003 | INFORMATIVE* |
| 6 | eq_pos_norm | 0.281 | 2.842 | 188051 | 0.7% | 0.4975 | -0.015 | -0.030 | 0.0005 | WEAK |
| 7 | sweep_dir | -0.049 | 0.999 | 3 | 0.1% | 0.5075 | +0.015 | +0.029 | 0.0001 | WEAK |
| 8 | choch_dir | 0.074 | 0.997 | 3 | 0.1% | 0.4978 | -0.005 | -0.009 | 0.0000 | NEAR-DEGENERATE |
| 9 | choch_confirms | 0.266 | 0.442 | 2 | 73.4% | 0.4991 | +0.002 | +0.005 | 0.0000 | NEAR-DEGENERATE |
| 10 | eqh_swept | 0.0045 | 0.067 | 2 | 99.6% | 0.5084 | -0.009 | -0.017 | 0.0000 | NEAR-DEGENERATE |
| 11 | eql_swept | 0.0082 | 0.090 | 2 | 99.2% | 0.5077 | -0.002 | -0.004 | 0.0000 | NEAR-DEGENERATE |
| 12 | delta_sign | 0.063 | 0.998 | 3 | 0.0% | 0.4936 | -0.016 | -0.032 | 0.0001 | WEAK |
| 13 | delta_mag | 0.015 | 0.191 | 197031 | 0.0% | 0.4888 | -0.020 | -0.040 | 0.0003 | WEAK |
| 14 | dist_high_atr | 4.050 | 6.109 | 130676 | 0.1% | 0.5035 | +0.001 | +0.002 | 0.0004 | UNKNOWN** |
| 15 | dist_low_atr | 5.505 | 5.347 | 117777 | 0.1% | 0.4993 | -0.008 | -0.017 | 0.0002 | WEAK |
| 16 | mom20_atr | 0.189 | 2.844 | 196516 | 0.1% | 0.4862 | -0.023 | -0.047 | 0.0003 | INFORMATIVE* |
| 17 | range_atr | 14.614 | 12.066 | 194430 | 0.7% | 0.5004 | +0.001 | +0.002 | 0.0004 | UNKNOWN** |
| 18 | confluence | 50.945 | 16.710 | 10 | 0.0% | 0.5080 | +0.014 | +0.028 | 0.0003 | WEAK |
`*` INFORMATIVE only statistically marginal (|AUC-0.5| ≤ 0.014, |d| ≤ 0.052) — **not
practically relevant**. `**` UNKNOWN = continuous feature with no detected effect (interpretation:
NO-EVIDENCE).
### A.2 Conclusion A
```text
NO feature has material information content against the current label.
Max |AUC - 0.5| = 0.0138 (f16 mom20). Max |r_pb| = 0.0258 (f5 chart_bias).
All effects are "negligible-to-small" (Cohen's d ≤ 0.05).
No CONSTANT feature, but 6 NEAR-DEGENERATE features (including f10/f11).
```
Research Question #1 answer (partial): **the 19 runtime-consistent features carry NO
meaningful univariate information against the current label.** (Caveat: a univariate
audit does not capture interactions — addressed in sections G/H.)
---
## B. F10/F11 EQH/EQL — SPECIAL AUDIT
```text
f10 eqh_swept : n_on=877 (0.45% of bars), P(+1|on)=0.4595 vs off=0.5256 -> lean BEARISH (-0.066)
P(-1|on)=0.4185 (label audit)
co-occurrence: sweep 100%, choch 100%, confirm 23.5%
year distribution: ~0 in 2017/2019/2020/2021, rising 2025-2026 (0.19%, 0.12%)
f11 eql_swept : n_on=1617 (0.82% of bars), P(+1|on)=0.5147 vs off=0.5254 (-0.011)
co-occurrence: sweep 100%, choch 100%, confirm 10.0%
year distribution: cluster 2019-2022, ~0 in 2025-2026
Both : concentration at 17-19 UTC (US session)
```
Conclusion B:
```text
Under legacy v4.4 semantics, EQH/EQL is a VERY RARE and NON-STATIONARY state
(frequency changes drastically across years). Not a recurring "setup trigger";
almost inert as a feature. The EQH direction association is bearish (logical:
equal-high swept = sell-side liquidity taken), but with 877-1617 events over
9.6 years and a small effect, it cannot be relied upon.
-> P2.6 was right: corrected f10/f11 ~constant 0; this is not a bug but the
character of the legacy semantics itself. Semantics NOT changed (P3 rules).
```
---
## C. HTF BIAS (f0-f2) + CONFLUENCE (f18) — SPECIAL AUDIT
```text
f0 htf1_bias (D1): always ±1 (0% neutral), balance +58.3%/-41.7%,
inter-bar persistence = 0.9996, mean run = 2526 bars (~26 M15 days!)
P(+1|state+1)=0.5278 vs P(+1|state-1)=0.5220
f1 htf2_bias (H4): balance +54.0%, persistence 0.9976, mean run 410 bars
P(+1|+) = 0.5302 vs P(+1|-) = 0.5196
f2 htf3_bias (H1): balance +54.6%, persistence 0.9908, mean run 109 bars
P(+1|+) = 0.5211 vs P(+1|-) = 0.5305 (H1 weakly contrarian)
f18 confluence : min 25, max 90, 10 unique values; P(+1|>=50)=0.5306 vs uncond 0.5253
```
Conclusion C:
```text
HTF bias is a VERY SLOWLY MOVING state (D1 persists on average 2526 bars).
At the 24-bar label horizon, the static f0-f2 values are almost constant -> practically
cannot carry 24-bar information except through TRANSITIONS (rare: ~0.04%-0.9%
of bars). The state-age analysis (section I) shows no run-duration effect.
-> H4/Higher-TF bias provides no measurable information content at the 24-bar
horizon. Consistent with P2.1 (the old 50-bar-lag bug) which used to
"inject" false information through the wrong window.
```
---
## D. STRUCTURE FEATURES (f3-f9, f12-f17) — SPECIAL AUDIT
```text
f3-f8, f12-f17 : active >99.9% of bars (NOT rare events; these are persistent states)
-> the hypothesis "event-driven SMC features with an informative
subset" is NOT supported: active-gap = 1 bar, cluster5 = 1.0 (always adjacent)
f9 choch_confirms : the only truly event-like feature: active 26.6%,
mean gap 3.8 bars, cluster5 0.984; yet P(+1|on)=0.5272 vs off=0.5247
(diff +0.003) -> NO label association.
The large P1(active) vs P1(inactive) for f12/f13/f5 (diff -0.47/-0.47/-0.16) is an
ARTIFACT of the very small inactive samples (2-360 bars) -> not a signal.
```
Conclusion D:
```text
Structure features are almost always active (state, not event). The only
event-like feature (f9) is not associated with the label. The pattern
"feature active -> label outcome" is not found at the univariate level.
```
---
## E. LABEL FORENSICS (P3.1b — label_audit.py)
### E.1 Current label characteristics (24-bar / 0.75 ATR)
```text
P(+1)=0.4144 P(-1)=0.3744 P(0)=0.2113 | labeled=78.9% (155405 bars)
entropy labeled 2-class = 0.692 ; 3-class = 1.061 (fairly balanced for 2-class)
```
### E.2 Decomposition (H=24, thr=0.75 ATR, ALL bars; no look-ahead)
```text
fwd/ATR : mean 0.181, median 0.143 (small final movement)
MFE/ATR : median 3.82 (typical favorable excursion 3.8 ATR!)
MAE/ATR : median 1.24 (typical adverse excursion 1.2 ATR)
TP-before-SL : 0.482 | SL-before-TP: 0.284 | no-hit: 0.234
median time-to-hit : 2.0 BARS (0.75 ATR threshold reached within ~2 bars)
per direction (label != 0):
[long] MFE med 4.29, MAE med 0.98, TP 0.610, SL 0.361, nohit 0.029
[short] MFE med 4.46, MAE med 0.99, TP 0.613, SL 0.358, nohit 0.029
[flat] MFE med 2.22, MAE med 2.22, TP 0, SL 0, nohit 1.000
```
### E.3 Interpretation
```text
1. The 0.75 ATR threshold is very small relative to path noise (typical MFE 3.8 ATR):
the label is not "did the setup produce a move", but "does the close 24 bars
ahead drift > 0.75 ATR" — dominated by running drift/trend.
2. Even flat (label=0) breaks 0.75 ATR inside the path (MFE 2.2) and closes back
-> the point-to-point label does not capture path quality.
3. TP-before-SL 0.48 vs SL-before-TP 0.28 (all bars): win/loss path odds are
almost symmetric; no exploitable asymmetry under this label.
4. Label x SMC state (H=24): almost ALL lifts ≈ 1.0:
f5 bull 0.981 | f5 bear 1.021 | f7 bull grab 1.017 | f8 bull 1.001
f9 confirm 0.997 | conf>=70 1.031 | conf>=80 1.026
eqh_on 0.859 (bearish) | eql_on 0.912
-> the label is NOT materially conditioned on SMC state.
```
Research Question #2 answer (partial): **the 24-bar/0.75-ATR label is NOT a good
operationalization of "SMC setup outcome"** — the threshold is reached within ~2 bars,
the horizon is dominated by overlap/trend, and there is no SMC conditioning.
---
## F. HORIZON SENSITIVITY (descriptive; not optimization)
| H | P(+1) | P(-1) | P(0) | ent3 | MFE med | MAE med | TP | SL | nohit | t_hit med |
|----|-------|-------|------|------|---------|---------|-----|-----|-------|-----------|
| 12 | 0.363 | 0.331 | 0.306 | 1.096 | 2.52 | 0.95 | 0.472 | 0.201 | 0.328 | 2.0 |
| 24 | 0.414 | 0.374 | 0.211 | 1.061 | 3.82 | 1.24 | 0.482 | 0.284 | 0.234 | 2.0 |
| 36 | 0.440 | 0.390 | 0.170 | 1.030 | 4.78 | 1.50 | 0.484 | 0.322 | 0.194 | 2.0 |
| 48 | 0.456 | 0.397 | 0.147 | 1.007 | 5.64 | 1.73 | 0.485 | 0.344 | 0.171 | 2.0 |
```text
- outcome rate rises with horizon (69% -> 85%); entropy drops slightly.
- time-to-hit STAYS 2 bars at ALL horizons -> the threshold is always reached fast;
longer horizons only add MAE noise & SL-before-TP.
- No horizon shows a qualitatively different "SMC persistence" profile. SMC signal
(if any) does not appear to "live" at a specific horizon.
```
---
## G. ALL-BAR vs CANDIDATE-SETUP (P3.1c — candidate_setup_audit.py)
Model: **frozen P2.6 MLP** (no retraining), purged TEST split, n=38828.
| stratum | n_test | P(+1) | AUC_L | AUC_S | dAUC_L vs D0 |
|---------|--------|-------|-------|-------|--------------|
| D0 all bars | 38828 | 0.5504 | **0.5105** | **0.5063** | — (P2.6 reproduction) |
| D1 any_event (f7\|f8\|f9\|f10\|f11) | 38828 | 0.5504 | 0.5105 | 0.5063 | 0.000 (rate 99.95%!) |
| BUY_candidate (f9&f7>0) | 5203 | 0.5933 | 0.5163 | 0.5369 | +0.006 |
| SELL_candidate (f9&f7<0) | 4312 | 0.5401 | 0.5133 | 0.5204 | +0.003 |
| BUY_strict (+conf>=60) | 2664 | 0.6141 | 0.4894 | 0.4857 | -0.021 |
| SELL_strict (+conf>=60) | 1355 | 0.5446 | 0.4902 | 0.4978 | -0.020 |
| MTF_aligned | 14172 | 0.5674 | 0.4862 | 0.4793 | -0.024 |
| MTF_conflicting | 24656 | 0.5406 | 0.5196 | 0.5187 | +0.009 |
| HTF_bull | 10558 | 0.5820 | 0.4766 | 0.4719 | -0.034 |
| HTF_bear | 3614 | 0.5246 | 0.5118 | 0.5069 | +0.001 |
| Fuzzy_high (conf>=60) | 10936 | 0.5749 | 0.4815 | 0.4826 | -0.029 |
| Fuzzy_low (conf<40) | 6763 | 0.5215 | 0.5003 | 0.5384 | -0.010 |
| EQH_event (f10) | 526 | 0.4620 | **0.5508** | 0.4317 | +0.040 |
| EQL_event (f11) | 3 | — | — | — | (too few) |
```text
Key findings:
1. The "ANY EVENT" filter is meaningless: 99.95% of bars have an SMC event
(features are almost always active) -> D1_any_event == D0.
2. Conditioning generally does NOT raise discrimination:
MTF_aligned -0.024, HTF_bull -0.034, Fuzzy_high -0.029, BUY/SELL_strict -0.02.
3. A real BASE-RATE SHIFT (P(+1) 0.55 -> 0.61 in BUY_strict) but the frozen model
CANNOT discriminate within that subset (AUC drops) -> base-rate shift is not an edge.
4. The only material AUC improvement: EQH_event (+0.040 LONG) with n=526 and
bearish class balance (P(+1)=0.462); BUY_candidate SHORT +0.031 (n=5203).
Both small, partly opposite in direction, and prone to multiple testing.
```
Research Question #3 answer: **candidate-setup conditioning does NOT consistently
increase information density.** Meta-labeling is not supported by this evidence
under the current semantics.
---
## H. CONDITIONAL SMC ANALYSIS (strata)
```text
P(Y=1 | stratum) relative to the unconditional 0.5253 (labeled):
conf>=70 : +1.031 | conf>=80 : +1.026 (best, still weak)
BUY cand : +1.017 | f7 bull grab : +1.017
MTF bull : +1.003 | f9 confirm : +0.997
EQH on : 0.859 (BEARISH - significant but small n)
SELL cand & ATR high : 0.974 ; conf>=60 & ATR low : 0.983
No combined state (SMC BUY + H4 bull + M30 bull + M15 confirm) has an outcome
distribution MATERIALLY different from baseline.
```
---
## I. TEMPORAL PERSISTENCE (P3.1d — temporal_diagnostic.py)
```text
Continuous-feature autocorrelation (lag1): eq_pos 0.988, dist_high 0.991, dist_low 0.989,
range 0.977, conf 0.964, mom20 0.930, delta_mag 0.867 (decays fastest;
lag10 -> -0.01).
Discrete-state persistence: 0.83 (delta_sign) to 1.00 (eqh 0.9997, htf1 0.9996).
Run length: htf1 2526 bars, htf2 410, swing 315, choch 391, htf3 109, internal 40.
LAGGED INFORMATION AUC(feature[t-k] vs label[t]): almost FLAT for all k
(htf1 0.5107->0.5095 ; chart_bias 0.4890->0.4996 ; mom20 0.4872->0.4977)
-> because features are so persistent, lagged values == current values; sequence
representation adds no new information.
LABEL persistence: P(Y_t=1 | Y_{t-1}=1) = 0.987 (lift 1.88) BUT
P(Y_t=1 | Y_{t-24}=1) = 0.512 ~ unconditional 0.525
-> label persistence is a WINDOW-OVERLAP ARTIFACT (labels t and t+1 share
23/24 outcome bars), NOT a real temporal structure.
STATE-AGE: P(Y=1) flat vs the running state age (htf1: 0.46/0.45/0.54/0.54/0.53
for age 1-3/4-10/11-30/31-100/101+; chart_bias: 0.52-0.53 in all buckets)
-> state duration is NOT informative.
```
Research Question #4 answer: **NO evidence of exploitable temporal dependency.**
The observed persistence = overlap artifact; lagged info flat; state-age flat.
**LSTM has no empirical basis on this evidence.**
---
## J. REGIME DIAGNOSTIC (P3.1e — regime_diagnostic.py)
Proxy: ATR percentile (500-bar), RealizedVol20, |f16|, f17. Terciles low/med/high.
```text
P(+1|regime): all lifts 0.99-1.01 (ATR pct 0.990/1.005/1.006; RV 0.996/0.994/1.011;
trend 0.999/1.003/0.997; range 1.000/0.995/1.005) -> NO EFFECT.
AUC of key features per regime: all 0.48-0.52, no material change across
terciles (mom20 low 0.4906/med 0.4796/high 0.4881; conf 0.5052/0.5170/0.4983)
Regime x SMC combo: lift 0.974-1.025 (BUY&ATR_low 1.025; SELL&ATR_high 0.974)
-> no standout combination.
Note: htf1_bias is often NaN per regime because it has ONLY 1 unique value in the
subset (D1 bias constant for long periods) -> further evidence of extreme persistence.
```
Research Question #5 answer: **NO evidence of regime dependency** under these simple
proxies. The feature→label relationship is uniformly weak in all market states.
(Regime models — MS-GARCH/HMM/gate — are not supported.)
---
## K. FEATURE-GROUP ABLATION (diagnostic proxy; not feature selection)
| group | features | max\|AUC-0.5\| | mean\|AUC-0.5\| | max\|r_pb\| | n INFORMATIVE |
|------|-------|---------------|----------------|------------|---------------|
| A_HTF | f0,f1,f2 | 0.0108 | 0.0078 | 0.0105 | 0 |
| B_Context | f6,f14,f15,f17 | 0.0035 | 0.0018 | 0.0149 | 0 |
| C_Entry | f7,f8,f9,f10,f11 | 0.0084 | 0.0053 | 0.0145 | 0 |
| D_PriceAction | f12,f13,f16 | 0.0138 | 0.0105 | 0.0234 | 1 |
| E_Structural | f3,f4,f5 | 0.0121 | 0.0090 | 0.0258 | 2 |
| F_Confluence | f18 | 0.0080 | 0.0080 | 0.0142 | 0 |
```text
The families with (a little) univariate information are only Structural & PriceAction —
and even those are below the practical-relevance threshold. HTF/Context/Entry/Confluence ≈ 0.
```
---
## L. MULTIPLE-TESTING CAVEATS
```text
- Tests this session: 19 features x ~6 metrics + 13 strata + 8 horizons/lags +
8 regime combos + 4 proxies x 3 terciles x 6 features ≈ 250+ statistics.
- n=155405 labeled -> even tiny effects are "statistically significant";
this session uses EFFECT SIZE (AUC dev, Cohen's d, rank-biserial) as the reference.
- Only a few results exceed |AUC-0.5| >= 0.02: EQH (+0.040, n=526 test),
BUY_candidate SHORT (+0.031, n=5203), MTF_aligned (-0.024, n=14172),
HTF_bull (-0.034, n=10558). Mixed directions; some subsets overlap (BUY_strict
is a subset of BUY_candidate); real false-positive risk.
- The 24-bar label overlap makes the effective sample size far smaller than n
(0.94+ persistence between labeled bars); all CIs narrow but dependent.
- CONCLUSIONS are not built from a single p-value / single split / single stratum.
```
---
## M. ROOT-CAUSE CLASSIFICATION
```text
A — FEATURE FAILURE : STRONGLY SUPPORTED
19 runtime-consistent features carry negligible information vs the current label
(max |AUC-0.5| = 0.014); f10/f11 near-degenerate; HTF too slow;
no informative feature group.
B — LABEL MISALIGNMENT : MODERATELY SUPPORTED
threshold reached ~2 bars vs the 24-bar horizon; label = 24-bar drift-sign with
overlap persistence; no SMC conditioning; path noise (MFE 3.8 ATR)
far above the threshold.
C — SETUP-CONDITIONING : NOT SUPPORTED (as the cause of the lost signal)
D1_any_event == D0 (universal events); conditioning shifts base rates
but does not raise discrimination; only EQH/BUY_cand minor positive.
D — TEMPORAL INFORMATION : NOT SUPPORTED
persistence = overlap artifact; lagged info flat; state-age flat.
E — REGIME DEPENDENCY : NOT SUPPORTED
P(Y|regime) lift ~1.0; feature AUC stable across regimes.
F — MULTIPLE FACTORS : OFFICIAL CLASSIFICATION
A dominant + B secondary. A and B are related: we cannot yet separate
"bad features" from "label does not measure the SMC outcome" without a
label-redesign experiment.
G — INCONCLUSIVE : not chosen (sufficient evidence for F)
```
**Official P3.1 classification: `F — MULTIPLE FACTORS (A dominant, B secondary)`.**
C/D/E are rejected by the evidence. This is NOT a license for more complex models —
quite the opposite (see section O).
---
## N. EVIDENCE CLASSIFICATION
```text
TRUSTED : P2 checkpoint (SHAs re-verified, dataset hash matches)
frozen-model reproduction (0.5105/0.5063 == P2.6)
all tables A-J (deterministic scripts, provenance recorded,
JSON outputs: ml/p3/output/*.json)
INCONCLUSIVE : EQH +0.040 & BUY_cand +0.031 (small n, multiple testing)
whether features are informative against a REDESIGNED label (not yet tested)
INVALID : none
SUPERSEDED : old baseline 0.671/0.658 (semantics bug) — still NOT used
```
---
## O. NEXT RESEARCH GATE
```text
P3.1 gate results (rule §23):
D (temporal) -> LSTM : NOT VIABLE (evidence rejects)
E (regime) -> regime ML : NOT VIABLE (evidence rejects)
C (conditioning) -> meta-label: NOT VIABLE (weak/negative evidence)
B (label misalign) -> LABEL CONTRACT MUST BE REDISCUSSED (supported)
A (feature failure) -> FEATURE ARCHITECTURE REVISION (supported)
DECISION:
1. STOP all model development (LSTM/Informer/MS-GARCH/gate/ensemble) —
no empirical foundation.
2. Next priority = LABEL CONTRACT REDESIGN (B), not a new model:
- define the outcome RECORDED ON THE EVENT SETUP (sweep/choch/eqh-eql),
not on all bars; horizon matched to the time-to-hit distribution (e.g. 2-8 bars)
- evaluate asymmetric TP/SL & path-based outcomes (MFE/MAE) as alternatives
- test whether features become informative against the new label (re-run feature audit)
3. In parallel (A): audit which features SHOULD capture "SMC setup quality"
(e.g. sweep size, distance to liquidity, clean CHoCH, order-block) —
revise FEATURE_CONTRACT only if the new label yields results (not now).
4. Only after the new label is proven informative should the next architecture
be considered (meta-label / temporal / regime) with a valid baseline.
```
---
## P. SUCCESS CRITERIA — STATUS
| # | Criterion | P3.1 Answer |
|---|----------|--------------|
| 1 | Do the 19 features have information content? | NO (negligible; max \|AUC-0.5\|=0.014) |
| 2 | Is the 24-bar ATR label fit for the SMC purpose? | PARTLY NO (t_hit 2 bars; overlap; no SMC conditioning) |
| 3 | Does candidate conditioning raise density? | NO (consistently) |
| 4 | Is the signal temporal? | NO EVIDENCE (overlap artifact) |
| 5 | Does the relationship change per regime? | NO EVIDENCE |
| 6 | Which architecture part is worth testing next? | Label contract redesign → feature re-audit (not a new model) |
---
## Appendix — Artifacts & Provenance
```text
Scripts (ml/p3/) : p3_common.py, feature_audit.py, label_audit.py,
candidate_setup_audit.py, temporal_diagnostic.py,
regime_diagnostic.py
Outputs (ml/p3/output/) : feature_audit.json, label_audit.json,
candidate_setup_audit.json, temporal_diagnostic.json,
regime_diagnostic.json
This report : docs/P3_FEATURE_LABEL_FORENSIC.md
Data : Files\AlgoForge\Data\XAUUSD_*.npz (hash D4DB132D.../d25763da.../d0ce3fd1.../f2969c61...)
P2.6 dataset (F+label+A+t) : e85a08616191fb7ce4bbe79c1754228ddbcf54562034a322747501c7aabe124e
Frozen model : 06df8452a112290ecb9bae1a3dbff6df492e084f8b872885d808bb0303ec6a70
Random seed : not used in diagnostics (all deterministic)
Multiple testing : documented (section L)
```
**Closing note:** P3.1 completed without production changes. FEATURE_CONTRACT,
runtime semantics, and the baseline model were NOT touched. All findings are
diagnostic; no "new feature" or "new threshold" is promoted.