SniperGold_ML/docs/P3_SMC_SEMANTIC_GOLDEN_DATASET.md

14 KiB

P3 SMC SEMANTIC GOLDEN DATASET — LIQUIDITY SWEEP (PHASE 1)

Date       : 2026-08-22
Session    : P3 — SMC Semantic Validation, Phase 1: LIQUIDITY SWEEP
Status     : RESEARCH / ANNOTATION ONLY — NO production change
Checkpoint : b41059d (P3.2.2) + 1e80833 (handover) — verified, working tree clean
Provenance : P2 8d330343 | contract C44CC6F2 | model 06df8452 | dataset e85a0861

0. EXECUTIVE SUMMARY

This session answers the first question of the validation architecture:

Does the current MQL5 implementation semantically represent "Liquidity Sweep" the same way humans mean it?

Interim answer (to be confirmed by the human golden test):

1. Actual code definition (f7 grab): wick breaks the internal swing + close-back, M15,
   PERMANENT state (never resets) — 99.95% of bars "sweep active".
2. Actual code definition (f10/f11 EQ): equal high/low swing len=50 pairs,
   wick break WITHOUT close-back (rejection withdrawn due to the AUC test),
   monotonic state (~649 bars max).
3. EVENT-VS-STATE BUG CONFIRMED at the population level (hypothesis D):
   f7 = 2.835 onsets -> 196.928 active bars (69.5x); f10 = 25 onsets -> 877 bars (35x);
   f11 = 38 onsets -> 1.617 bars (42.6x).
4. Golden dataset of 60 cases ready (stratified by state age + vol + regime + year),
   machine annotation parity-verified (0 mismatches vs the P2.6 F cache).
5. HUMAN ANNOTATION NOT YET DONE (awaiting reviewers) — comparison.py ready.

1. EXISTING IMPLEMENTATION LOCATIONS (Task 6)

Three Liquidity Sweep implementations found in the codebase:

# Implementation Location Used by ML?
1 DetectLiquidityGrabs (f7 sweep_dir) AlgoForge_Backtest_Baseline.mq5, SniperGold_SMC_ProPlus_v4_4/4_5.mq5, build_features_p2.py YES (FEATURE_CONTRACT f7)
2 DetectEQ (f10 eqh_swept / f11 eql_swept) same as #1 YES (f10/f11)
3 AF_DetectSweep (Engine 2 AlgoForge) Include\AlgoForge\AF_Engine2_Agents.mqh NO (fuzzy agents N/E)

AF_DetectSweep (#3) is used by the Narrative/Entry Engine-2 agents: reference = min/max of the two last fractal 2/2 swings, 8-bar window, with close-back, no ATR, on 4 agent slots (H4/M30/M15/M3). This is NOT the primitive used by ML features.


2. ACTUAL CODE DEFINITION — FORMALIZATION (Task 7)

The following definitions are what the code REALLY executes (not assumptions; extracted from the EA runtime + training parity-verified, see audit_liquidity_sweep.json).

2.1 f7 — DetectLiquidityGrabs (internal liquidity grab)

reference        : internal swing pivot (fractal len=5) — ProcessStructure(INTERNAL_LEN=5)
sweep condition  : bearish grab: isHigh && high[b] > lvl && close[b] < lvl   (buy-side swept)
                   bullish grab: !isHigh && low[b] < lvl && close[b] > lvl   (sell-side swept)
lookback         : 8-bar window after the pivot (b in (p, p+8])
minimum excess   : 0 (every wick > lvl counts; no ATR threshold)
rejection        : PRESENT — close returns below/above the level (part of the condition)
ATR usage        : NO
timeframe        : M15 (chart TF; closed-bar lock)
bar indexing     : i=0 = NEWEST closed bar (Engine 1); absolute [t-649, t-50] for pivots
event emission   : STATE — g_swpDir/g_swpBar persist until a newer grab;
                   NO state-age bound

2.2 f10/f11 — DetectEQ (equal highs/lows swept)

reference        : CONSECUTIVE SWING pivot pairs (fractal len=50), same type (HH / LL)
pair requirement : |bar(p2)-bar(p1)| >= EQ_BARS(3) ; |price(p2)-price(p1)| <= 0.10*ATR(row)
sweep condition  : EQH: any bar b in (p2, r] with high[b] > pr2
                   EQL: any bar b in (p2, r] with low[b] < pr2
lookback         : p1 >= r-649 AND p2 <= r-50 (absolute valid pivots)
minimum excess   : 0
rejection        : NONE (close-back WITHDRAWN — see §5.2)
ATR usage        : YES — tolerance = 0.10 * ATR(row bar r), NOT pivot ATR
timeframe        : M15
bar indexing     : absolute (same as the EA)
event emission   : MONOTONIC STATE — eqh_swept/eql_swept = 1 for r >= first_cross
                   until the pair leaves the window [r-649, r-50]

3. EVENT VS STATE TEST — RESULTS (Task 17, HYPOTHESIS D)

Quantitative test on the full XAUUSD M15 population 2017–2026 (197.032 bars):

f7 (grab)        : onset events = 2.835 | active bars = 196.928 (99.95%)
                   repetition ratio = 69.5x
                   state NEVER returns to 0 after the first grab
                   a single grab dominates the state: median 94 bars, max 996 bars
f10 (EQH)        : onsets = 25 | active bars = 877 | repetition 35.1x
                   state run median 14 bars, max 384 bars
f11 (EQL)        : onsets = 38 | active bars = 1.617 | repetition 42.6x
                   state run median 14 bars, max 504 bars
f9-confirm stream: E_BUY 27.096 bars in 437 runs (median run 40, 62x)
                   E_SELL 25.259 bars in 400 runs (median run 42, 63x)

Conclusion: SEMANTIC STATE/EVENT BUG — CONFIRMED (hypothesis D).

A condition becomes TRUE at t0 (onset) and stays TRUE at t1..tn.
Expected : ONE event at the onset.
Current  : n events (state counted as a repeated event).

This explains the P3.2.2 findings (retention 1.5%, cluster 46 bars, E_EQH asymmetry artifact): the existing event stream = state counted repeatedly, not discrete setup units.


4. GOLDEN DATASET — 60 CASES (Tasks 8-9)

4.1 Sampling (sample_cases.py, seed 42)

N             : 60 (within the 50-100 range)
Symbol/TF     : XAUUSD / M15 (decision); H4/M30/M15/M3 context for humans
Min spacing   : 96 bars (24 hours) between cases
Distribution  : strata of state age + EQ; vol LOW/MID/HIGH; regime TREND_UP/DN/
                SIDEWAYS/RANGE; years 2018-2026
Stratum n Semantic meaning
FRESH_GRAB 15 machine YES — fresh grab (age<8), rejection PRESENT
AGED_GRAB 10 machine YES — state 16-40 bars (past setup age)
STALE_GRAB 15 machine YES — state >=40 bars (stale)
EQH_SWEPT 6 machine YES — EQH, rejection ABSENT
EQL_SWEPT 6 machine YES — EQL, rejection ABSENT
EQ_NEARCROSS 6 machine NO (EQ) — valid pair not yet crossed
EQ_MARGINAL 2 machine YES —

Note: a pure "machine NEGATIVE" stratum (f7=0 & f10=0 & f11=0) does NOT exist in the population (only 104 bars, 0.05%, all before the first 2017 grab) — that fact itself is a semantic finding (permanent state).

4.2 Machine annotation (machine_annotator.py)

Output        : output/machine_annotations.csv (60 rows)
Parity        : 0/60 mismatches vs the P2.6 F cache (replication = existing code)
machine_decision : YES = 60/60 (f7 active 99.95% of the population)
primitives    : f7_grab 46, f11_eql 8, f10_eqh 6 (primary: EQ > grab)
Fields        : case_id, decision_timestamp, symbol, decision_tf,
                reference_type, reference_level, sweep_price, excess,
                close_price, rejection, timeframe, machine_decision,
                machine_reason + per-primitive JSON detail
N/A rule      : fields not used by the code = "N/A" (e.g. rejection for f10/f11)

5. ADDITIONAL SEMANTIC FINDINGS

5.1 Rejection (close-back) — inconsistency across primitives

f7 grab  : rejection PRESENT (close returns below/above the level)
f10/f11  : rejection ABSENT — wick break alone = swept

Both primitives coexist in one model with DIFFERENT "sweep" definitions. f7 uses the classic SMC definition (sweep + rejection), f10/f11 use the "break" definition chosen by an AUC test.

5.2 Definition chosen based on backtest (STOP CONDITION §21)

DESAIN_MTF_v45.md records:

"Uji close-back (wick-break + close-balik) -> eqh_swept/eql_swept = 0 di 60.000 bar,
 AUC test 0.6795 -> 0.6664. KESIMPULAN: pemaksaan close-back MENURUNKAN kualitas.
 DIPULIHKAN ke definisi 'tembus' (konsisten model v1815)."

This is exactly the brief's stop condition: a definition chosen by backtest performance, not SMC semantics. It must be adjudicated whether "break without close-back" = Liquidity Sweep in the SMC definition.

5.3 Timeframe — M15 only (hypothesis C)

f7/f10/f11 are only computed on M15. HTF (D1/H4/H1) provides only direction bias (f0-f2), NOT liquidity references. If a human sees a sweep on M30/H4 (HTF liquidity level), the machine cannot capture it -> TIMEFRAME SEMANTIC GAP (tested in comparison.py).

5.4 Windowed vs full-feed f7 — empirically a non-issue

EA replication (700-bar window) vs training (full feed): 0 divergence on 1.500 probe bars. Theoretical divergence exists but does not occur on XAUUSD M15 (grab density keeps the last grab fresh).


6. HUMAN ANNOTATION PROTOCOL (Tasks 10-12)

  1. Blinding: annotators ONLY see cases.csv (case_id, timestamp) + the MT5 chart up to decision_timestamp. cases_meta.json and machine_annotations.csv MUST NOT be opened before the humans finish. (Task 11)
  2. Template: human_annotation_template.csv — columns: liquidity_sweep (YES/NO/AMBIGUOUS), reference (EQH/EQL/Swing/Other/None), direction (Bullish/Bearish/None), timeframe (H4/M30/M15/M3/None), reason, confidence (HIGH/MEDIUM/LOW) + visual->numeric facts: ref_equal_highs, ref_price, sweep_occurred, sweep_price, sweep_excess, close_returned, displacement. (Tasks 10 & 15)
  3. Multi-annotator: at least 2 (Human A / Human B) when possible. Disagreement -> ADJUDICATION REQUIRED (Task 12).
  4. Machine annotation NOT shown before the humans finish. (Task 11)

7. COMPARISON (Tasks 13-14, 16) — ready to run after humans

comparison.py human_A.csv human_B.csv produces:

- Human A vs Human B : agreement rate, disagreement cases (adjudication)
- Machine vs Human   : TP/FP/FN/TN, Precision/Recall/F1 (AMBIGUOUS separate)
- Reference agreement: EQ vs Swing vs Other
- Timeframe agreement: human TF vs machine TF (M15) -> TIMEFRAME SEMANTIC MISMATCH
- Direction agreement: Bullish/Bearish
- SEMANTIC FALSE AGREEMENT: Human=YES & Machine=YES with DIFFERENT reference/
  timeframe/rejection semantics -> NOT full validation

Semantic agreement levels (from the P3 handover):

Level 1 — Outcome agreement     : Human YES & Machine YES
Level 2 — Structural agreement  : reference level, timeframe, direction, rejection, context IDENTICAL
Level 3 — Reason agreement      : human reasoning & machine primitives substantively IDENTICAL
ONLY Level 3 = STRONG SEMANTIC AGREEMENT

8. REGRESSION TEST SPECS (Task 18) — spec only, NOT yet fixed

ID Given When Expected Current Severity
R1 valid EQH pair, fc<=r r=fc..fc+649 ONE event at onset f10==1 for ~649 bars (state) STATE/EVENT BUG (D)
R2 wick sweeps & close returns machine evaluates f10 close-back required f10=1 without rejection DEFINITION AMBIGUITY (B)
R3 grab at bar b r=b+40 sweep considered stale f7 still active (stale state) EVENT/STATE (D)
R4 grab from pivot < r-699 EA vs training f7 parity 0/1500 divergence (empirically OK) LOW
R5 wick breaks by 0.001 ATR f10/f11 evaluation min excess required min excess = 0 SIMPLIFICATION (B)
R6 human sweep on M30/H4 machine evaluates M15 detection on the same TF M15 only; HTF = bias only TIMEFRAME GAP (C)

These specs are NOT implemented in production this session (per protocol).


9. SUCCESS CRITERIA (Task 20) — status

# Criterion Status
1 Definition of Liquidity Sweep used by the code? DEFINED (§2)
2 Same as the reviewer definition? AWAITING human annotation
3 Same timeframe? awaiting humans (machine: M15 only)
4 Same liquidity reference? awaiting humans (machine: internal swing / EQ pair)
5 Same rejection semantics? ⚠️ machine inconsistent (f7 YES, f10/11 NO)
6 Event only at onset? NO — PERMANENT/MONOTONIC STATE
7 Precision/Recall/F1? awaiting humans
8 Main semantic mismatch? ⚠️ state/event + rejection + timeframe gap

10. STOP CONDITIONS (Task 21)

Met (stopped before code modification):

[x] event-vs-state bug (R1/R3) — CONFIRMED
[x] rejection definition ambiguity (R2) — CONFIRMED (chosen via AUC)
[x] timeframe ambiguity (R6) — indicated (machine M15 only)
[ ] high human disagreement — NOT YET MEASURED (awaiting annotators)
[ ] reference ambiguity — awaiting humans

NO definition was chosen based on backtests in this session.


11. NEXT STEPS (after human annotation)

1. Human A & B fill the template (60 cases x 2, blind).
2. comparison.py -> TP/FP/FN/TN, P/R/F1, agreement, false-agreement, TF audit.
3. LIQUIDITY SWEEP SEMANTIC VERDICT: PASS / PASS WITH CAVEAT / FAIL / INCONCLUSIVE.
4. Only then decide whether CHoCH is the next concept (NOT directly OB/FVG).
5. NO MLP / LSTM / regime / meta-label / backtest optimization.

12. ARTIFACTS

docs/P3_SMC_SEMANTIC_GOLDEN_DATASET.md          (this document)
ml/p3/smc_semantic/smc_semantic_common.py
ml/p3/smc_semantic/sample_cases.py
ml/p3/smc_semantic/machine_annotator.py
ml/p3/smc_semantic/human_annotation_template.csv
ml/p3/smc_semantic/comparison.py
ml/p3/smc_semantic/audit_liquidity_sweep.py
ml/p3/smc_semantic/README.md
ml/p3/smc_semantic/output/cases.csv               (for humans; no machine decisions)
ml/p3/smc_semantic/output/cases_meta.json         (BLINDED)
ml/p3/smc_semantic/output/machine_annotations.csv (BLINDED from humans)
ml/p3/smc_semantic/output/audit_liquidity_sweep.json