# P3-S22.2 — DISCREPANCY IMPACT ASSESSMENT ```text Date : 2026-08-25 Session : P3-S22.2 — DISCREPANCY IMPACT ASSESSMENT & REPAIR DECISION Status : COMPLETE (research-only; NO repair, NO production modification) Policy : docs/CODE_VERIFICATION_AND_SILENT_BUG_POLICY_v1.md Repair : NONE implemented in this phase (repair decision matrix only) Handover : docs/SESSION_HANDOVER_2026-08-25_P3_S22_1_VERIFICATION_DEBT_CLOSURE.md (newest authoritative handover by Git ancestry; read completely FIRST; predecessor chain traced through P3-S22 -> S21.R -> ...) Starting SHA : 0f1421b0d0bd6e9c5c858ee9bc1fb02a18f3599f (local==origin/main) ``` ## Scope Assess, with independent machine-readable evidence, whether the two discrepancies found in P3-S22.1 materially affect: 1. F2 zone state semantics, 2. F3 Candidate Setup creation, 3. F4 MTF gating, 4. Candidate Setup population, 5. P3-S16 v1 labels, 6. P3-S18/P3-S20 baseline ML results, 7. production-facing behavior. The two discrepancies: - **ADJ-1** — production `AF_FVGZoneState`/`AF_OBZoneState` never assign `mit_state=PARTIALLY_FILLED(1)`; partial fills remain `UNMITIGATED(0)` with the auxiliary `partial_filled` flag. - **ADJ-2** — the committed research M30 gate is built by index-pair resampling of the M15 feed (`m30_from_m15`), which drifts from the runtime UTC-clock M30 at session/data gaps. ## Method Each discrepancy was assessed independently with: - an independent oracle (fresh implementation in this phase, not the audited code under test), - controlled synthetic fixtures (human-auditable), - full-scope deterministic measurement on the frozen research feed, - counterfactual replay of the FROZEN F3 chain (canonical_oracle / F3SetupEngine, P3-S17R2 FULL PARITY) with only the disputed input changed, - mutation tests on disposable copies, - two-run byte reproducibility (runner). Evidence: `ml/p3/p3_s222_discrepancy_impact/output/*.json/csv` + `ml/p3/p3_s222_discrepancy_impact/README.md`. --- ## PART A — ADJ-1 IMPACT ASSESSMENT ### Independent oracle and cases A fresh canonical zone-contract evaluator was written in this namespace (partial fill = `mit=1`, ACTIVE; full fill = terminal; FVG wick full-fill; OB close-through strict). 14 cases cover the 12 mandated scenarios (untouched, fully mitigated, partial, partial-then-further, partial-then-invalidated, pre-entry partial, post-entry partial, narrow/wide zone, boundary touch, bull/bear symmetry, multiple consecutive partial fills, OB partial and OB close-exactly-on-boundary). For each case the following were compared: - production state representation (`{mit, partial, invalidated}`), - oracle semantic state (contract `{mit=1 for partial}`), - consumer availability (IsActive / invalidation), - F3 Candidate Setup creation with BOTH encodings. ### Results (s222_partA_adj1.json) ```text cases_total : 14 consumer_aligned : 14/14 f3_creation_identical: 14/14 exact_label_equal : 6/14 adj1_label_departures: 8/14 (the ADJ-1 label itself) mutations : 6/6 detected (M_A1..M_A6) ``` The 8 label departures are exactly the partial-fill cases: production encodes `mit=0 + partial_filled=true`, the contract oracle encodes `mit=1`. In every one of those cases the consumer availability and the F3 creation result are **identical**. ### Downstream determination - F2 consumers (`AF_FindFVG`/`AF_FindOrderBlock`) and the F3 layer use `IsActive()`/`zone_consumable`, which treat `UNMITIGATED(0)` and `PARTIALLY_FILLED(1)` identically. - The F3 `zone_mit_at_creation` field is informational; no feature/label uses `mit_state` (FEATURE_COLS is geometry-based: width/offset/dist/ages/gates). - Mutation suite confirms the oracle DETECTS label-level mutations (force/suppress PARTIALLY_FILLED, alter flag) and semantic mutations (mitigation-bar shift ±1, boundary condition), proving the verification has power: the ADJ-1 label difference is real but behaviorally inert. ### Classification ```text ADJ-1 : A — REPRESENTATIONAL ONLY / NO DOWNSTREAM IMPACT ``` --- ## PART B — ADJ-2 IMPACT ASSESSMENT ### Independent UTC-clock M30 oracle `utc_clock_m30` (s222_common.py) aggregates the M15 feed into fixed UTC 30-minute buckets by wall-clock open time (open = floor(t/1800)*1800, close = open+1800). The frozen Context-agent rule is then evaluated per UTC M30 bucket and mapped to a per-M15-bar gate by the frozen as-of rule (newest CLOSED M30 bar with close_time <= t+900). Oracle self-test: bucket opens are multiples of 1800; 8 consecutive M15 bars form exactly 4 UTC buckets. Agent pipeline validation: feeding the oracle the committed index-pair M30 OHLC reproduces the committed per-bucket dirs with **0 mismatches** (98,516/98,516) — the oracle machinery is byte-faithful and the only difference is the M30 bar construction. ### Measurements (s222_partB_adj2.json) ```text feed bars : 197,032 (scope 194,985) M30 bars index-pair : 98,516 M30 bars UTC-clock : 100,764 (diff +2,248) full-feed gate diffs : 38,399 bars (19.49%) scope gate diffs : 37,651 bars (19.31%) diff attribution (full): continuous (after a phase shift) : 36,320 session_gap : 2,079 mutations : 5/5 detected (M_B1..M_B5) ``` Gap taxonomy of the feed: intraday breaks (3,600/4,500/7,200 s), weekend gaps (~48-51 h), holiday/DST (~24-25 h) — the index-pair construction keeps its pairing phase fixed in data-index space, so after any missing-bar/odd-gap the synthesized M30 bucket boundaries no longer align to wall-clock 30-min marks, and the phase persists across continuous stretches (which is why the majority of gate diffs sit inside "continuous" segments). ### Propagation question **Does the M30 construction difference propagate into F4 gate decisions or Candidate Setup creation?** Answered by Part C replay: the F4/F3 layer reads the m30 gate as a semantic input, and the gate values differ on 19.5% of bars; the counterfactual (below) shows this DOES change which Candidate Setups are created. H4 and M15 gates are not affected. ### Classification ```text ADJ-2 : C — HISTORICAL BEHAVIOR IMPACT CONFIRMED (Candidate Setup creation / population level under runtime-consistent UTC-clock M30 semantics) ``` --- ## PART C — HISTORICAL IMPACT REPLAY (deterministic counterfactual) The frozen chain builder `VPR.chain_parity` (same kernels, same F3 engines) reproduces the committed population exactly when fed the committed gates (694 created / 686 in-scope, per-bar mismatch 0), confirming the replay baseline. Two independent counterfactuals were then applied to the SAME feed: ### C-A — ADJ-1: production mit encoding Every partially-filled active zone re-encoded to production form (`mit=0` + partial flag), everything else identical. ```text entity comparison : 686/686 common, 0 only-base, 0 only-adj1 outcome signature : identical feature rows : 0 changed verdict : IDENTICAL ``` ### C-B — ADJ-2: UTC-clock M30 gate The committed index-pair m30 gate replaced by the independent UTC-clock gate; all other inputs (h4, m15, sweep, choch, zone) byte-identical. ```text entity comparison (chain identity): n_a(committed)=686 n_b(utc)=707 common chain identity : 484 committed only : 202 (would NOT be created under UTC M30) utc only : 223 (would be created instead) outcome signature : DIFFERS shared m30_gate feature changed rows : 0 verdict : DIFFERS ``` Spot-check (`s222_partC_spotcheck.json`): for the 484 shared setups the m30 gate at the creation bar is IDENTICAL (0% diff) and only ~10% of window bars differ; for the 202 committed-only setups 62% have a different creation-bar gate and ~42% of window bars differ; the 223 UTC-only setups formed under similarly differing gates. This confirms the population change is driven by the m30 gate VALUE differences, not an artifact. **Interpretation.** The research chain is the historical source of the 686-row P3-S18/P3-S20 population. Under runtime-consistent (UTC-clock) M30 semantics the frozen F3 rules create a different population (707 in-scope identities; 202 committed identities not reproduced; 223 new). Therefore ADJ-2 is not representational: it changes Candidate Setup creation on the validated historical scope. --- ## PART D — ML IMPACT - **No retraining, no calibration, no artifact rewrite** (frozen P3-S16/P3-S18/P3-S20 results byte-identical; git status confirms only new research artifacts). - **ADJ-1**: 0 feature rows, 0 label rows affected (Part C 686/686 identical; mit_state not a feature/label). - **ADJ-2**: the FROZEN P3-S18/P3-S20 rows are unchanged. However, the historical Candidate Setup population they derive from is M30-construction-dependent: - 202 of 686 committed in-scope chain identities (29.4%) are NOT reproduced under the independent UTC-clock oracle; - 223 different identities appear; - the m30_gate FEATURE of surviving (shared) setups is unchanged (structural identity h4==m30==direction holds), so the row-LEVEL change is a population-membership change, not a feature-value change on survivors. If the research M30 construction were ever repaired (UTC-grid rebuild), the Candidate Setup population and therefore P3-S18/P3-S20 input rows would change; any model result re-derivation is a separate authorized experiment and was NOT performed here. --- ## PART F — SILENT-BUG PROTECTION (summary) | Dimension | Applied | |---|---| | Independent oracle | yes (fresh contract oracle A; UTC-clock M30 oracle B; agent-pipeline 0-mismatch validation) | | Boundary conditions | yes (14 synthetic cases; bucket fixtures; boundary touch) | | Temporal isolation | yes (as-of rule; closed-bar; no future bar) | | Row alignment | yes (chain-identity key; CSV row diff; spot-check) | | Metamorphic | yes (bull/bear symmetry; pair-phase shift; continuous-feed identity index-pair==UTC) | | Mutation | 6/6 (Part A) + 5/5 (Part B) | | Reproducibility | 2 runs byte-identical except generated_utc (runner) | Evidence grades: independent verification (oracles, mutations, replay) and validated evidence (agent pipeline reproduction 0-mismatch + Part C spot-check mechanism) are recorded in the JSON outputs. --- ## External-data / production / deployment status ```text Production MQL5 : NONE changed (verified by git). FEATURE_CONTRACT / labels / TP-SL-horizon : NONE changed. ML : NONE (no retraining / calibration / nonlinear / deployment). External data : NONE (no Tickstory/Dukascopy / substitution). Terminal/tester : NONE executed in this phase (all evidence is Python on committed frozen research artifacts). Deployment/trading : NONE. ``` ## Final classification ```text ADJ-1 : A — REPRESENTATIONAL ONLY / NO DOWNSTREAM IMPACT ADJ-2 : C — HISTORICAL BEHAVIOR IMPACT CONFIRMED (Candidate Setup creation/population under runtime-consistent UTC-clock M30 semantics; frozen P3-S18/P3-S20 rows unchanged) Overall decision on repair: see docs/P3_S22_2_REPAIR_DECISION.md ``` *End of P3-S22.2 impact assessment.*