SniperGold_ML/docs/P3_S22_3_RESEARCH_M30_REPAIR.md

7.5 KiB

P3-S22.3 — RESEARCH M30 CONSTRUCTION REPAIR

Date       : 2026-08-26
Session    : P3-S22.3 — RESEARCH M30 CONSTRUCTION REPAIR &
             POPULATION REDERIVATION
Status     : COMPLETE (research-only; corrected pipeline + verified)
Policy     : docs/CODE_VERIFICATION_AND_SILENT_BUG_POLICY_v1.md
Handover   : docs/SESSION_HANDOVER_2026-08-25_P3_S22_2_DISCREPANCY_IMPACT_ASSESSMENT.md
             (newest authoritative handover by Git ancestry; read completely
              FIRST) -> predecessor chain through P3-S22.2, .1, P3-S21.R ...
Starting SHA : 77764ba4a67f59d5867e5c7b8992289983869606

Objective

Repair the RESEARCH-side M30 gate construction so it is semantically consistent with the verified runtime-consistent UTC-clock M30 semantics, and deterministically re-derive the downstream Candidate Setup population and label/feature population. This is a DATA / PIPELINE CORRECTION phase, not an ML experiment.

PART A — the exact research M30 construction vs runtime

Item Old (committed research) Runtime / corrected
Construction index-pair resample m30_from_m15: 2 x M15 per M30, t30=t15[::2] fixed UTC 30-min buckets t30_open = floor(t15/1800)*1800
Grouping data-index phase (2k,2k+1) wall-clock 30-min grid
At gaps pairing phase drifts from wall clock; persists over continuous stretches bucket identity/phase is fixed to UTC grid
OHLC high=max, low=min of its (tick) M15 pair high/low of every M15 in [open,open+1800)
as-of frozen: newest CLOSED M30 with close_time<=t+900 unchanged (frozen)

Which timestamps are affected: every M15 bar in or after a session/data/ weekend/holiday/DST gap, and the continuous stretches that follow a phase shift. Only the TIMESTAMP/GROUP-ASSIGNMENT differs — the OHLC aggregation mathematics (open-first, high-max, low-min, close-last) is the same rule; the difference is WHICH M15 bars belong to a bucket.

  • Session gaps: intraday breaks (3600/4500/7200 s), weekend (~48-51 h), holiday/DST (~24-25 h or 1 h offset). In every such case index-pair phase breaks from the UTC grid. Handled by the corrected construction via the UTC-clock floor formula (no interpolation, no forward fill).
  • Missing M15 bar: corrected bucket simply omits it; the bucket OHLC reflects the M15 bars actually present. Incomplete (unclosed) final bucket is handled by the as-of closed-bar lock (never emitted as a closed decision input).
  • DST: the feed is UTC-anchored; a +1 h drift test verifies the grid stays on the UTC 30-min grid (regression case "dst_shift").

The invariant (proved)

For every research M15 bar belonging to a COMPLETE UTC M30 bucket:

M30_open = floor(M15_open / 1800) * 1800     (UTC epoch seconds)

Verified on all 100,764 UTC buckets of the corrected construction: 0 mismatches (PART A, output/s223_main_pipeline.json). This is the exact boundary convention of the verified runtime/oracle (P3-S22.2 Part B).

PART B — corrected research M30 construction

Implemented in ml/p3/p3_s223_m30_repair/s223_common.py + s223_partB_oracle.py. Requirements honoured: fixed UTC buckets, deterministic chronological order, explicit incomplete-bucket handling, NO forward fill, NO look-ahead, NO future bars, NO change to F1..F4 / label / features except the direct M30 gate consequence. Isolated & auditable; does not edit committed p3_s17r1_gate_series_cache.npz.

PART C — independent oracle (100% agreement)

oracle_clock_m30 (from-scratch in-order accumulation) + searchsorted as-of is a SEPARATE implementation. It independently computes timestamps / open / high / low / close / bar completeness / M30 direction / M30 gate. Agreement: buckets 100,764 == 100,764; timestamps T, open T, high T, low T, close T, dirs T, gate T -> 100% agreement on all applicable rows (output/s223_partC_agreement.json). The corrected gate also matches the P3-S22.2 validated UTC oracle cache exactly and reproduces the ADJ-2 diff (38,399 full / 37,651 scope).

PART D — regression / mutation

12 regression scenarios (continuous, missing M15, weekend, intraday, holiday, DST, first/last bucket bar, exact boundary, incomplete bucket, out-of-order, duplicate ts) — all invariant_ok, deterministic fingerprints. 6 mutations (1-bar ts shift, row delete, row insert, wrong M30 boundary, index-pair substitution, future-bar injection) all DETECTED. Test suite run twice, byte-identical (output/s223_reproducibility.json).

PART E — frozen baseline reproduction

Committed (index-pair) gates in the SAME frozen chain reproduce 694 created / 686 in-scope / per-bar mismatch 0 EXACTLY. The old research pipeline remains the frozen historical reference; corrected construction does not overwrite it.

PART F — counterfactual corrected replay

Substituting ONLY the corrected M30 gate into the frozen chain:

committed (old) in-scope : 686
corrected in-scope       : 707
shared identity          : 484
old-only                 : 202   (would NOT be created under UTC M30)
corrected-only           : 223   (would be created instead)

This MATCHES the independent P3-S22.2 counterfactual (707 / 202 / 223). It is the direct, deterministic population change caused by the M30 correction.

PART G — label / feature impact bucket

Only the M30 gate input changes. Under the frozen P3-S16 label contract:

shared identity (484) identical label changed label
484 484 0
  • previously absent (corrected-removed): 202
  • newly available (corrected-added): 223
  • UNRESOLVED corrected: 28 ; AMBIGUOUS corrected: 6.
  • The corrected population is larger (707 vs 686) and has more UNRESOLVED and AMBIGUOUS rows. Feature rows: 202 removed / 484 retained / 223 added relative to the frozen old rows (feature contract unchanged).

PART H — population audit

old        : total 686 = leads 594 + follow-ons 92     (reconciles)
corrected  : total 707 = leads 610 + follow-ons 97     (reconciles)

No row duplicated or silently lost.

PART I — old vs corrected dataset

identity Count
intersection 484
old-only 202
corrected-only 223
M30 gate diff full 38,399 bars
M30 gate diff scope 37,651 bars

Formal: old = frozen historical reference; corrected = freshly derived research candidate (NOT called production / final ML dataset — pending validated).

PART J — machine verification

Independent oracle (100%), 12 regression fixtures, 6 mutation tests (all detected), deterministic heavy replay (hash-stable), Part D test suite run twice byte-identical. Evidence machine-readable (JSON + CSV + hashes); no image/OCR.

PART K — deliverables & namespace

  • docs: this file, P3_S22_3_POPULATION_REDERIVATION.md, P3_S22_3_REPAIR_DECISION.md, SESSION_HANDOVER_2026-08-26_P3_S22_3_RESEARCH_M30_REPAIR.md (this handover
    • 3 reports).
  • namespace ml/p3/p3_s223_m30_repair/ (README, construction, oracle, replay runner, tests, mutation, JSON + CSV evidence).

PART L — repair decision

C — CORRECTION VERIFIED, MATERIAL POPULATION CHANGE.
    (686 -> 707 in-scope; 202 removed; 223 added; gate diff ~19.5%)

Production / ML / external-data status

Production MQL5 : NONE changed (verified by git).
F1/F2/F3/F4 / labels / FEATURE_CONTRACT / TP-SL-horizon : NONE changed.
ML  : NONE (no retrain / calibrate / nonlinear / deployment).
Historical P3-S16..P3-S22 artifacts : NOT modified (additive only).
External data : NONE (no Tickstory/Dukascopy).
Terminal / tester : NONE executed (all Python on frozen artifacts).
Deployment / trading : NONE.

End of P3-S22.3 research M30 repair.