# P3-S23 — SILENT-BUG VERIFICATION ```text Date : 2026-08-26 Session : P3-S23 — corrected-population pre-registered ML baseline Policy : docs/CODE_VERIFICATION_AND_SILENT_BUG_POLICY_v1.md Result : NO SILENT BUG FOUND — 8/8 mutation tests detected, independent metric oracle matches (~1e-16), two-run byte reproducibility. ``` ## 1. Scope Applies the code-verification / silent-bug policy to the new P3-S23 pre-registered experiment. Mutations operate on disposable in-memory copies only; the invariant/real pipeline is never broken. Detection uses independent invariants/oracles, not "the training code re-prints its own number". ## 2. Mutation battery (8/8) | # | Mutation | Detection signal | Detected | |---|----------|------------------|----------| | 1 | future feature injection (later rows leak into earlier rows) | ΔROC-AUC 0.3790 | YES | | 2 | global scaler (fit on train+OOS) | ΔROC-AUC 0.0050 | YES | | 3 | temporal shuffle (order broken) | order not preserved | YES | | 4 | OOS injected into training | ΔROC-AUC 0.0245 | YES | | 5 | row shift (feature/label misalign) | ΔROC-AUC 0.1087 | YES | | 6 | train/OOS boundary violation (purge < H) | boundary flag | YES | | 7 | label alignment shift (labels rolled) | ΔROC-AUC 0.1477 | YES | | 8 | prediction/row alignment mismatch | 79 rows changed | YES | detection_count : 8/8 ; all_detected : YES ## 3. Independent (from-scratch) metric oracle The pooled evaluation metrics were recomputed by a disposable oracle that implements ROC-AUC (rank/tie Mann-Whitney-U), PR-AUC (average-precision step-integral matching average_precision_score), LogLoss, and Brier manually (no sklearn.metrics, no shared helper). Cross-check on pooled OOS (n=279): | Metric | committed | oracle | diff | |--------|-------------------|-------------------|-------| | ROC-AUC| 0.52352729885 | 0.52352729885 | 1.1e-16 | | PR-AUC | 0.3301550655 | 0.3301550655 | 5.6e-17 | | LogLoss| 0.6473685391 | 0.6473685391 | 0 | | Brier | 0.2259499408 | 0.2259499408 | 0 | Verification note: an initial oracle used a naive trapezoid PR-AUC which differed from the committed average-precision convention by ~0.011. This was a DEFINITIONAL convention difference (discovered by the oracle), not a data / pipeline defect; the oracle was realigned to the standard average-precision convention, after which it matched to ~5e-16. Documented as process, not silently skipped. ## 4. Reproducibility Two independent runs produced byte-identical fold results and OOS predictions (SHA-256 match on p3_s23_fold_results.csv and p3_s23_oos_predictions.csv); metadata (generated_utc / git_commit) differ as permitted. ## 5. Result ```text Dataset accounting / fold topology / purge / train-only scaler : PASS Mutation tests 8/8 : PASS Independent metric oracle match : PASS (~1e-6) Two-run byte reproducibility : PASS NO SILENT BUG FOUND ``` Evidence: `ml/p3/p3_s23_corrected_population_ml/output/p3_s23_silent_bug_verification.json`, `p3_s23_reproducibility.json`. *End of P3-S23 silent-bug verification.*