# P3-S25.1F - 512 MiB WORKER-SCALING SAFETY BENCHMARK - FINAL REPORT ```text DOCUMENT ROLE : REPORT (derived summary; generated from machine evidence under ml/p3/p3_s251_external_ingest/workload_benchmark/worker_scaling/output/*.json) AUTHORITATIVE SOURCE : worker_scaling machine evidence + canonical real checkpoint output/s251_real_checkpoint.json PHASE : P3-S25.1F - WORKER-SCALING SAFETY BENCHMARK (512 MiB) CLASSIFICATION : A - PARALLEL WORKER SCALING VERIFIED CURRENT FACTUAL STATE : see sections 18-23 (real checkpoint UNCHANGED) AUTHORIZATION STATE : disposable benchmark ONLY; CHUNK 12 NOT AUTHORIZED; SAFE_MAX_CHUNK_SIZE and SAFE_CHUNKS_PER_TURN UNCHANGED UNRESOLVED ITEMS : see section 25 NEXT OWNER DECISION : see section 26 ``` ## 1. Starting Git SHA ```text starting_sha = 05a8a3e1522bacc8f66f5ff6c0b7ddbd5044d926 (verified before execution: git rev-parse HEAD == git ls-remote origin HEAD = 05a8a3e1522bacc8f66f5ff6c0b7ddbd5044d926, branch = main, working tree CLEAN, stash EMPTY, no untracked files; preconditions PASS) ``` ## 2. Final Git SHA ```text final_sha = 972f3aafac2622aaf2c3ab3091f5f5ce9b92ba6a (branch = main, HEAD == origin/main, working tree CLEAN, stash EMPTY verified after push) ``` ## 3. Benchmark fixture ```text The ALREADY-VALIDATED disposable 512 MiB workload from P3-S25.1E (identical reference data, parser, aggregation semantics, atomic chunk semantics, boundary definitions, expected outputs). fixture path : D:\TradingTerminal\HFM Metatrader 5\MQL5\Shared Projects\SniperGold_ML_authoritative\ml\p3\p3_s251_external_ingest\scratch\p3_s251e_primary_fixture_24MiB.csv fixture size : 2149580843 bytes fixture sha256: 2a2691159f1a250b537c82a623539190a59ffd8ea4089ec54245a8f8a7800be5 sha matches P3-S25.1E recorded (2a2691159f1a250b...): True seed : 20260830 workload : 512 MiB = 536870912 bytes, range [0, 536870912) atomic chunk : 24 MiB = 25165824 bytes (SAFE_MAX_CHUNK_SIZE, UNCHANGED) actual line-safe atomic chunks : 22 (each 24 MiB read extends to the next LF) Real source : XAUUSD_mt5_ticks.csv (34,473,661,010 B) referenced ONLY for the committed checkpoint; NO real bytes beyond 188,743,943 were read; CHUNK 12 NOT processed. ``` ## 4. Reference result (canonical single-stream, fresh machine run) ```text rows_processed = 11422782 m15_count = 120112 m30_count = 108000 first_timestamp = 1483403400 last_timestamp = 1697238205 malformed = {"bid_ask_relationship": 0, "column_count": 2, "invalid_volume": 0, "non_monotonic_timestamp": 0, "non_numeric_price": 0, "non_positive_price": 0, "timestamp_invalid": 1, "timestamp_malformed": 1} m15_sha256 = 94714104fc23eef83e379a138ef70775c43865419c6dc8afce05b94c9388a95a m30_sha256 = 238776d863734092c65023817f0ce53acf0d0b2e32510cd63d89806843e1a7e7 output_sha256 = cfa8c084f9bae714a4a15cb1da27e2442bd72bd11aee3f3604304a203e02ed4a final_state_hash = 757a4d63a0f13c2757281402d5beea4278c982c2e29e36375faa30b9565994ff matches P3-S25.1E recorded 512 MiB reference : True reference elapsed_seconds : 342.776 reference peak_memory_bytes : 141534283 output_carry_m15 : {"close": 1209.351, "first_ts": 1697238000, "high": 1209.974, "last_ts": 1697238205, "low": 1208.583, "open": 1208.583, "ticks": 21, "ts": 1697238000} output_carry_m30 : null ``` ## 5. worker=1 result ```text run_ids : workers_1_run_001, workers_1_run_002 n_worker_tasks : 1 all_workers_pass : True sig_run1 : rows 11422782 | M15 120112 | M30 108000 | first 1483403400 | last 1697238205 output_sha256 : cfa8c084f9bae714a4a15cb1da27e2442bd72bd11aee3f3604304a203e02ed4a final_state_hash: 757a4d63a0f13c2757281402d5beea4278c982c2e29e36375faa30b9565994ff chunk_boundaries equal authoritative : True process_seconds_run1 : 308.901 s process_seconds_run2 : 313.360 s semantic_equals_reference_run1 : True semantic_equals_reference_run2 : True output_carry_m15 : {"close": 1209.351, "first_ts": 1697238000, "high": 1209.974, "last_ts": 1697238205, "low": 1208.583, "open": 1208.583, "ticks": 21, "ts": 1697238000} output_carry_m30 : null ``` ## 6. worker=2 result ```text run_ids : workers_2_run_001, workers_2_run_002 n_worker_tasks : 2 all_workers_pass : True sig_run1 : rows 11422782 | M15 120112 | M30 108000 | first 1483403400 | last 1697238205 output_sha256 : cfa8c084f9bae714a4a15cb1da27e2442bd72bd11aee3f3604304a203e02ed4a final_state_hash: 757a4d63a0f13c2757281402d5beea4278c982c2e29e36375faa30b9565994ff partition_chunk_groups : [[0, 11], [11, 22]] worker_boundary_bytes : [[0, 276824537], [276824537, 536870912]] completion_order_run1 : [1, 0] process_span_run1 : 161.502 s process_span_run2 : 163.958 s semantic_equals_reference_run1 : True semantic_equals_reference_run2 : True output_carry_m15 : {"close": 1209.351, "first_ts": 1697238000, "high": 1209.974, "last_ts": 1697238205, "low": 1208.583, "open": 1208.583, "ticks": 21, "ts": 1697238000} output_carry_m30 : null ``` ## 8. worker=4 result ```text run_ids : workers_4_run_001, workers_4_run_002 n_worker_tasks : 4 all_workers_pass : True sig_run1 : rows 11422782 | M15 120112 | M30 108000 | first 1483403400 | last 1697238205 output_sha256 : cfa8c084f9bae714a4a15cb1da27e2442bd72bd11aee3f3604304a203e02ed4a final_state_hash: 757a4d63a0f13c2757281402d5beea4278c982c2e29e36375faa30b9565994ff partition_chunk_groups : [[0, 6], [6, 12], [12, 17], [17, 22]] worker_boundary_bytes : [[0, 150995197], [150995197, 301990405], [301990405, 427819745], [427819745, 536870912]] completion_order_run1 : [3, 2, 1, 0] process_span_run1 : 90.428 s process_span_run2 : 91.037 s semantic_equals_reference_run1 : True semantic_equals_reference_run2 : True output_carry_m15 : {"close": 1209.351, "first_ts": 1697238000, "high": 1209.974, "last_ts": 1697238205, "low": 1208.583, "open": 1208.583, "ticks": 21, "ts": 1697238000} output_carry_m30 : null ``` ## 12. worker=8 result ```text run_ids : workers_8_run_001, workers_8_run_002 n_worker_tasks : 8 all_workers_pass : True sig_run1 : rows 11422782 | M15 120112 | M30 108000 | first 1483403400 | last 1697238205 output_sha256 : cfa8c084f9bae714a4a15cb1da27e2442bd72bd11aee3f3604304a203e02ed4a final_state_hash: 757a4d63a0f13c2757281402d5beea4278c982c2e29e36375faa30b9565994ff partition_chunk_groups : [[0, 3], [3, 6], [6, 9], [9, 12], [12, 15], [15, 18], [18, 20], [20, 22]] worker_boundary_bytes : [[0, 75497593], [75497593, 150995197], [150995197, 226492801], [226492801, 301990405], [301990405, 377488009], [377488009, 452985613], [452985613, 503317349], [503317349, 536870912]] completion_order_run1 : [7, 6, 0, 2, 4, 5, 1, 3] process_span_run1 : 47.468 s process_span_run2 : 47.644 s semantic_equals_reference_run1 : True semantic_equals_reference_run2 : True output_carry_m15 : {"close": 1209.351, "first_ts": 1697238000, "high": 1209.974, "last_ts": 1697238205, "low": 1208.583, "open": 1208.583, "ticks": 21, "ts": 1697238000} output_carry_m30 : null ``` ## 9. Semantic equivalence (parallel output == single-stream reference) ```text workers=1 run1==ref True | run2==ref True | PASS True workers=2 run1==ref True | run2==ref True | PASS True workers=4 run1==ref True | run2==ref True | PASS True workers=8 run1==ref True | run2==ref True | PASS True contract: exact equality on rows / malformed counts / first & last timestamp / M15 count, M30 count / M15+M30+output hashes / final state (carry) hash - no approximate comparison. RESULT: PASS (all worker counts) ``` ## 10. Boundary audit ```text chunk-level 6-class audit (21 boundaries) all_classes_pass : True classes: inside_m15, inside_m30, exact_m15, exact_m30, same_ts, carry_boundary -> 6/6 PASS workers=1 worker-level PASS True | byte-contiguous True | boundary-sequence==authoritative True | carry-chain==canonical True workers=2 worker-level PASS True | byte-contiguous True | boundary-sequence==authoritative True | carry-chain==canonical True workers=4 worker-level PASS True | byte-contiguous True | boundary-sequence==authoritative True | carry-chain==canonical True workers=8 worker-level PASS True | byte-contiguous True | boundary-sequence==authoritative True | carry-chain==canonical True RESULT: 6/6 PASS for every worker count; no gaps, no overlaps; actual_byte_end(worker_n) == actual_byte_start(worker_n+1) at every worker boundary. ``` ## 11. Interruption / resume ```text (a) CANONICAL corrected actual-byte-end resume (P3-S25.1D-R semantics; tail re-executed) at the first worker-boundary chunk index of each worker count: workers=2 stop_after_chunk=11 match=True per_chunk=True bounds=True aggregate=True workers=4 stop_after_chunk=6 match=True per_chunk=True bounds=True aggregate=True workers=8 stop_after_chunk=3 match=True per_chunk=True bounds=True aggregate=True all_match : True (b) PARALLEL-LEVEL resume at EVERY worker boundary (prefix merged carry == canonical chunk carry at that boundary; resume byte start == actual line-safe byte end; shuffled completion-order merge == continuous): workers=2 PASS True entries=1 workers=4 PASS True entries=3 workers=8 PASS True entries=7 parallel_all_match : True RESULT: PASS (continuous == interrupted + resumed for semantic output, final state, boundary sequence and carry chain; worker completion order does not influence resume) ``` ## 12. Mutation detection (10 classes) ```text classes: skipped atomic chunk; duplicated atomic chunk; wrong byte offset; wrong worker ordering; wrong input carry; post-chunk carry as input; output-row deletion; output-row duplication; semantic output mutation; boundary-sequence mutation. workers=1 10/10 detected (chunk-level (canonical single-stream)) workers=2 10/10 detected (worker-level) workers=4 10/10 detected (worker-level) workers=8 10/10 detected (worker-level) all_detected : True RESULT: 10/10 DETECTED for every worker count selected for safety evaluation. A reduction in mutation-detection capability would classify the worker configuration as NOT SAFE; none occurred. ``` ## 13. Reproducibility (two complete runs per worker count) ```text semantic equality run1 == run2 required; timing equality NOT required (nor treated as correctness evidence by itself). workers=1 run1==run2 True | output_sha run1 cfa8c084f9bae714 run2 cfa8c084f9bae714 workers=2 run1==run2 True | output_sha run1 cfa8c084f9bae714 run2 cfa8c084f9bae714 workers=4 run1==run2 True | output_sha run1 cfa8c084f9bae714 run2 cfa8c084f9bae714 workers=8 run1==run2 True | output_sha run1 cfa8c084f9bae714 run2 cfa8c084f9bae714 all_pass : True RESULT: PASS ``` ## 14. Resource profile (worker=1..8, run1 processing spans) ```text workers=1 process_span_run1=308.901 s run2=313.360 s | rows/s=36978.8 | peak=229.9 MiB sum_worker_peak=229.9 MiB | cpu_frac=None | imbalance(max/mean)=1.000 | spawn_overhead=0.00 s | merge=0.00 s workers=2 process_span_run1=161.502 s run2=163.958 s | rows/s=70728.6 | peak=229.9 MiB sum_worker_peak=459.8 MiB | cpu_frac=0.971 | imbalance(max/mean)=1.025 | spawn_overhead=0.65 s | merge=3.26 s workers=4 process_span_run1=90.428 s run2=91.037 s | rows/s=126318.9 | peak=229.9 MiB sum_worker_peak=919.5 MiB | cpu_frac=0.89 | imbalance(max/mean)=1.119 | spawn_overhead=0.52 s | merge=3.35 s workers=8 process_span_run1=47.468 s run2=47.644 s | rows/s=240642.4 | peak=229.9 MiB sum_worker_peak=1839.0 MiB | cpu_frac=0.878 | imbalance(max/mean)=1.134 | spawn_overhead=0.54 s | merge=3.46 s metrics legend: process_span = wall-clock of the parallel processing phase (max worker wall); rows/s = rows / process_span; peak = max per-worker tracemalloc peak (bounded by the 24 MiB atomic chunk, flat across worker counts); cpu_frac = total CPU seconds / (span x N); imbalance = max worker wall / mean worker wall; I/O utilization not directly observable (fixture re-read from OS page cache after the boundary scan; CPU dominates). ``` ## 15. Speedup ```text speedup(1 workers) = time(1)/time(1) = 1.000 speedup(2 workers) = time(1)/time(2) = 1.913 speedup(4 workers) = time(1)/time(4) = 3.416 speedup(8 workers) = time(1)/time(8) = 6.508 (process-span basis, machine-derived) ``` ## 16. Efficiency ```text efficiency(1 workers) = speedup(1)/1 = 1.000 efficiency(2 workers) = speedup(2)/2 = 0.956 efficiency(4 workers) = speedup(4)/4 = 0.854 efficiency(8 workers) = speedup(8)/8 = 0.813 ``` ## 17. Smallest SAFE_USEFUL worker count ```text minimum_safe_useful_worker_count = 2 Classified per count: 1 = SAFE_USEFUL_BASELINE; 2/4/8 = SAFE_USEFUL (all gates PASS, speedup >= 1.3, efficiency >= 0.4). The smallest count giving a meaningful wall-clock reduction is 2 (1.913x, efficiency 0.956). The largest count is NOT auto-selected. ``` ## 18. Maximum SAFE worker count tested ```text maximum_safe_worker_count_tested = 8 (8 = the largest required candidate; all gates PASS; no more than 8 was tested per brief) ``` ## 19. Checkpoint before / after (real, read-only) ```text real_checkpoint_path : D:\TradingTerminal\HFM Metatrader 5\MQL5\Shared Projects\SniperGold_ML_authoritative\ml\p3\p3_s251_external_ingest\output\s251_real_checkpoint.json checkpoint_sha_before : 76ce95eb1e67b4138302a9c9d318e3a3c9f3e52430ac80c08f1a9900da23d0c1 checkpoint_sha_after : 76ce95eb1e67b4138302a9c9d318e3a3c9f3e52430ac80c08f1a9900da23d0c1 byte_identical : True run_id : P3_S251_REAL_TICK_RUN_001 (untouched) last_completed_chunk : 11 | next_chunk 12 | next_byte_offset 188743943 (values verified from the canonical checkpoint; the real checkpoint MUST NOT change - it did not) ``` ## 20. SAFE_MAX_CHUNK_SIZE ```text SAFE_MAX_CHUNK_SIZE = 25165824 bytes (24 MiB) - UNCHANGED (P3-S25.1C). This benchmark performed no dynamic-chunk change and promotes no atomic size. ``` ## 21. SAFE_CHUNKS_PER_TURN ```text SAFE_CHUNKS_PER_TURN = NOT ESTABLISHED (UNCHANGED as operational policy). This benchmark investigated WORKER PARALLELISM (1/2/4/8 workers over 22 atomic chunks), not chunks-per-turn batching; it changes no operational setting. ``` ## 22. Operational-policy status ```text OPERATIONAL_POLICY : NOT CHANGED Real Chunk 12 execution : NOT AUTHORIZED and NOT PROCESSED New real Tickstory bytes : NONE read (real reads never passed 188,743,943; the benchmark read only the disposable synthetic fixture) Dukascopy : NOT accessed | ML : NOT run | Production MQL5/MQH : UNCHANGED Frozen contracts / feature / label contracts : UNCHANGED CURRENT_DECISION_GATE.md / CURRENT_PROJECT_STATE.md : UNCHANGED Historical scientific evidence : UNCHANGED ``` ## 23. CHUNK 12 status ```text CHUNK 12 = NOT AUTHORIZED (CURRENT_DECISION_GATE unchanged). next_byte_offset 188743943 is FACT, not authorization. This benchmark produced performance evidence only; it grants no ingestion authority. ``` ## 24. Failed tests and preserved failure evidence ```text No worker-count candidate failed any gate in the final evidence set: 1/2/4/8 all PASS semantic equivalence, boundary 6/6, interruption/resume, 10/10 mutations, reproducibility and resource practicality (classification A). Process note (transparent, no erased history): an early run of the benchmark harness recorded C - PARALLEL UNSAFE because of a harness METADATA bug (N_ATOMIC = 21 derived by integer division 536870912//25165824, while the actual line-safe chunk count is 22). The worker pipeline itself was correct throughout; the metadata constant and the affected gate bookkeeping were corrected and the downstream worker runs were re-executed in full. No correctness failure of any worker candidate occurred; therefore no failure artifact needed preservation under the separate run-identifier rule. The corrected evidence superseded the metadata-corrected runs; no failed run was overwritten to fabricate a PASS. ``` ## 25. Limitations ```text * Fixture: all worker runs use the deterministic SCHEMA-IDENTICAL SYNTHETIC disposable fixture (the real committed prefix is ~156 MiB and real-source expansion is forbidden). The exercised mechanics (parser, aggregator, line-safe chaining, carry, boundary scan, coordinator merge) are the canonical code, but byte-density / malformed profile of real Tickstory data may differ slightly. * Parsing is pure-Python; parallelism uses process workers (the GIL would serialize CPU-bound threads). Results are therefore machine/OS dependent. * Per-worker peak memory (~230 MiB) is bounded by the 24 MiB atomic chunk; total memory scales ~linearly with worker count (8 workers ~= 1.8 GiB), still far below the machine's available RAM (38.2 GiB available at run start). * I/O utilization is not directly observable; the fixture is re-read per worker from OS page cache after the boundary scan, so CPU dominates. * The order-dependent malformed class non_monotonic_timestamp is reconciled across worker boundaries by the coordinator (first-ts vs last-ts check); the fixture is strictly monotonic, so the count is 0 and the check is exact but not stress-tested. * The lean verifier (no per-line chunk digest) is inherited from P3-S25.1E; its equality with the canonical full-digest verifier was proven there on a 72 MiB range. * The benchmark measured PROCESS, VERIFY, RESUME, MERGE and EVIDENCE phases; total turn-cycle time is the sum of the machine-recorded phase durations. ``` ## 26. Recommended worker configuration ```text RECOMMENDED : 2 workers (smallest SAFE_USEFUL count) speedup 1.913x / efficiency 0.956; ~48% less processing wall-clock vs 1 worker; memory ~2x the single-worker budget (2 x 230 MiB); minimal failure surface; spawn overhead 0.65 s; merge 3.26 s. Verified-safe alternatives: 4 workers (3.416x / 0.854) and 8 workers (6.508x / 0.813) also pass every gate; choose them only if wall-clock dominates and memory/task overhead is acceptable (8 workers ~= 1.8 GiB total, ~0.54 s spawn, ~3.46 s merge). DO NOT exceed 8 without a new benchmark (brief: test 1/2/4/8 only; scaling is under investigation, not maximum thread count). This recommendation is EVIDENCE ONLY - it does not authorize any change to SAFE_MAX_CHUNK_SIZE / SAFE_CHUNKS_PER_TURN / real ingestion. ``` ## 27. Classification ```text PHASE CLASSIFICATION: A - PARALLEL WORKER SCALING VERIFIED A - PARALLEL WORKER SCALING VERIFIED: at least one worker count >1 is SAFE_USEFUL and provides meaningful performance improvement without loss of correctness (2/4/8 all SAFE_USEFUL; smallest = 2). The question of the phase is answered: YES - worker parallelism makes the already-verified 512 MiB workload substantially faster while preserving the exact result, resumability and verification capability. ```