# P3-S25.1Q — WORKER-SCALING OPTIMIZATION EXPERIMENT — FINAL REPORT ```text DOCUMENT ROLE : REPORT (derived summary; generated from machine evidence under ml/p3/p3_s251_external_ingest/workload_benchmark/ worker_scaling_q/output/*.json) AUTHORITATIVE SOURCE: P3-S25.1Q machine evidence + canonical real checkpoint output/s251_real_checkpoint.json (READ-ONLY) PHASE : P3-S25.1Q - WORKER-SCALING OPTIMIZATION EXPERIMENT CLASSIFICATION : A - OPTIMIZATION CANDIDATE IDENTIFIED GATES PASS : correctness 14/14, boundary, state/carry, interruption/resume, mutation 10/10, reproducibility, checkpoint protection STARTING GIT SHA : 1e92c291e16de78c2fab562ae7aeba49b18063c2 (branch main, == origin/main, CLEAN, stash EMPTY, no untracked) ``` ## 1. Machine CPU topology (programmatically read, never hard-coded) ```text CPU name : Intel(R) Xeon(R) Gold 5220 CPU @ 2.20GHz (2 sockets) sockets : 2 cores per socket : [12, 12] physical cores : 24 (sum of sockets) logical processors : 24 (os.cpu_count() == 24; no hyper-threading) python : 3.14.6 (AMD64) ``` ## 2. Total RAM (programmatically read) ```text ram_total_bytes = 51,468,738,560 (~47.93 GiB) ram_available at gate start = 35,185,541,120 (~32.8 GiB) ram_available at gate close = 35,148,988,416 (~32.7 GiB) min ram available during 24-worker run (sampled) = 25,218,666,496 (~23.5 GiB) ``` ## 3. Baseline worker count ```text BASELINE_WORKERS = 8 (already-validated configuration: 8 workers, 24 MiB atomic chunks, 512 MiB workload) ``` ## 4. Candidate worker counts ```text 8 (baseline), 12, 16, 20, 24 (exactly the required set; no intermediate counts tested, no expansion beyond 24) ``` ## 5. Disposable dataset identity ```text path : ml/p3/p3_s251_external_ingest/scratch/ p3_s251e_primary_fixture_24MiB.csv (gitignored scratch) size : 2,149,580,843 bytes schema : Tickstory XAUUSD tick schema - YYYY.MM.DD,HH:MM:SS,bid, ask,last,volume (CFG.NCOL=6, UTC, epoch seconds) = schema-identical to the real Tickstory source generation : deterministic; seed 20260830; rows_written 45,735,763 workload : 512 MiB = 536,870,912 bytes, byte range [0, 536870912) atomic semantics : 24 MiB = 25,165,824 bytes (SAFE_MAX_CHUNK_SIZE) actual line-safe atomic chunks : 22 (each 24 MiB read extends to the next LF; last boundary [528483217, 536870912)) immutability : ONE dataset used for EVERY candidate and BOTH runs; a different dataset was never generated per worker count. ``` ## 6. Dataset SHA ```text fixture sha256 : 2a2691159f1a250b537c82a623539190a59ffd8ea4089ec54245a8f8a7800be5 (matches P3-S25.1E/F recorded; verified preflight) raw range sha256 [0,536870912) : 00c0421ac60d4851dbfe79a10cbb0f2fcdc51d978414934f0d761f40514de88f ``` ## 7. Reference result (canonical single-stream, ONE per experiment) ```text rows_processed = 11,422,782 m15_count = 120,112 m30_count = 108,000 first_timestamp = 1,483,403,400 last_timestamp = 1,697,238,205 malformed = {"column_count": 2, "timestamp_malformed": 1, "timestamp_invalid": 1, "non_numeric_price": 0, "invalid_volume": 0, "bid_ask_relationship": 0, "non_positive_price": 0, "non_monotonic_timestamp": 0} m15_sha256 = 94714104fc23eef83e379a138ef70775c43865419c6dc8afce05b94c9388a95a m30_sha256 = 238776d863734092c65023817f0ce53acf0d0b2e32510cd63d89806843e1a7e7 output_sha256 = cfa8c084f9bae714a4a15cb1da27e2442bd72bd11aee3f3604304a203e02ed4a final_state_hash = 757a4d63a0f13c2757281402d5beea4278c982c2e29e36375faa30b9565994ff output carry M15 = {"ts":1697238000,"open":1208.583,"high":1209.974, "low":1208.583,"close":1209.351,"ticks":21, "first_ts":1697238000,"last_ts":1697238205} output carry M30 = null matches P3-S25.1E recorded (known-answer) : True reference elapsed (single stream) : 366.765 s reference peak memory : 141,534,283 B (~135.0 MiB) Canonical atomic chained reference (22 x 24 MiB, lean verifier): sig == single-stream : True | boundary sequence == authoritative : True Canonical full-digest atomic reference (22 x 24 MiB, ChunkParser digests): all status PASS : True | == single-stream : True raw_chunk_sha256[22] + parsed_chunk_sha256[22] recorded parsed_chain_sha256 = ded07b41892d5e3c0e1277865c43a7d1dba0ee32e8c7f803187a8348c4ea9073 ``` ## 8. Correctness result per candidate (14-point contract) ```text Every candidate == canonical single-stream reference (EXACT, no approximation) checks: rows | malformed counts | first ts | last ts | M15 output | M30 output | raw hash | parsed hash | row hash | output hash | final state hash | boundary sequence | byte coverage | carry/state chain workers=8 14/14 PASS (raw hash implied by byte coverage; parsed hash by boundary identity + canonical full-digest equality) workers=12 14/14 PASS workers=16 14/14 PASS workers=20 14/14 PASS workers=24 14/14 PASS ``` ## 9. Boundary result per candidate ```text no gaps / no overlaps : PASS (all candidates) source-byte ordering : PASS (worker_id ascending == source byte order) line-safe boundaries : PASS (authoritative 22-boundary sequence used) boundary sequence == authoritative : PASS (all) worker partition coverage == workload coverage [0,536870912) : PASS (all) ``` ## 10. State / carry result per candidate ```text input_state(N) == predecessor output_state : PASS (prefix merged carry == canonical chunk carry at every worker boundary) output_state(N) == input_state(N+1) : PASS (canonical 22-chunk chain on ACTUAL byte_end; candidate boundary entry hashes == canonical chunk hashes at every boundary) final worker result state == reference : PASS (final_state_hash 757a4d63... for every candidate) worker boundaries verified : 8->7 | 12->11 | 16->15 | 20->19 | 24->21 ``` ## 11. Interruption / resume result per candidate ```text (a) CANONICAL genuine resume (P3-S25.1D-R corrected; tail re-executed) at first worker-boundary chunk (stop_after_chunk=3): per_chunk_equal True | boundary_sequence_identical True | aggregate_equal True | continuous_output_hash == resumed_output_hash == cfa8c084...e4a | continuous_final_state == resumed_final_state == 757a4d63... -> match True (b) PARALLEL resume at EVERY worker boundary: prefix merged carry == canonical chunk carry; resume byte start == actual line-safe byte end; shuffled completion-order merge == continuous merge == reference: workers=8 True | 12 True | 16 True | 20 True | 24 True (c) GENUINE worker-level interruption + resume (tail re-executed): 8 workers resume at byte 75,497,593 -> equals continuous True 16 workers resume at byte 50,331,725 -> equals continuous True continuous == interrupted + resumed : PASS (all) The real production checkpoint was NOT interrupted (disposable only). ``` ## 12. Mutation result per candidate ```text 10/10 detected for EVERY candidate (skipped-atomic-chunk, duplicated-atomic-chunk, wrong-byte-offset, wrong-worker-ordering, wrong-input-carry, post-chunk-carry-as-input, output-row-deletion, output-row-duplication, semantic-output-mutation, boundary-sequence-mutation) The 5 minimum required brief classes (reordered results, wrong boundary, wrong input carry, omitted range, duplicated range) are all DETECTED. ``` ## 13. run_1 vs run_2 reproducibility ```text semantic equality run1 == run2 (timing equality NOT required): 8 PASS | 12 PASS | 16 PASS | 20 PASS | 24 PASS output_sha run1 == output_sha run2 == cfa8c084...e4a for all candidates ``` ## 14. Wall-clock per candidate (process span, final coherent pass) ```text workers=8 span1 = 48.86 s | span2 = 48.15 s workers=12 span1 = 33.73 s | span2 = 32.81 s workers=16 span1 = 33.58 s | span2 = 33.02 s workers=20 span1 = 33.65 s | span2 = 33.59 s workers=24 span1 = 19.54 s | span2 = 19.25 s total elapsed (dispatch wall): 8->53.47 s | 12->36.32 s | 16->36.33 s | 20->37.00 s | 24->21.49 s (run1) ``` ## 15. Rows/sec (span basis, run1) ```text 8 -> 233,790 | 12 -> 338,666 | 16 -> 340,177 | 20 -> 339,496 | 24 -> 584,490 rows = 11,422,782 for every candidate ``` ## 16. Bytes/sec (span basis, run1) ```text 8 -> 10,988,134 | 12 -> 15,917,298 | 16 -> 15,988,325 | 20 -> 15,956,312 | 24 -> 27,471,018 bytes = 536,870,912 for every candidate ``` ## 17. Peak memory per worker (tracemalloc; bounded by 24 MiB atomic chunk) ```text max per-worker peak: 241,050,426 B (~229.9 MiB) for EVERY candidate (worker memory is bounded by the 24 MiB atomic chunk, flat across N) ``` ## 18. Total Python memory ```text total worker peak (sum of per-worker peaks): 8->1,839.0 MiB | 12->2,604.5 MiB | 16->3,524.0 MiB | 20->4,443.6 MiB | 24->4,903.3 MiB total worker RSS (measured, psapi): 8->945.5 MiB | 12->1,303.2 MiB | 16->1,697.6 MiB | 20->2,088.7 MiB | 24->2,282.6 MiB coordinator RSS during run (measured): 8->103.7 MiB | 12->246.7 MiB | 16->246.7 MiB | 20->247.1 MiB | 24->248.2 MiB aggregate Python RSS (workers + coordinator): 8->1,049.2 MiB | 12->1,549.9 MiB | 16->1,944.3 MiB | 20->2,335.8 MiB | 24->2,530.8 MiB NOTE: 1839 MiB x multipliers for the 24 MiB-chunk working set are NOT assumed; every value is measured. ``` ## 19. CPU utilization (machine-wide, sampled during worker phase) ```text workers=8 average 0.616 | peak 0.676 workers=12 average 0.693 | peak 0.760 workers=16 average 0.715 | peak 0.890 workers=20 average 0.754 | peak 0.998 workers=24 average 0.961 | peak 1.000 ``` ## 20. Coordinator overhead (merge + dispatch/spawn) ```text merge (run1): 8->3.41 s | 12->3.45 s | 16->3.31 s | 20->3.39 s | 24->3.27 s dispatch/spawn: 8->0.71 s | 12->0.69 s | 16->0.77 s | 20->0.75 s | 24->1.10 s ``` ## 21. Worker imbalance (max wall / mean wall, run1) ```text 8->1.1345 | 12->1.1363 | 16->1.4846 | 20->1.7611 | 24->1.0638 (12/16/20 are quantisation-bound by 2-chunk stragglers: 22 chunks do not divide evenly; 24 -> 22 workers x 1 chunk -> best balance) ``` ## 22. Actual process count ```text requested=8 actual=8 non-empty partitions=8 (match) requested=12 actual=12 non-empty partitions=12 (match) requested=16 actual=16 non-empty partitions=16 (match) requested=20 actual=20 non-empty partitions=20 (match) requested=24 actual=22 non-empty partitions=22, empty partitions (worker ids 22,23) -> DISCREPANCY RECORDED (22 atomic chunks cannot populate 24 workers; the executor legitimately created 22 processes) Worker PIDs recorded per run; completion order is never used for semantic ordering (merge is by source byte order). ``` ## 23. Speedup (vs 8-worker baseline, span run1) ```text speedup(8) = 1.000 speedup(12) = 48.859 / 33.729 = 1.449 speedup(16) = 48.859 / 33.579 = 1.455 speedup(20) = 48.859 / 33.646 = 1.452 speedup(24) = 48.859 / 19.543 = 2.500 ``` ## 24. Efficiency (vs 8-worker baseline) ```text efficiency(8) = 1.000 efficiency(12) = speedup / (12/8) = 1.449 / 1.500 = 0.966 efficiency(16) = 1.455 / 2.000 = 0.728 efficiency(20) = 1.452 / 2.500 = 0.581 efficiency(24) = 2.500 / 3.000 = 0.833 ``` ## 25. MAX_SAFE_TESTED ```text MAX_SAFE_TESTED = 24 (24 is the largest required candidate; all gates PASS: correctness 14/14, boundary, state/carry, interruption/resume, mutation 10/10, reproducibility, checkpoint byte-identical.) ``` ## 26. PREFERRED_WORKER_COUNT ```text PREFERRED_WORKER_COUNT = 24 Evidence: 2.500x speedup vs the 8-worker baseline at 0.833 efficiency is the ONLY candidate that breaks the 2-chunk-per-worker plateau (12/16/20 all sit at ~1.45x because 22 atomic chunks quantise into 2-chunk straggler workers). 24 also has the best balance (imbalance 1.0638), the highest CPU utilisation (0.961 avg), and the lowest coordinator merge time (3.27 s). Memory 4,903 MiB total worker peak / 2,531 MiB aggregate RSS is ~10% of the 47.93 GiB machine RAM and the minimum available RAM stayed >= 23.5 GiB throughout. The 'prefer fewer workers when close' rule does not bind: 12 (0.966 eff) is the most efficient, but its 1.449x speedup is NOT close to 24's 2.500x, and 12/16/20 are mutually close (~1.45x) so the fewest of THOSE (12) is the natural pick within the plateau cluster. ``` ## 27. Resource trade-off ```text wall speed eff tot worker agg RSS cpu avg imb workers=8 48.86 s 1.000 1.000 1839 MiB 1049 MiB 0.616 1.135 workers=12 33.73 s 1.449 0.966 2605 MiB 1550 MiB 0.693 1.136 workers=16 33.58 s 1.455 0.728 3524 MiB 1944 MiB 0.715 1.485 workers=20 33.65 s 1.452 0.581 4444 MiB 2336 MiB 0.754 1.761 workers=24 19.54 s 2.500 0.833 4903 MiB 2531 MiB 0.961 1.064 24 workers buys 32.3 s less wall-clock per 512 MiB workload than the baseline for +2,664 MiB total worker peak (+1,482 MiB aggregate RSS); 8->24 = -59.9% wall-clock; 12->24 = -42.1% wall-clock for +2,298 MiB peak. ``` ## 28. Failed candidates ```text NONE. All five candidates (8/12/16/20/24) passed every gate. No failure artifact required preservation. ``` ## 29. Checkpoint SHA before / after (real, READ-ONLY) ```text real_checkpoint_path : ml/p3/p3_s251_external_ingest/output/s251_real_checkpoint.json checkpoint_sha_before = 8c6cad1ca561308e0ba91fcdca3f8ebb95e6d79abc55edaf514d6565f6ce4010 checkpoint_sha_after = 8c6cad1ca561308e0ba91fcdca3f8ebb95e6d79abc55edaf514d6565f6ce4010 byte_identical : True run_id : P3_S251_REAL_TICK_RUN_001 (untouched) last_completed_chunk : 187 | next_chunk 188 | next_byte_start 4483711426 (UNCHANGED) ``` ## 30. Operational-policy status ```text OPERATIONAL_POLICY : NOT CHANGED SAFE_MAX_CHUNK_SIZE : 24 MiB (UNCHANGED) SAFE_CHUNKS_PER_TURN : NOT ESTABLISHED (UNCHANGED) production worker configuration : UNCHANGED CURRENT_DECISION_GATE.md / CURRENT_PROJECT_STATE.md : NOT MODIFIED production MQL5/MQH : UNCHANGED | frozen contracts : UNCHANGED historical evidence : UNCHANGED real checkpoint : byte-identical before/after CHUNK 188 : NOT PROCESSED / NOT AUTHORIZED new real Tickstory bytes : NONE read (benchmark read only the disposable synthetic fixture [0, 536870912)) Dukascopy : NOT accessed | ML : NOT run ``` ## 31. Recommendation for possible future production use ```text EVIDENCE ONLY - no operational promotion is made by this experiment. If the owner later authorises a production worker-count change for the validated ingestion architecture, the evidence supports 24 workers as the highest-throughput safe point (2.500x vs 8, eff 0.833, all gates PASS), with 12 workers as the most memory-efficient safe point (0.966 eff) for environments where memory saving matters more than wall-clock. A later owner authorization is required before any promotion; the real SAFE_* values, the real checkpoint and the decision gate remain untouched. ``` ## 32. Limitations ```text 1. The parallel engine is the P3-S25.1F lean verifier (no per-line chunk digest); raw/parsed hashes are therefore proven by byte-coverage + boundary identity + the canonical full-digest atomic reference equality (the F/G contract: lean == full-digest was proven on a 72 MiB range in P3-S25.1E and re-asserted here on the 512 MiB canonical atomic reference, all 22 chunks PASS). 2. All worker runs executed warm-cache: the 2.1 GiB fixture was fully resident in the OS page cache after the reference pass; no cold-cache runs were performed this session. Batch-to-batch timing variance (8-worker span 52.76 s in the first batch vs 46.6-48.9 s in later batches) is attributed to cache warmth / system noise; the decision uses the final coherent pass (batch 3). Timing repeatability is evidence of stability, not a correctness argument. 3. N=24 requested 24 but only 22 non-empty worker partitions exist (22 atomic chunks); the 2 empty partition ids (22, 23) are recorded and the gate evaluation uses the actual 22-worker set. This is the legitimate 'fewer processes due to empty tasks' case; it does not weaken any gate. 4. 24 workers saturate the 24 logical CPUs (machine-wide CPU average 0.961, peak 1.000) - a near-saturation condition; further scaling above 24 physical/logical CPUs was not tested (brief: do not expand beyond 24). 5. Parsing is pure-Python and CPU-bound; results are machine/OS dependent (2-socket Xeon Gold 5220, 24 logical, no HT). 6. The synthetic fixture is schema-identical but byte-density may differ slightly from the real 34.5 GB Tickstory source; real-source expansion is forbidden for this experiment. 7. Duplicated worker set mutation for N=24 appends a NEW worker id (22), keeping the worker-identity audit meaningful for the 24->22 process discrepancy case. ``` ## 33. Final classification ```text CLASSIFICATION : A - OPTIMIZATION CANDIDATE IDENTIFIED MAX_SAFE_TESTED = 24 ; PREFERRED_WORKER_COUNT = 24 Every value in this report is machine-derived (harness p3_s251q_harness.py + evidence workload_benchmark/worker_scaling_q/output/*.json). A was not forced: all five candidates passed every gate; 24 was selected because it is the only candidate that breaks the 2-chunk imbalance plateau (2.500x speedup, 0.833 efficiency, imbalance 1.064, CPU 0.961) while remaining far inside machine memory limits. 12 workers remains the back-up preference for memory-constrained operation (0.966 efficiency). ``` *End of P3-S25.1Q final report (generated from machine evidence).*