18 KiB
18 KiB
P3-S25.1F - 512 MiB WORKER-SCALING SAFETY BENCHMARK - FINAL REPORT
DOCUMENT ROLE : REPORT (derived summary; generated from machine
evidence under
ml/p3/p3_s251_external_ingest/workload_benchmark/worker_scaling/output/*.json)
AUTHORITATIVE SOURCE : worker_scaling machine evidence + canonical real
checkpoint output/s251_real_checkpoint.json
PHASE : P3-S25.1F - WORKER-SCALING SAFETY BENCHMARK (512 MiB)
CLASSIFICATION : A - PARALLEL WORKER SCALING VERIFIED
CURRENT FACTUAL STATE : see sections 18-23 (real checkpoint UNCHANGED)
AUTHORIZATION STATE : disposable benchmark ONLY; CHUNK 12 NOT
AUTHORIZED; SAFE_MAX_CHUNK_SIZE and
SAFE_CHUNKS_PER_TURN UNCHANGED
UNRESOLVED ITEMS : see section 25
NEXT OWNER DECISION : see section 26
1. Starting Git SHA
starting_sha = 05a8a3e1522bacc8f66f5ff6c0b7ddbd5044d926
(verified before execution: git rev-parse HEAD == git ls-remote origin HEAD = 05a8a3e1522bacc8f66f5ff6c0b7ddbd5044d926,
branch = main, working tree CLEAN, stash EMPTY, no untracked files; preconditions PASS)
2. Final Git SHA
final_sha = 972f3aafac2622aaf2c3ab3091f5f5ce9b92ba6a
(branch = main, HEAD == origin/main, working tree CLEAN, stash EMPTY verified after push)
3. Benchmark fixture
The ALREADY-VALIDATED disposable 512 MiB workload from P3-S25.1E (identical reference data,
parser, aggregation semantics, atomic chunk semantics, boundary definitions, expected outputs).
fixture path : D:\TradingTerminal\HFM Metatrader 5\MQL5\Shared Projects\SniperGold_ML_authoritative\ml\p3\p3_s251_external_ingest\scratch\p3_s251e_primary_fixture_24MiB.csv
fixture size : 2149580843 bytes
fixture sha256: 2a2691159f1a250b537c82a623539190a59ffd8ea4089ec54245a8f8a7800be5
sha matches P3-S25.1E recorded (2a2691159f1a250b...): True
seed : 20260830
workload : 512 MiB = 536870912 bytes, range [0, 536870912)
atomic chunk : 24 MiB = 25165824 bytes (SAFE_MAX_CHUNK_SIZE, UNCHANGED)
actual line-safe atomic chunks : 22 (each 24 MiB read extends to the next LF)
Real source : XAUUSD_mt5_ticks.csv (34,473,661,010 B) referenced ONLY for the committed
checkpoint; NO real bytes beyond 188,743,943 were read; CHUNK 12 NOT processed.
4. Reference result (canonical single-stream, fresh machine run)
rows_processed = 11422782
m15_count = 120112
m30_count = 108000
first_timestamp = 1483403400
last_timestamp = 1697238205
malformed = {"bid_ask_relationship": 0, "column_count": 2, "invalid_volume": 0, "non_monotonic_timestamp": 0, "non_numeric_price": 0, "non_positive_price": 0, "timestamp_invalid": 1, "timestamp_malformed": 1}
m15_sha256 = 94714104fc23eef83e379a138ef70775c43865419c6dc8afce05b94c9388a95a
m30_sha256 = 238776d863734092c65023817f0ce53acf0d0b2e32510cd63d89806843e1a7e7
output_sha256 = cfa8c084f9bae714a4a15cb1da27e2442bd72bd11aee3f3604304a203e02ed4a
final_state_hash = 757a4d63a0f13c2757281402d5beea4278c982c2e29e36375faa30b9565994ff
matches P3-S25.1E recorded 512 MiB reference : True
reference elapsed_seconds : 342.776
reference peak_memory_bytes : 141534283
output_carry_m15 : {"close": 1209.351, "first_ts": 1697238000, "high": 1209.974, "last_ts": 1697238205, "low": 1208.583, "open": 1208.583, "ticks": 21, "ts": 1697238000}
output_carry_m30 : null
5. worker=1 result
run_ids : workers_1_run_001, workers_1_run_002
n_worker_tasks : 1
all_workers_pass : True
sig_run1 : rows 11422782 | M15 120112 | M30 108000 | first 1483403400 | last 1697238205
output_sha256 : cfa8c084f9bae714a4a15cb1da27e2442bd72bd11aee3f3604304a203e02ed4a
final_state_hash: 757a4d63a0f13c2757281402d5beea4278c982c2e29e36375faa30b9565994ff
chunk_boundaries equal authoritative : True
process_seconds_run1 : 308.901 s
process_seconds_run2 : 313.360 s
semantic_equals_reference_run1 : True
semantic_equals_reference_run2 : True
output_carry_m15 : {"close": 1209.351, "first_ts": 1697238000, "high": 1209.974, "last_ts": 1697238205, "low": 1208.583, "open": 1208.583, "ticks": 21, "ts": 1697238000}
output_carry_m30 : null
6. worker=2 result
run_ids : workers_2_run_001, workers_2_run_002
n_worker_tasks : 2
all_workers_pass : True
sig_run1 : rows 11422782 | M15 120112 | M30 108000 | first 1483403400 | last 1697238205
output_sha256 : cfa8c084f9bae714a4a15cb1da27e2442bd72bd11aee3f3604304a203e02ed4a
final_state_hash: 757a4d63a0f13c2757281402d5beea4278c982c2e29e36375faa30b9565994ff
partition_chunk_groups : [[0, 11], [11, 22]]
worker_boundary_bytes : [[0, 276824537], [276824537, 536870912]]
completion_order_run1 : [1, 0]
process_span_run1 : 161.502 s
process_span_run2 : 163.958 s
semantic_equals_reference_run1 : True
semantic_equals_reference_run2 : True
output_carry_m15 : {"close": 1209.351, "first_ts": 1697238000, "high": 1209.974, "last_ts": 1697238205, "low": 1208.583, "open": 1208.583, "ticks": 21, "ts": 1697238000}
output_carry_m30 : null
8. worker=4 result
run_ids : workers_4_run_001, workers_4_run_002
n_worker_tasks : 4
all_workers_pass : True
sig_run1 : rows 11422782 | M15 120112 | M30 108000 | first 1483403400 | last 1697238205
output_sha256 : cfa8c084f9bae714a4a15cb1da27e2442bd72bd11aee3f3604304a203e02ed4a
final_state_hash: 757a4d63a0f13c2757281402d5beea4278c982c2e29e36375faa30b9565994ff
partition_chunk_groups : [[0, 6], [6, 12], [12, 17], [17, 22]]
worker_boundary_bytes : [[0, 150995197], [150995197, 301990405], [301990405, 427819745], [427819745, 536870912]]
completion_order_run1 : [3, 2, 1, 0]
process_span_run1 : 90.428 s
process_span_run2 : 91.037 s
semantic_equals_reference_run1 : True
semantic_equals_reference_run2 : True
output_carry_m15 : {"close": 1209.351, "first_ts": 1697238000, "high": 1209.974, "last_ts": 1697238205, "low": 1208.583, "open": 1208.583, "ticks": 21, "ts": 1697238000}
output_carry_m30 : null
12. worker=8 result
run_ids : workers_8_run_001, workers_8_run_002
n_worker_tasks : 8
all_workers_pass : True
sig_run1 : rows 11422782 | M15 120112 | M30 108000 | first 1483403400 | last 1697238205
output_sha256 : cfa8c084f9bae714a4a15cb1da27e2442bd72bd11aee3f3604304a203e02ed4a
final_state_hash: 757a4d63a0f13c2757281402d5beea4278c982c2e29e36375faa30b9565994ff
partition_chunk_groups : [[0, 3], [3, 6], [6, 9], [9, 12], [12, 15], [15, 18], [18, 20], [20, 22]]
worker_boundary_bytes : [[0, 75497593], [75497593, 150995197], [150995197, 226492801], [226492801, 301990405], [301990405, 377488009], [377488009, 452985613], [452985613, 503317349], [503317349, 536870912]]
completion_order_run1 : [7, 6, 0, 2, 4, 5, 1, 3]
process_span_run1 : 47.468 s
process_span_run2 : 47.644 s
semantic_equals_reference_run1 : True
semantic_equals_reference_run2 : True
output_carry_m15 : {"close": 1209.351, "first_ts": 1697238000, "high": 1209.974, "last_ts": 1697238205, "low": 1208.583, "open": 1208.583, "ticks": 21, "ts": 1697238000}
output_carry_m30 : null
9. Semantic equivalence (parallel output == single-stream reference)
workers=1 run1==ref True | run2==ref True | PASS True
workers=2 run1==ref True | run2==ref True | PASS True
workers=4 run1==ref True | run2==ref True | PASS True
workers=8 run1==ref True | run2==ref True | PASS True
contract: exact equality on rows / malformed counts / first & last timestamp / M15 count,
M30 count / M15+M30+output hashes / final state (carry) hash - no approximate comparison.
RESULT: PASS (all worker counts)
10. Boundary audit
chunk-level 6-class audit (21 boundaries) all_classes_pass : True
classes: inside_m15, inside_m30, exact_m15, exact_m30, same_ts, carry_boundary -> 6/6 PASS
workers=1 worker-level PASS True | byte-contiguous True | boundary-sequence==authoritative True | carry-chain==canonical True
workers=2 worker-level PASS True | byte-contiguous True | boundary-sequence==authoritative True | carry-chain==canonical True
workers=4 worker-level PASS True | byte-contiguous True | boundary-sequence==authoritative True | carry-chain==canonical True
workers=8 worker-level PASS True | byte-contiguous True | boundary-sequence==authoritative True | carry-chain==canonical True
RESULT: 6/6 PASS for every worker count; no gaps, no overlaps; actual_byte_end(worker_n)
== actual_byte_start(worker_n+1) at every worker boundary.
11. Interruption / resume
(a) CANONICAL corrected actual-byte-end resume (P3-S25.1D-R semantics; tail re-executed)
at the first worker-boundary chunk index of each worker count:
workers=2 stop_after_chunk=11 match=True per_chunk=True bounds=True aggregate=True
workers=4 stop_after_chunk=6 match=True per_chunk=True bounds=True aggregate=True
workers=8 stop_after_chunk=3 match=True per_chunk=True bounds=True aggregate=True
all_match : True
(b) PARALLEL-LEVEL resume at EVERY worker boundary (prefix merged carry == canonical chunk
carry at that boundary; resume byte start == actual line-safe byte end; shuffled
completion-order merge == continuous):
workers=2 PASS True entries=1
workers=4 PASS True entries=3
workers=8 PASS True entries=7
parallel_all_match : True
RESULT: PASS (continuous == interrupted + resumed for semantic output, final state,
boundary sequence and carry chain; worker completion order does not influence resume)
12. Mutation detection (10 classes)
classes: skipped atomic chunk; duplicated atomic chunk; wrong byte offset; wrong worker
ordering; wrong input carry; post-chunk carry as input; output-row deletion; output-row
duplication; semantic output mutation; boundary-sequence mutation.
workers=1 10/10 detected (chunk-level (canonical single-stream))
workers=2 10/10 detected (worker-level)
workers=4 10/10 detected (worker-level)
workers=8 10/10 detected (worker-level)
all_detected : True
RESULT: 10/10 DETECTED for every worker count selected for safety evaluation.
A reduction in mutation-detection capability would classify the worker configuration
as NOT SAFE; none occurred.
13. Reproducibility (two complete runs per worker count)
semantic equality run1 == run2 required; timing equality NOT required (nor treated as
correctness evidence by itself).
workers=1 run1==run2 True | output_sha run1 cfa8c084f9bae714 run2 cfa8c084f9bae714
workers=2 run1==run2 True | output_sha run1 cfa8c084f9bae714 run2 cfa8c084f9bae714
workers=4 run1==run2 True | output_sha run1 cfa8c084f9bae714 run2 cfa8c084f9bae714
workers=8 run1==run2 True | output_sha run1 cfa8c084f9bae714 run2 cfa8c084f9bae714
all_pass : True
RESULT: PASS
14. Resource profile (worker=1..8, run1 processing spans)
workers=1 process_span_run1=308.901 s run2=313.360 s | rows/s=36978.8 | peak=229.9 MiB
sum_worker_peak=229.9 MiB | cpu_frac=None | imbalance(max/mean)=1.000 | spawn_overhead=0.00 s | merge=0.00 s
workers=2 process_span_run1=161.502 s run2=163.958 s | rows/s=70728.6 | peak=229.9 MiB
sum_worker_peak=459.8 MiB | cpu_frac=0.971 | imbalance(max/mean)=1.025 | spawn_overhead=0.65 s | merge=3.26 s
workers=4 process_span_run1=90.428 s run2=91.037 s | rows/s=126318.9 | peak=229.9 MiB
sum_worker_peak=919.5 MiB | cpu_frac=0.89 | imbalance(max/mean)=1.119 | spawn_overhead=0.52 s | merge=3.35 s
workers=8 process_span_run1=47.468 s run2=47.644 s | rows/s=240642.4 | peak=229.9 MiB
sum_worker_peak=1839.0 MiB | cpu_frac=0.878 | imbalance(max/mean)=1.134 | spawn_overhead=0.54 s | merge=3.46 s
metrics legend: process_span = wall-clock of the parallel processing phase (max worker wall);
rows/s = rows / process_span; peak = max per-worker tracemalloc peak (bounded by the
24 MiB atomic chunk, flat across worker counts); cpu_frac = total CPU seconds / (span x N);
imbalance = max worker wall / mean worker wall; I/O utilization not directly observable
(fixture re-read from OS page cache after the boundary scan; CPU dominates).
15. Speedup
speedup(1 workers) = time(1)/time(1) = 1.000
speedup(2 workers) = time(1)/time(2) = 1.913
speedup(4 workers) = time(1)/time(4) = 3.416
speedup(8 workers) = time(1)/time(8) = 6.508
(process-span basis, machine-derived)
16. Efficiency
efficiency(1 workers) = speedup(1)/1 = 1.000
efficiency(2 workers) = speedup(2)/2 = 0.956
efficiency(4 workers) = speedup(4)/4 = 0.854
efficiency(8 workers) = speedup(8)/8 = 0.813
17. Smallest SAFE_USEFUL worker count
minimum_safe_useful_worker_count = 2
Classified per count: 1 = SAFE_USEFUL_BASELINE; 2/4/8 = SAFE_USEFUL (all gates PASS,
speedup >= 1.3, efficiency >= 0.4). The smallest count giving a meaningful wall-clock
reduction is 2 (1.913x, efficiency 0.956). The largest count is NOT auto-selected.
18. Maximum SAFE worker count tested
maximum_safe_worker_count_tested = 8
(8 = the largest required candidate; all gates PASS; no more than 8 was tested per brief)
19. Checkpoint before / after (real, read-only)
real_checkpoint_path : D:\TradingTerminal\HFM Metatrader 5\MQL5\Shared Projects\SniperGold_ML_authoritative\ml\p3\p3_s251_external_ingest\output\s251_real_checkpoint.json
checkpoint_sha_before : 76ce95eb1e67b4138302a9c9d318e3a3c9f3e52430ac80c08f1a9900da23d0c1
checkpoint_sha_after : 76ce95eb1e67b4138302a9c9d318e3a3c9f3e52430ac80c08f1a9900da23d0c1
byte_identical : True
run_id : P3_S251_REAL_TICK_RUN_001 (untouched)
last_completed_chunk : 11 | next_chunk 12 | next_byte_offset 188743943
(values verified from the canonical checkpoint; the real checkpoint MUST NOT change - it did not)
20. SAFE_MAX_CHUNK_SIZE
SAFE_MAX_CHUNK_SIZE = 25165824 bytes (24 MiB) - UNCHANGED (P3-S25.1C).
This benchmark performed no dynamic-chunk change and promotes no atomic size.
21. SAFE_CHUNKS_PER_TURN
SAFE_CHUNKS_PER_TURN = NOT ESTABLISHED (UNCHANGED as operational policy).
This benchmark investigated WORKER PARALLELISM (1/2/4/8 workers over 22 atomic chunks),
not chunks-per-turn batching; it changes no operational setting.
22. Operational-policy status
OPERATIONAL_POLICY : NOT CHANGED
Real Chunk 12 execution : NOT AUTHORIZED and NOT PROCESSED
New real Tickstory bytes : NONE read (real reads never passed 188,743,943; the benchmark
read only the disposable synthetic fixture)
Dukascopy : NOT accessed | ML : NOT run | Production MQL5/MQH : UNCHANGED
Frozen contracts / feature / label contracts : UNCHANGED
CURRENT_DECISION_GATE.md / CURRENT_PROJECT_STATE.md : UNCHANGED
Historical scientific evidence : UNCHANGED
23. CHUNK 12 status
CHUNK 12 = NOT AUTHORIZED (CURRENT_DECISION_GATE unchanged).
next_byte_offset 188743943 is FACT, not authorization. This benchmark produced
performance evidence only; it grants no ingestion authority.
24. Failed tests and preserved failure evidence
No worker-count candidate failed any gate in the final evidence set: 1/2/4/8 all PASS
semantic equivalence, boundary 6/6, interruption/resume, 10/10 mutations, reproducibility
and resource practicality (classification A).
Process note (transparent, no erased history): an early run of the benchmark harness
recorded C - PARALLEL UNSAFE because of a harness METADATA bug (N_ATOMIC = 21 derived by
integer division 536870912//25165824, while the actual line-safe chunk count is 22). The
worker pipeline itself was correct throughout; the metadata constant and the affected gate
bookkeeping were corrected and the downstream worker runs were re-executed in full. No
correctness failure of any worker candidate occurred; therefore no failure artifact needed
preservation under the separate run-identifier rule. The corrected evidence superseded the
metadata-corrected runs; no failed run was overwritten to fabricate a PASS.
25. Limitations
* Fixture: all worker runs use the deterministic SCHEMA-IDENTICAL SYNTHETIC disposable
fixture (the real committed prefix is ~156 MiB and real-source expansion is forbidden).
The exercised mechanics (parser, aggregator, line-safe chaining, carry, boundary scan,
coordinator merge) are the canonical code, but byte-density / malformed profile of real
Tickstory data may differ slightly.
* Parsing is pure-Python; parallelism uses process workers (the GIL would serialize
CPU-bound threads). Results are therefore machine/OS dependent.
* Per-worker peak memory (~230 MiB) is bounded by the 24 MiB atomic chunk; total memory
scales ~linearly with worker count (8 workers ~= 1.8 GiB), still far below the
machine's available RAM (38.2 GiB available at run start).
* I/O utilization is not directly observable; the fixture is re-read per worker from OS
page cache after the boundary scan, so CPU dominates.
* The order-dependent malformed class non_monotonic_timestamp is reconciled across worker
boundaries by the coordinator (first-ts vs last-ts check); the fixture is strictly
monotonic, so the count is 0 and the check is exact but not stress-tested.
* The lean verifier (no per-line chunk digest) is inherited from P3-S25.1E; its equality
with the canonical full-digest verifier was proven there on a 72 MiB range.
* The benchmark measured PROCESS, VERIFY, RESUME, MERGE and EVIDENCE phases; total
turn-cycle time is the sum of the machine-recorded phase durations.
26. Recommended worker configuration
RECOMMENDED : 2 workers (smallest SAFE_USEFUL count)
speedup 1.913x / efficiency 0.956; ~48% less processing wall-clock vs 1 worker;
memory ~2x the single-worker budget (2 x 230 MiB); minimal failure surface;
spawn overhead 0.65 s; merge 3.26 s.
Verified-safe alternatives: 4 workers (3.416x / 0.854) and 8 workers (6.508x / 0.813)
also pass every gate; choose them only if wall-clock dominates and memory/task overhead is
acceptable (8 workers ~= 1.8 GiB total, ~0.54 s spawn, ~3.46 s merge).
DO NOT exceed 8 without a new benchmark (brief: test 1/2/4/8 only; scaling is under
investigation, not maximum thread count). This recommendation is EVIDENCE ONLY - it does
not authorize any change to SAFE_MAX_CHUNK_SIZE / SAFE_CHUNKS_PER_TURN / real ingestion.
27. Classification
PHASE CLASSIFICATION: A - PARALLEL WORKER SCALING VERIFIED
A - PARALLEL WORKER SCALING VERIFIED: at least one worker count >1 is SAFE_USEFUL and
provides meaningful performance improvement without loss of correctness (2/4/8 all
SAFE_USEFUL; smallest = 2). The question of the phase is answered: YES - worker
parallelism makes the already-verified 512 MiB workload substantially faster while
preserving the exact result, resumability and verification capability.