2 KiB
VALIDATION_PROTOCOL.md
1. Primary method: chronological walk-forward
training window
-> forecast future observation/window
-> observe actual outcome at origin+H
-> advance origin
-> repeat
supported windows: expanding, rolling, with configurable train length and horizon.
All data used in a forecast is strictly before the forecast origin. The outcome is revealed only after the prediction is frozen (origin + H).
2. Baselines
| Variant | Description |
|---|---|
| A | Naive (random-walk / last value), normalized zero-return where appropriate |
| B | ARIMA only |
| C | SAX only |
| D | ARIMA + SAX |
| E | Agreement-only filter |
| F | Disagreement/uncertainty rejection |
No variant may claim superiority unless it beats the same protocol baselines.
3. Metrics
- Forecast: MAE, RMSE, MASE (scale-aware), directional accuracy, calibration of directional probability.
- Distribution/Evidence: analog count dist, interval widths, SAX p25-p75, ARIMA PI width, agreement/disagreement/no-edge/insufficient rates.
- Economic (only after forecast metrics): gross/net expected return, expectancy, profit factor, max drawdown, Sharpe/Sortino where appropriate, trade frequency, turnover, cost sensitivity. Cost assumptions explicit.
4. Ablation matrix and experiments
E1: ARIMA > naive? E2: SAX > naive? E3: SAX + info beyond ARIMA? E4: ARIMA + info beyond SAX? E5: agreement -> improved conditional performance? E6: disagreement -> poor/unstable regime? E7: survives costs? E8: stability across walk-forward periods?
Do not proceed to optimization before E1–E8 are answered.
5. Reproducibility
Every report records: commit, source version, config, symbol, timeframe, data range,
data source, horizon H, model parameters, seeds, cost assumptions, validation
protocol, artifacts.
6. No-lookahead enforcement
A single integration test drives the whole pipeline; it fails hard on any forward-contamination. A failing no-lookahead test is a HARD STOP.