Reference predictors priced under BOTH execution models on the exact decisive OOS universe: pooled walk-forward validation blocks, fold-1 calibration burn-in excluded (B3), strict-drop seam (B7). Every net-expectancy number below is directly comparable to the ML G5 stream to come. taker is decisive; maker is φ-weighted (φ = 0.7) reported-only. Held-out sealed.
The number every ML config must beat. Net = r_h − round_trip − funding(long pays). taker CI-lower floor = -0.00055.
| h | model | n_trades | mean_net | CI95 low | CI95 high | > taker floor? |
|---|---|---|---|---|---|---|
| 8 | taker | 5,030 | -0.00085 | -0.00137 | -0.00033 | no |
| 8 | maker_entry | 5,030 | -0.00034 | -0.00070 | +0.00003 | — |
| 16 | taker | 2,515 | -0.00060 | -0.00159 | +0.00046 | no |
| 16 | maker_entry | 2,515 | -0.00016 | -0.00086 | +0.00058 | — |
| 24 | taker | 1,676 | -0.00036 | -0.00192 | +0.00112 | no |
| 24 | maker_entry | 1,676 | +0.00001 | -0.00108 | +0.00105 | — |
| 48 | taker | 838 | +0.00039 | -0.00303 | +0.00385 | no |
| 48 | maker_entry | 838 | +0.00053 | -0.00186 | +0.00295 | — |
momentum = last-bar sign; markov = 4-state train-quantile. Not selected, not tuned — a floor of cheap directional skill. verdict per taker G5 contract.
| baseline | h | model | n_trades | mean_net | CI95 low | verdict |
|---|---|---|---|---|---|---|
| momentum | 8 | taker | 5,029 | -0.00109 | -0.00159 | fail |
| momentum | 8 | maker_entry | 5,029 | -0.00051 | -0.00086 | reported_only |
| markov | 8 | taker | 0 | n/a | n/a | insufficient_evidence |
| markov | 8 | maker_entry | 0 | n/a | n/a | reported_only |
| momentum | 16 | taker | 2,514 | -0.00121 | -0.00223 | fail |
| momentum | 16 | maker_entry | 2,514 | -0.00059 | -0.00130 | reported_only |
| markov | 16 | taker | 0 | n/a | n/a | insufficient_evidence |
| markov | 16 | maker_entry | 0 | n/a | n/a | reported_only |
| momentum | 24 | taker | 1,676 | -0.00120 | -0.00275 | fail |
| momentum | 24 | maker_entry | 1,676 | -0.00058 | -0.00166 | reported_only |
| markov | 24 | taker | 521 | -0.00084 | -0.00437 | fail |
| markov | 24 | maker_entry | 521 | -0.00033 | -0.00280 | reported_only |
| momentum | 48 | taker | 838 | -0.00100 | -0.00397 | fail |
| momentum | 48 | maker_entry | 838 | -0.00044 | -0.00252 | reported_only |
| markov | 48 | taker | 405 | -0.00150 | -0.00731 | fail |
| markov | 48 | maker_entry | 405 | -0.00079 | -0.00486 | reported_only |
Amendment 2: accuracy/F1/AUC never decide; G5 does. Shown to locate cheap directional skill, not to license a trade.
| baseline | h | accuracy | edge vs majority (pp) | n_val |
|---|---|---|---|---|
| majority | 8 | 0.4085 | -0.58 | 40,239 |
| momentum | 8 | 0.3836 | -3.07 | 40,239 |
| markov | 8 | 0.4085 | -0.58 | 40,239 |
| majority | 16 | 0.4419 | -0.70 | 40,231 |
| momentum | 16 | 0.4233 | -2.56 | 40,231 |
| markov | 16 | 0.4419 | -0.70 | 40,231 |
| majority | 24 | 0.4573 | -0.68 | 40,223 |
| momentum | 24 | 0.4338 | -3.03 | 40,223 |
| markov | 24 | 0.4573 | -0.68 | 40,223 |
| majority | 48 | 0.4804 | -1.03 | 40,199 |
| momentum | 48 | 0.4521 | -3.85 | 40,199 |
| markov | 48 | 0.4804 | -1.03 | 40,199 |
- Buy-and-hold under taker costs is net-positive at some h on this OOS universe — this is how much of the +1>−1 label skew is just 2020–25 drift, and it is the bar the ML must exceed, not merely match.
- The maker column shows the same streams φ-weighted; it is context for the reopen-licensing lever, never a verdict.
- Directional accuracy (§3) is diagnostic; a baseline can look 'skilled' on accuracy and still be net-negative once costs and the non-overlap subsample are applied — which is the entire point of pricing it here first.