Skip to content

Latest commit

 

History

History
66 lines (54 loc) · 3.93 KB

File metadata and controls

66 lines (54 loc) · 3.93 KB

Cycle-2 Baselines Report — the honest bar

Reference predictors priced under BOTH execution models on the exact decisive OOS universe: pooled walk-forward validation blocks, fold-1 calibration burn-in excluded (B3), strict-drop seam (B7). Every net-expectancy number below is directly comparable to the ML G5 stream to come. taker is decisive; maker is φ-weighted (φ = 0.7) reported-only. Held-out sealed.

1. Honest bar — costed buy-and-hold (long-only) net expectancy

The number every ML config must beat. Net = r_h − round_trip − funding(long pays). taker CI-lower floor = -0.00055.

h model n_trades mean_net CI95 low CI95 high > taker floor?
8 taker 5,030 -0.00085 -0.00137 -0.00033 no
8 maker_entry 5,030 -0.00034 -0.00070 +0.00003
16 taker 2,515 -0.00060 -0.00159 +0.00046 no
16 maker_entry 2,515 -0.00016 -0.00086 +0.00058
24 taker 1,676 -0.00036 -0.00192 +0.00112 no
24 maker_entry 1,676 +0.00001 -0.00108 +0.00105
48 taker 838 +0.00039 -0.00303 +0.00385 no
48 maker_entry 838 +0.00053 -0.00186 +0.00295

2. Tradeable baselines — net expectancy at reference p*=0.50 (pooled OOS)

momentum = last-bar sign; markov = 4-state train-quantile. Not selected, not tuned — a floor of cheap directional skill. verdict per taker G5 contract.

baseline h model n_trades mean_net CI95 low verdict
momentum 8 taker 5,029 -0.00109 -0.00159 fail
momentum 8 maker_entry 5,029 -0.00051 -0.00086 reported_only
markov 8 taker 0 n/a n/a insufficient_evidence
markov 8 maker_entry 0 n/a n/a reported_only
momentum 16 taker 2,514 -0.00121 -0.00223 fail
momentum 16 maker_entry 2,514 -0.00059 -0.00130 reported_only
markov 16 taker 0 n/a n/a insufficient_evidence
markov 16 maker_entry 0 n/a n/a reported_only
momentum 24 taker 1,676 -0.00120 -0.00275 fail
momentum 24 maker_entry 1,676 -0.00058 -0.00166 reported_only
markov 24 taker 521 -0.00084 -0.00437 fail
markov 24 maker_entry 521 -0.00033 -0.00280 reported_only
momentum 48 taker 838 -0.00100 -0.00397 fail
momentum 48 maker_entry 838 -0.00044 -0.00252 reported_only
markov 48 taker 405 -0.00150 -0.00731 fail
markov 48 maker_entry 405 -0.00079 -0.00486 reported_only

3. Directional diagnostics (accuracy / edge over majority) — DIAGNOSTIC ONLY

Amendment 2: accuracy/F1/AUC never decide; G5 does. Shown to locate cheap directional skill, not to license a trade.

baseline h accuracy edge vs majority (pp) n_val
majority 8 0.4085 -0.58 40,239
momentum 8 0.3836 -3.07 40,239
markov 8 0.4085 -0.58 40,239
majority 16 0.4419 -0.70 40,231
momentum 16 0.4233 -2.56 40,231
markov 16 0.4419 -0.70 40,231
majority 24 0.4573 -0.68 40,223
momentum 24 0.4338 -3.03 40,223
markov 24 0.4573 -0.68 40,223
majority 48 0.4804 -1.03 40,199
momentum 48 0.4521 -3.85 40,199
markov 48 0.4804 -1.03 40,199

4. Reading the bar (factual)

  • Buy-and-hold under taker costs is net-positive at some h on this OOS universe — this is how much of the +1>−1 label skew is just 2020–25 drift, and it is the bar the ML must exceed, not merely match.
  • The maker column shows the same streams φ-weighted; it is context for the reopen-licensing lever, never a verdict.
  • Directional accuracy (§3) is diagnostic; a baseline can look 'skilled' on accuracy and still be net-negative once costs and the non-overlap subsample are applied — which is the entire point of pricing it here first.