Skip to content

Latest commit

 

History

History
123 lines (96 loc) · 7.44 KB

File metadata and controls

123 lines (96 loc) · 7.44 KB

Cycle 2 Closeout — btc-longhorizon-predictor

Verdict: KILL-for-now. Held-out SEALED (never opened). D9 exit TRIGGERED — the quant-research line is shelved as a solved question.

Registered at cycle2-registration-v1 (2026-07-14), executed and closed 2026-07-15. Zero frozen literals changed across the build; every mid-build mechanic is logged in docs/registration_log.md (B1–B16). This memo is the settled reading; the numbers live in reports/validation_folds.md and reports/validation_results.csv.


1. What the decisive table says

240 configurations (3 tree families × the frozen hyperparameter grid × 4 horizons, 6 purged folds each), calibrated out-of-sample, priced under both a taker and a modeled-maker cost model. 0/240 pass all validation gates (G1 & G2 & G3 & G4 & taker-G5) with ≥200 out-of-sample trades → the registered insufficiency rule → KILL-for-now, held-out sealed. The held-out block was never opened because no config qualified to open it.

gate plain question pass / 240
G5-taker makes money after costs? 40
G1 beats "always guess majority" by ≥3pp? 0
G3 are its probabilities trustworthy OOS? 0
G4 positive in ≥5 of 6 time folds? 0 (best = 4)
ALL 0

2. The mechanism, named: cost hypothesis CONFIRMED, signal hypothesis REFUTED — same table

This is the reading that matters, and it is precise.

Cycle 1 could not distinguish "no signal" from "signal hiding below the cost floor." Everything died at G5: 0/600 configs cleared 11 bps taker costs, so the flat null was ambiguous — was the market unpredictable, or was the edge real but smaller than retail fees?

Cycle 2 was built to resolve exactly that ambiguity, and it did. The maker model was built, the cost floor was roughly halved as designed, and 40 configs stepped over the lower bar. So the cost hypothesis is confirmed: costs really were the thing killing G5, and lowering them lets configs pass it.

But what stepped over the bar turned out to be drift-harvesting, not prediction. The same table that confirmed the cost hypothesis refuted the signal hypothesis, on four independent axes:

  • 0/240 at G1 — no config has directional skill over the majority-class baseline.
  • 0/240 at G3 — every config's probabilities are uninformative out-of-sample.
  • max 4/6 positive folds — the pooled "profit" is regime-dependent, never persistent across ≥5 folds; it comes from a couple of lucky stretches.
  • 35 of the 40 G5-passers fail to beat costed buy-and-hold (only 5 have a drift-difference CI above zero, B14/D8).

That last number is the cleanest sentence in the memo: what cleared the lower cost bar was mostly the long drift the D8 baseline was registered to expose — being long in a rising 2020–2025 market, not predicting it. The residual pooled expectancy that survives at 8–24h is drift, not skill.

The two cycles jointly answer the full question. Not only "is there a directional edge at retail-accessible costs?" (no), but also "were costs merely hiding an edge?" (also no — once the floor was halved, what emerged was drift, not prediction). A two-cycle proof that closes both the question and its most obvious escape hatch.

3. Escape hatch — D9 reopening conditions, now PARTIALLY CONSUMED

D9's registered reopening conditions were (a) materially lower costs, (b) — [reserved], and (c) a categorically different input class. Condition (a) is now spent. This cycle is the lower-cost test: the maker model was built, the floor halved, and the verdict survived it. Future-you may not re-cite "lower costs" as fresh grounds to reopen — that lever has been pulled and the null held.

What remains live is only the categorically-different-input list, and only that:

  • measured fills from real paper-trading / execution infrastructure (replacing the φ=0.7 assumption with a measurement) — note this can only ever license a paper/measurement spec, never live capital, per D9;
  • trade-level / order-book data (a genuinely different information source than OHLCV klines);
  • a different market (different asset or venue with different efficiency).

Absent one of those, the question is closed. "Try more models," "tune the grid," "another horizon," "lower costs again" are all explicitly out of scope by registration.

4. Honesty ledger — a pipeline that never inflated

Cycle 1's ledger recorded inflation removed — it caught its own fold-1 in-sample-calibration leak (+1.80e-3 in-sample vs −5.36e-4 honest) and corrected it. Cycle 2's ledger is different in kind: it records a pipeline that never inflated in the first place.

  • Determinism resolved honestly, against convenience. The GPU was available (RTX 4060) and the standing instruction preferred it, but the R2.2 gate found XGBoost GPU≠CPU (bit-diff 0.295) and LightGBM had no OpenCL device — so both ran CPU, because a result a CPU re-run cannot reproduce is not admissible evidence. The full 240-config grid then reproduced bit-identically on an independent re-run. Determinism was claimed and then demonstrated.
  • The leak lesson was promoted to design. Cycle 1's fold-1 in-sample fallback (the leak) was replaced by the D6 burn-in scheme; fold-1's burn-in trades were excluded from every decisive statistic (B3), enforced in code and tests. The 2024-10-28 exchange-outage bar was NaN-guarded and disclosed, not silently absorbed.
  • Named risks were scored as registered, including the one that didn't fire. B12 pre-named an h=48 trade-count handicap; it did not bind (every h=48 config traded ≥205). That non-binding was recorded as a non-cause, not quietly retrofitted into the narrative. B13's location bet (h=16–24) was scored as weakly-confirmed-but-moot. B8's fold-1 triple handicap was pre-named so weak fold-1 numbers were never "fixed" mid-run.
  • The verdict came from a pre-committed rule with zero mid-cycle judgment calls. No gate was softened after seeing results; no floor was lowered when h=48 looked thin; the insufficiency rule fired mechanically.

The reusable asset is the methodology, now proven twice: once by catching its own leak (Cycle 1), once by running clean end-to-end (Cycle 2). That is the transferable output — a validation discipline that answers honestly whether or not the answer is the one you wanted.

5. The remarkable non-event

After two full cycles and 840 configurations, the held-out block was never opened. The final exam was never needed, because the honest validation process kept answering first — every time, the pooled out-of-sample gates delivered a verdict before any config earned the right to touch the sealed data. A held-out set that stays sealed through two complete research cycles is not a failure to use it; it is the process working as designed.

6. The shelf

Per D9, and per the decision made while neutral (before any results existed): the line goes on the shelf. The settled question is recorded in project memory in its full citable form. The energy goes where the base rates are positive — the ventures with distribution and a non-adversarial market. The discipline that made these two kills trustworthy is fully transferable; its next application should be somewhere the market is not an efficient adversary priced to erase exactly the edge you seek.

Closed 2026-07-15. Tag: modeling-cycle-2-closed. Held-out sealed. Reopen only under D9 condition (c), never (a).