You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Ledger 127: ownership stickiness at rebuild does not reduce migration on the rung deck and is not byte-identical (few-ULP owner dependence in the stage-1 advance); negative result, not landed
Copy file name to clipboardExpand all lines: docs/documentation/amr_action_plan.md
+10Lines changed: 10 additions & 0 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -226,6 +226,16 @@ possible while AMR aborts on the target machine at 1 rank, and every increment b
226
226
on a compiler that does not reproduce it. It also means the ladder should add a CCE arm as soon as one
227
227
exists, or the same class of breakage will keep accumulating undetected.
228
228
229
+
## 2026-09-09 (127) — OWNERSHIP STICKINESS DOES NOT PAY ON THE RUNG DECK AND IS NOT BYTE-IDENTICAL (GOAL v7 item 3b, step 1, NEGATIVE): a regrid box identical to a previous-generation block keeping that block's owner (inc/sticky 87400436, accepted when the total weight imbalance stays within 10 % of the plain cut's) fired at one of the four rebuilds of the np8 rung deck (320 of 576 boxes kept at step 220; 0 at the two mesh-growth rebuilds, steps 20 and 40, where nothing pre-exists at 20 and the set grows 128 -> 576 at 40 with no box identical; 0 at step 240, where the envelope shifted and no box was identical) and left the migration where it was, in fact marginally worse (975 vs 967 blocks moved and 72.5 vs 69.9 GB over the four rebuilds -- the counters are cumulative and exact; at the rebuild where it fired, 419 vs 416 blocks and 26.6 vs 25.2 GB: keeping 320 of 576 boxes on their owner did not remove a single net block move; rg:mig 32.0 vs 32.3 s per 240 steps, read on the session node k004-002 with builds running and against a base binary five commits back, so the seconds are indicative only), because the re-cut of the remaining 256 boxes moved at least as many blocks as the plain cut of all 576 did (no instrument attributes moves to kept vs re-cut boxes; this is inference from the totals); and it broke byte identity at the few-ULP level (density in 21770 scattered cells at step 41 of the S0 deck, max relative 4.4e-16, ~2x double eps, amplified in the near-zero momenta; base grid identical at step 40, first divergence in the stage-1 advance, not the fills; persists with the batched advance off), which the pre-registration had defined as a bug to localize rather than a tolerance to accept -- not landed
230
+
231
+
**Pre-registration (notes/ledger_drafts/l127_prereg.md).** (1) kept >= 60 % on the late rebuilds, 0 on the seeds: NOT MET on either late rebuild (55.6 % at step 220, 0 % at step 240); the seed clause met (0 at both). (2) blocks_moved 967 -> < 400, bytes 70 -> < 30 GB, rg:mig 129 -> < 70 ms/step: NOT MET (975 blocks, 72.5 GB cumulative; rg:mig 133 vs 134 ms/step on the k004-002 pair, contaminated but equal). (3) accepted within 1.10 on every rebuild: met (imb_sticky 1.090 vs the plain cut's 1.011, ratio 1.078; accepted T at all four) -- but the absolute balance did degrade 1.011 -> 1.090 at that rebuild, i.e. the lever bought no migration and cost 8 % of balance. (4) np8 step -1 to -3 %: NOT READ -- the pin's only pair ran on the session node k004-002 while builds ran on it (discarded, as every k004-002 pair is in ledger 130: sticky 820.2 s vs base 879.0 s per 240 steps exist in logs/hold-np8pair-{sticky,base}.log and are not an A/B for stickiness), and the base binary 96966782 is five commits back (the coarse-halo hoist and the fold merge are in the sticky pin: halo 1380 vs 720 calls, b:halo 3.9 vs 49.2 s); the decision rests on stickiness's own rows (kept, blocks moved, bytes), which are exact counts and did not move -- the first two rebuilds are byte-equal across the arms (109 blocks, 16842660480 B in both), the evidence that the box sets agree. [amr-cad] escaped 0 on both arms and both stops. Falsifier ("kept < 30 % on the late rebuilds -> the migration is the envelope's motion"): fired on the second late rebuild.
232
+
233
+
**Identity.** ident2 87400436 vs 96966782 on the S0 deck np8: DIFFER on both restart files. Bisection (notes/item3b_affinity_stickiness_design.md, 2026-09-09 09:00 section): step 40 base grid IDENTICAL, the AMR file differs; step 41 both differ; density differs in 0.3 % of the refined region's cells at the few-ULP level (max rel 4.4e-16, ~2x double eps), momenta near zero amplify it; persists with amr_batched_advance = F; per-phase XOR fingerprints (probe/ckxor) put the first divergence in stage 1's advance of step 40, after identical regrid output and identical fill interiors (the fingerprint excludes ghost shells; the ghost-inclusive probe did not run). Ownership therefore enters the arithmetic at the round-off level in the advance path or in the ghost data it reads; the reflux apply is excluded (its fingerprint at step 40 stage 1 is identical, and the divergence precedes it), and the surviving suspects are the ghost-shell contents set by the fills and an owner-dependent path inside the advance itself, e.g. a kernel pair with different FMA contraction on the co-located vs cross-rank path. Not localized; the goldens subset (6/6) and CPU AMR goldens (71/71) pass, i.e. the difference is within golden tolerance.
234
+
235
+
**What was learned.** On this deck the snap already turns 8 of the 12 regrids into no-ops and the remaining rebuilds are envelope shifts where few boxes are identical; by the equality of the counts, the migration cost is the envelope's motion, not gratuitous re-cutting (an inference from totals, not an attribution per box). The few-ULP owner dependence is a latent property of the code (any ownership change, including the Morton/Cartesian alignment noted in notes/item4_interior_first_design.md, will show it) and needs its own localization before any ownership lever can be gated by identity.
236
+
237
+
**Deferred, stated per GOAL v7 item 6:** localizing the owner dependence (per-block fingerprints inside the stage-1 advance); parent affinity (i3) and the Morton/Cartesian alignment, both blocked on the same identity question; the branch inc/sticky stays in the gate tree.
238
+
229
239
## 2026-09-09 (119) — THE NP16 REBUILD-FREE WINDOW PROBE (GOAL v6 item 1): the per-interval print splits the two-node wait into a PERSISTENT part present in rebuild-free windows and ONE OBSERVED post-rebuild spike -- at np16 (two nodes, one of them GPU-faulting) the interval after the late rebuild at step 220 carries 22-30 % of every wait row of the run (b:halo 30 %, reflux 28 %, restr 25 %, halo 23 %, seam 22 %) at 2.5-4.1x the mean of the seven rebuild-free intervals, with an rhs max/mean of 1.85, while those seven intervals still run at rhs max/mean 1.12-1.24 and a 2.9x spread in b:halo between them, and the interval after the seed rebuild at step 40 is not elevated in its waits; at np8 on a healthy node the same post-rebuild interval carries 16.5-20.5 % of each row at 1.8-2.8x the rebuild-free mean with NO rhs skew (1.09; the flat intervals 1.04-1.08) -- so the aftermath is real at both scales but is an exchange-side wait at np8 and, at np16, either rank-work skew or the faulting node (n = 1, not separable here), and the larger, persistent part of the np16 wait exists with no rebuild in sight; GOAL v6 item 3 (the level-2 detail-preserving overlap copy) is DEFERRED, a stated departure from the pre-registered decision rule
230
240
231
241
**Pre-registration (notes/ledger_drafts/l119_prereg.md, saved 2026-09-08 22:06 before the instrument was built).** (1) at np16 the intervals right after a rebuild carry wait rows 1.5-3x the flat intervals', flat intervals within 20 % of each other: PARTLY -- the one observed post-rebuild interval is 2.5-4.1x (above the range), the post-seed interval is not elevated, and the flat intervals are NOT within 20 % (b:halo 7.6-22.4 s, 2.9x). (2) np8 shows a smaller jump (< 1.5x): NOT MET as stated -- on the healthy node the post-rebuild interval is 1.8-2.8x the rebuild-free mean; the relative statement (np8 below np16's 2.5-4.1x) holds. (3) rhs max/mean > 1.1 in rebuild intervals and < 1.07 in flat ones: HALF at np16 (1.85 after the rebuild; 1.12-1.24 in the seven flat intervals, 1.30 in the post-seed interval), NOT MET at np8 (1.09 after the rebuild, 1.04-1.08 flat: no rhs skew either side). (4) migration bytes nonzero only in rebuild intervals: met (44.1 GB at the seed, 62.9 and 67.0 GB at the two late rebuilds, zero elsewhere). The prereg expected three late rebuilds; two occurred (steps 220 and 240) and only the first's aftermath is inside the run, so every aftermath statement here is n = 1.
0 commit comments