You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Ledger 132: scorecard item 2 re-baselined on the fixed pin, one node, one window: MFC AMR excess 0.58 s/step vs AMReX 0.45 (1.30x; ledger 117: 1.94x); half wait, a third thin work, 0.09 residual per-block RHS inflation; the batching's fragmentation is the pre-registered lever (pad tolerance falsified, largest-first order identity-gated, wall read pending on the EOS pin)
Copy file name to clipboardExpand all lines: docs/documentation/amr_action_plan.md
+23Lines changed: 23 additions & 0 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -226,6 +226,29 @@ possible while AMR aborts on the target machine at 1 rank, and every increment b
226
226
on a compiler that does not reproduce it. It also means the ladder should add a CCE arm as soon as one
227
227
exists, or the same class of breakage will keep accumulating undetected.
228
228
229
+
## 2026-09-09 (132) — SCORECARD ITEM 2 RE-BASELINED ON THE FIXED PIN, ONE NODE, ONE WINDOW (GOAL v7 item 7a, the user-authorized single-node pivot): MFC's steady AMR excess is 0.58 s/step (3 reps, sd 0.04) against AMReX's 0.45 (sd 0.02) on k004-002 between 18:43 and 19:25 -- 1.30x, against ledger 117's 1.94x on k004-004 (0.70 vs 0.36) -- and the excess decomposes into ~0.30 s/step of MPI wait (differenced [mpiwait] rows, mean of 3 reps: reflux 0.11, restrict 0.07, gather 0.05, halo 0.04, parent 0.02, regrid 0.02), ~0.21 of thin AMR-family work spread over a dozen rows none above 0.06 (the gather host consume 0.054, the fill 0.026, reflux-to-parent 0.024, migration 0.022 for the one rebuild inside the differenced window, seam 0.016) and ~0.09 of residual per-block RHS inflation (the compute rows, fine advance 0.79 + base advance 0.21 + RK 0.03, sit within 4 % of the uniform-scaled whole-step ideal of 0.99 but 0.07-0.11 above the uniform coarse row scaled the way the pre-registration defined it -- about half of ledger 117's 0.1-0.2, not gone) -- so to reach the new target (excess <= 0.45, near AMReX) the 0.13 has to come out of the waits, the thin work rows and that residual, and the pre-registered candidate that could be closed without code was: the batched advance's swap/restore, 1 % of the advance on both pins, closed before any code
230
+
231
+
**Pre-registration (notes/ledger_drafts/l132_prereg_rebaseline.md, written before the run).** (1) MFC excess 0.50-0.62: MET (0.583). (2) AMReX 0.33-0.39: NOT MET (0.448; its uniform arm is where ledger 117 had it, 0.079-0.083 vs 0.080-0.082, its AMR arm is +10 %, 0.86-0.91 vs 0.78-0.82 s/step -- node k004-002 vs k004-004, or the day). (3) Ratio 1.4-1.7x: BELOW the band (1.30x) because AMReX's excess is higher here, not because MFC's is lower than predicted. (4) MFC uniform falls 20-25 % from 0.212-0.236: NOT MET (0.250-0.257, +11 %) -- the uniform arm on this node is slower than ledger 117's node, so the cross-node comparison of absolute steps is not readable; the same-node comparison against ledger 125's pin 96966782 is below. Falsifier "excess >= 0.68": did not fire (0.583). Falsifier "MFC AMR arms climb > 3 % monotone": FIRED at the letter (343.0 / 346.1 / 353.5 s, +3.1 %, monotone) and its instruction (re-run before reading) was NOT followed -- the arms are read as they stand, with the climb inside the sd (excess by rep 0.54 / 0.60 / 0.62); the 96966782 run that followed is a different pin and does not discharge it.
232
+
233
+
**Protocol.** twocode_fix.sh = twocode_clean.sh (ledger 117's protocol: MFC AMR 40/240 x3, MFC uniform 20/60 x3, AMReX AMR and uniform 40/240 x3, consecutive, differenced, excess = AMR - uniform x (cells advanced / 400^3), MFC 3.911, AMReX 5.377) writing to its own directory; the pin 7ae6f2af (up/mega tip + the EOS fix: not landed, ledger 131 blocked on the MHD GPU failures, but byte-identical to 96966782 on this ideal-gas deck); nothing else on the node (the bisect chain released the GPU lock at 18:41; no builds). logs/twocode-fix-7ae6f2af{,.log}; twocode_table.py prints the table and the differenced [phase] rows.
**The regressed pin on the same node, the next hour (20:03-20:51, logs/twocode-fix-96966782).** Ledger 125's pin 96966782 (the master-merge lineage without the halo merge, fold merge, rebuild interleave, seam early post or the EOS fix): MFC excess 0.758 (sd 0.083; AMR 1.908-2.024, uniform 0.284-0.330, no climb: the walls fell 441.6 / 428.7 / 418.6), AMReX 0.541 (sd 0.051; its AMR arm 0.93-1.03 against 0.86-0.91 an hour earlier, uniform 0.078-0.081 unchanged) -- 1.40x. Two readings: (i) AMReX's own excess moves 0.45 -> 0.54 (+21 %; its AMR arm +10 %) between consecutive hours on one node with nothing else running, and a third window the same evening (21:00-21:47, the sorted-batch pin, AMReX 0.549 with sd 0.135) makes the three ratios 1.30 / 1.40 / 1.24x, so a ratio here is good to about +-0.15x; (ii) the tip pin's 0.583 against 0.758 is -0.175 s/step (-23 %) for the four schedule increments plus the EOS fix together, not separable here, and the rows say where it went: rhs 0.955 -> 0.789 and coarse 0.325 -> 0.214 (the fix, compute), reflux 0.161 -> 0.115, regrid 0.111 -> 0.065 (the interleave), seam 0.072 -> 0.016 and b:halo 0.061 -> 0.0 (the seam post and the halo merge; gather rose 0.075 -> 0.115 as ledger 130 found), restr 0.128 -> 0.121; the uniform ideal fell 1.21 -> 0.99 with the fix.
245
+
246
+
**Rows (MFC AMR, 240-40 differenced, s/step, mean over ranks / slowest rank, mean of 3 reps).** rhs 0.789 / 0.846; coarse 0.214 / 0.220; restr 0.121 / 0.130 (its wave waits 0.057; reflux-to-parent rs:rfp 0.024); reflux 0.115 / 0.179 (rf:wait 0.106 -- 92 % wait); gather 0.115 / 0.123 (gw:wait 0.046; the host consume h:unpk + h:fill 0.054); regrid 0.065 (rg:mig 0.022, of which rg:move is the nested second half, rg:build 0.007; one rebuild inside the differenced window); halo 0.041 / 0.113; rk 0.030; gfill 0.026; seam 0.016; mg:wait 0.014; swap 0.007. [mpiwait] rows differenced (240-40, mean of 3 reps): reflux 0.106, restr 0.066, gather 0.046, halo 0.035, pgather 0.025, regrid 0.020, b:halo 0.000 (its 3.2 s per 240 steps is start-up, already accrued by step 40) -> 0.299 s/step, 19 % of the step, 51 % of the excess. Batched-advance overhead from amr_batch_r*.log on today's k004-009 pairs (both pins, 240 steps): swap 2.0-3.3 s + restore 0.3-0.8 s per rank against rhs 248-283 s + rk 20-22 s -> 0.9-1.3 % of the advance (1.0-1.4 % against rhs alone; k004-009's step reads are void for stalls but the fraction is intra-run and ~1 % on the unaffected ranks too): CLOSED, not a lever.
247
+
248
+
**Reading.** On one node the excess is half wait, a third thin work and the rest residual per-block RHS inflation. The wait is the sum-of-maxes of GOAL v7 at 8 ranks: the fine advance's slowest rank is 7 % over the mean (0.846 vs 0.789) and every family after it pays that skew at its own barrier (reflux 0.09, restrict 0.06, gather 0.04). The work has no single row worth a kernel campaign: the largest, the gather's host consume at 0.054, then migration at 0.022 for the one rebuild inside the window. AMReX's 0.45 is presumably its own fill + reflux + average-down + regrid at 5.4x base cells (inference: its run.log carries no profiler rows). The lever chosen from these rows and pre-registered the same evening (l133_prereg_batchdefrag.md) is the batching's fragmentation: half the ranks run 71-98 % more batches for the same cells (within 2 %) (rank 3: 4980 batches, 900 single-block launches, against 2520 on rank 1) and their advance is 7-14 % slower, which is the skew every barrier pays; the pad tolerance is not the cause (identical histograms at 0.10-1.00: prereg A's falsifier fired), the greedy leader order is part of it (largest-first order, c8e84bbb, byte-identical to the tip pin 98c2050f; it cuts the per-rank spread of advance compute by about a third, max/min 1.14 -> 1.08-1.12 and max/mean 1.08 -> 1.06 against prereg B's < 1.04, and calls 157 -> 135 per 40 steps, but singles do not go to 0), but its wall read is not comparable to the pin above: c8e84bbb branches from a35a8ad5 and does NOT carry the EOS fix, which is itself worth +20 % on rhs and on the uniform arm (its uniform 0.286-0.318 is the no-fix pin's, not 7ae6f2af's), so the sort must be rebuilt on the EOS pin before any wall read; against its own code parent 98c2050f the identity run at 19:57 shows +8.6 % wall (108.25 vs 99.65 s on the 60-step deck), an open cost -- ledger 133.
249
+
250
+
**Deferred, stated per GOAL v7 item 6:** the sort's wall read on the EOS pin and its +8.6 % against its parent; the covered level-0 cells (an absolute-step item, not an excess item under this metric); the np32/np64 curve and the np16 reads of the schedule increments (batch queue).
251
+
229
252
## 2026-09-09 (130) — THE CLEAN NP8 READS OF THE RENDEZVOUS AND REBUILD INCREMENTS (GOAL v7 items 2-3, timing addenda to ledgers 122, 123, 126, 129): on one healthy, otherwise idle node (k004-009, hold job 411158; every pair a fresh 40- and 240-step launch of the rung deck, differenced) against ledger 125's fix pin at 3.801 s/step -- the owner-interleaved rebuild walk (126) is -6.6 % (3.550 s/step; regrid row 133 -> 93 s per 240 steps, rg:build 87 -> 47, the per-rank parent-gather wait 6-67 s flattened to 17-29 and the rebuild's end barrier 69 -> 6 s on rank 0; wait total 317 -> 254 s per rank), the fold merge (123) is +1.3 % (3.851; waits unchanged at 317), the halo merge (122) is -7.5 % (3.515; b:halo wait 52 -> 4 s, halo 24 -> 37, reflux 59 -> 49; waits 317 -> 271), the pre-wire control pin 2e1c5356 on the same node is 2.949 s/step (-22.4 %; per-rank rhs 248-257 s per 240 steps, flat), and the up/mega tip with all four plus the seam early post (129) read -3.1 % (3.684; the seam wait 29.7 -> 0.1 s but gather 5.9 -> 26.3, pgather 16.0 -> 25.7, halo 24.2 -> 40.9, reflux 59.0 -> 65.9; waits 317 -> 270) on a single pair taken before 13:00 whose raw directory the 14:00 repetition then overwrote -- only its comparator row survives (logs/np8pair-seampost-411158-FIRSTREAD.txt), so it cannot be re-checked for the GPU stalls described below and stands as an UNVERIFIED single read -- single pairs each, at a floor of 0.7 % on a differenced step (ledger 120); whether the four are additive (interleave and halo merge alone -6.6 and -7.5 %, the union's unverified read -3.1 %) is OPEN until a healthy node repeats the union pin; second repetitions of the four pairs were run on the same node and are VOID: from ~14:00 k004-009 stalls two GPUs intermittently (the base pin itself re-read 5.155 s/step at 14:48, +36 %, with ranks 3 and 5 alone +13 % in pure compute, in bursts confined to a few 20-step blocks; temperatures 40-54 C, 185 W, no leaked process, ownership byte-equal; the repetitions of the halo merge, interleave, seam post and fold read +14, +36, +74 and +12 %), so the first reads of the fold, interleave and halo merge and the control pair, all taken before 14:00 with flat per-rank compute, stand as single reads; the seam-post repetition of 14:02-14:18 (4.097, +7.8 %) carries the two-rank burst signature (rhs 397/399 on ranks 0 and 3 against 338-353; batched-advance compute 424/427 against 361-378) and is void on the same criterion; nothing here is scored as final until a healthy node repeats them -- and item 4's own falsifier is structurally tripped on this deck (notes/item4_interior_first_design.md: ~1-6 % of blocks are free of cross-rank fill dependencies)
230
253
231
254
**Pre-registrations scored.** 122 (halo merge): (1) PH_HALO row < 20 ms/step: NOT MET (the [phase] row went 116 -> 168 ms/step: the hoisted exchange carries the wait, [mpiwait] halo 24.2 -> 36.8 s per 240 steps); (2) b:halo does not grow: met (52 -> 4); the falsifier "the wait merely moved" is PARTLY: half moved (halo +13, gather +7), half vanished (total -46 s); (3) step -2 to -3 %: EXCEEDED (-7.5 %, one pair). 123 (fold merge): (1) restr wait falls by >= P1's share: NOT MET (restr 43.7 -> 42.5, within noise); (2) step -1 to -2 %: NOT MET (+1.3 %); falsifier (the P2 WAITALL absorbs P1's wait unchanged) FIRED. 126 (interleave): (1) pg:recv per rank <= 15 s: NOT MET as a bound (17-29 s) but flat and no longer monotone in rank as predicted (the 126 falsifier did not fire); rb:xchg on rank 0 < 10 s: met (6 s); (2) rg:build 88.5 -> 30-45: NOT MET as a bound (47.5) though the direction and the per-rebuild halving (22 -> 12 s) landed; regrid 135 -> 80-95: met (93); (3) step -5 to -8 %: MET (-6.6 %). 129 (seam early post; read only in the union pin, on the unverified first read): (1) WT_SEAM falls >= 60 %: met (29.7 -> 0.1 s); (2) the SUM of the three fill waits falls by >= 40 % of the seam's share (>= 12 s): NOT MET -- gather + pgather + seam 51.6 -> 52.1 s against the base pin, 59.9 -> 52.1 against the halo-merge pin (which shares the merged halo; -7.8 s, 26 % of the seam's share): the seam's wait largely reappears in the gather and parent rows (the falsifier FIRED on this read; the void repetition shows the same rows at 29.0 / 27.3 / 0.2); (3) step -1 to -2 %: not separable in the union pin.
0 commit comments