You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Ledger 133: five np16 rungs on a healthy node pair (doublings 1.23-1.33x, the schedule increments do not move weak scaling); item 2(c) and item 3(c)/ledger 121 closed by written decision; the np32/np48 rungs void on an unexplained scale-dependent NaN with the mesh never refining, not node health; np64 cannot be placed on this machine
Copy file name to clipboardExpand all lines: docs/documentation/amr_action_plan.md
+8Lines changed: 8 additions & 0 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -226,6 +226,14 @@ possible while AMR aborts on the target machine at 1 rank, and every increment b
226
226
on a compiler that does not reproduce it. It also means the ladder should add a CCE arm as soon as one
227
227
exists, or the same class of breakage will keep accumulating undetected.
228
228
229
+
## 2026-09-10 (133) — THE NP16 READS OF THE SCHEDULE INCREMENTS, TWO WRITTEN CLOSURES, AND THE MACHINE'S CEILING ON THE CURVE (GOAL v7 items 2, 3, 5): five np16 rungs on one healthy node pair (k004-001 + k004-005, 2026-09-09 18:29-22:13, jobs 411779/411172/411316/411651/411780, each a fresh np8 pair on node A and an np16 pair across both, differenced, the per-rank fine-advance rows flat at 330-386 s per 240 steps on every run) put the doubling at 1.230x on ledger 125's pin, 1.254x with the fold merge, 1.256x with the rebuild interleave, 1.265x on the union pin with the seam early post, and 1.331x with the halo merge alone -- the schedule increments do not move weak scaling (GOAL v7's <= 1.30x holds on every pin but the halo merge's), the np16 step sits at 4.45-4.97 s on the lineage that carries ledger 131's unlanded +37 % RHS regression, and the np8 level on the same node varies by hour (3.52-4.04 s) more than by pin, so only each rung's own ratio and its wait rows are read; item 2(c) (the base-halo ISEND/IRECV design, notes/item2c_base_halo_isend_design.md) is CLOSED by written decision -- ledger 130's falsifier fired at np8 and the union pin's np16 rows repeat it (b:halo 118 -> 4 s per rank per 240 steps and seam 61 -> 0, while halo 58 -> 98, gather 16 -> 41, pgather 45 -> 64; the halo-merge pin alone, the clean single-increment control, shows the same: b:halo 118 -> 4 while halo 58 -> 107 with nothing else changed -- the wait moves to the next rendezvous), so a third rendezvous cut on the same path is not built; item 3(c) (ledger 121, the level-2 overlap copy) is CLOSED by written decision on ledger 119's finding that the larger, persistent part of the np16 wait exists with no rebuild in sight (its post-rebuild aftermath is real, 22-30 % of every np16 wait row, but 119 could not separate it from the faulting node and deferred item 3 on that basis); and item 5's np64 rung cannot be placed on this machine (the mi2508x partition has ten nodes, two down -- k004-007 failed its HPL test, k004-010 not responding -- and two unusable, k004-003 with the faulting GPU at PCI 14:00.0 and k004-006 with a dead IB port; k004-009 stalls two GPUs intermittently per ledger 130; so np48 on six nodes is the ceiling and only k004-001 and k004-005 are trusted for timing): the curve of ledger 128 is judged np8 -> np32 and np48, and the first rungs at both, run 2026-09-10 (np32 411707 at 07:35 on nodes 001/004/005/009; np48 412296 at 13:01 on 001/002/004/005/008/009), are VOID for a reason that is NOT node health -- both arms of both rungs ran to their final step on the base grid alone (Time step 221 of 241 at 0.28-0.33 s/step; rebuilds_incl_seed=1, no [amr-bat]/[amr-merge]/[amr-snap]/[amr-cad] rows, 25-31 % GPU memory against 78-98 % at np16: the mesh never refined past the seed) and then every rank NaN'd and invoked MPI_ABORT at the output step (32 of 32, 48 of 48, on all six nodes including the pair the np16 rungs ran clean on the previous evening), so the np32/np48 curve is blocked on an unexplained scale-dependent NaN at the 4-way decompositions (4x4x2, 4x4x3), a statement-3 matter, to be reproduced on one healthy node with four ranks per GPU (the runs used a quarter of device memory) before any rung is re-queued
230
+
231
+
**Data.** logs/np16-rung-{fix,fold,interleave,seam,halomerge2}-{411779,411172,411316,411651,411780} (ARM lines in logs/np16-rung-<job>.log; [mpiwait] and [phase-rank] rows in <dir>/np{8,16}/run240/run.log). np8 / np16 differenced steps: fix pin a6dd813c 4.040 / 4.969; fold e8e2486c 3.556 / 4.459; interleave 5749447e 3.824 / 4.804; seam-post union 98c2050f 3.518 / 4.449; halo merge 646bbf6d 3.630 / 4.831. np16 [mpiwait] mean per rank per 240 steps (halo, b:halo, gather, pgather, seam, reflux, restr, regrid): fix pin 58 / 118 / 16 / 45 / 61 / 64 / 89 / 87; fold 44 / 84 / 13 / 42 / 57 / 46 / 66 / 84; interleave 55 / 113 / 14 / 43 / 60 / 59 / 82 / 37; union 98 / 4 / 41 / 64 / 0 / 61 / 85 / 37; halo merge 107 / 4 / 23 / 59 / 67 / 70 / 89 / 86. The node-0 skew of ledger 119 survives every increment, moving rows with the merge: b:halo's slowest rank is rank 0 on the fix, fold and interleave pins (274 / 237 / 272 s against means 118 / 84 / 113), and on the union and halo-merge pins, where b:halo falls to ~4 s, it reappears as the halo row's rank 0 (253 and 269 s against means 98 and 107). The rebuild interleave alone takes the np16 regrid row from 87 to 37 s (-58 %; at np8 in these rungs 60 -> 36, -40 %; ledger 130's clean np8 read 133 -> 93).
232
+
233
+
**Rendezvous per step, tabled (notes/rendezvous_table.md):** 21 rendezvous / 51 blocking calls per step before GOAL v7; 20 / 32 on the union pin (halo merge 36 -> 18 SENDRECVs, the freg wave folded into the restrict wave); the seam early post removes none. Item 2's target of ~10 rendezvous is not reached and, by the falsifier, not pursued further.
234
+
235
+
**Deferred, stated per GOAL v7 item 6:** the np8/np16 reads of ledger 131's fix once it lands (the attributor-cap link flag is the next experiment for its MHD GPU failure); the batch-defragmentation lever pre-registered in ledger 132 (notes/ledger_drafts/l133_prereg_batchdefrag.md: pad tolerance falsified, largest-first order identity-gated with the per-rank spread cut by a third, wall read pending on the fixed lineage) moves to ledger 134; the np32 NaN reproduction and diagnosis.
236
+
229
237
## 2026-09-09 (132) — SCORECARD ITEM 2 RE-BASELINED ON THE FIXED PIN, ONE NODE, ONE WINDOW (GOAL v7 item 7a, the user-authorized single-node pivot): MFC's steady AMR excess is 0.58 s/step (3 reps, sd 0.04) against AMReX's 0.45 (sd 0.02) on k004-002 between 18:43 and 19:25 -- 1.30x, against ledger 117's 1.94x on k004-004 (0.70 vs 0.36) -- and the excess decomposes into ~0.30 s/step of MPI wait (differenced [mpiwait] rows, mean of 3 reps: reflux 0.11, restrict 0.07, gather 0.05, halo 0.04, parent 0.02, regrid 0.02), ~0.21 of thin AMR-family work spread over a dozen rows none above 0.06 (the gather host consume 0.054, the fill 0.026, reflux-to-parent 0.024, migration 0.022 for the one rebuild inside the differenced window, seam 0.016) and ~0.09 of residual per-block RHS inflation (the compute rows, fine advance 0.79 + base advance 0.21 + RK 0.03, sit within 4 % of the uniform-scaled whole-step ideal of 0.99 but 0.07-0.11 above the uniform coarse row scaled the way the pre-registration defined it -- about half of ledger 117's 0.1-0.2, not gone) -- so to reach the new target (excess <= 0.45, near AMReX) the 0.13 has to come out of the waits, the thin work rows and that residual, and the pre-registered candidate that could be closed without code was: the batched advance's swap/restore, 1 % of the advance on both pins, closed before any code
230
238
231
239
**Pre-registration (notes/ledger_drafts/l132_prereg_rebaseline.md, written before the run).** (1) MFC excess 0.50-0.62: MET (0.583). (2) AMReX 0.33-0.39: NOT MET (0.448; its uniform arm is where ledger 117 had it, 0.079-0.083 vs 0.080-0.082, its AMR arm is +10 %, 0.86-0.91 vs 0.78-0.82 s/step -- node k004-002 vs k004-004, or the day). (3) Ratio 1.4-1.7x: BELOW the band (1.30x) because AMReX's excess is higher here, not because MFC's is lower than predicted. (4) MFC uniform falls 20-25 % from 0.212-0.236: NOT MET (0.250-0.257, +11 %) -- the uniform arm on this node is slower than ledger 117's node, so the cross-node comparison of absolute steps is not readable; the same-node comparison against ledger 125's pin 96966782 is below. Falsifier "excess >= 0.68": did not fire (0.583). Falsifier "MFC AMR arms climb > 3 % monotone": FIRED at the letter (343.0 / 346.1 / 353.5 s, +3.1 %, monotone) and its instruction (re-run before reading) was NOT followed -- the arms are read as they stand, with the climb inside the sd (excess by rep 0.54 / 0.60 / 0.62); the 96966782 run that followed is a different pin and does not discharge it.
0 commit comments