You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Ledger (107): scorecard item 2 re-measured on the shipped defaults, three forms, three AMReX arms -- MFC's steady AMR excess 0.94 s/step (sd 0.06; ledger 86: 1.33) vs AMReX 0.36, 2.6x on the 2x target; relative-to-ideal indistinguishable; a ghost-width-matched AMReX build lowered its own excess 14% (the metric charges ghost width to the physics denominator), so halo width stays untested; two thirds of the excess is the AMR-only families and a sixth is unbracketed
Copy file name to clipboardExpand all lines: docs/documentation/amr_action_plan.md
+72Lines changed: 72 additions & 0 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -226,6 +226,78 @@ possible while AMR aborts on the target machine at 1 rank, and every increment b
226
226
on a compiler that does not reproduce it. It also means the ladder should add a CCE arm as soon as one
227
227
exists, or the same class of breakage will keep accumulating undetected.
228
228
229
+
## 2026-09-08 (107) — SCORECARD ITEM 2 RE-MEASURED ON THE SHIPPED DEFAULTS, THREE FORMS, THREE AMReX ARMS: MFC's steady AMR excess is 0.94 s/step (3 reps, sd 0.06; ledger 86: 1.33) against AMReX's 0.36, 2.6x on the 2x target; relative to each code's own ideal the two are indistinguishable (MFC 1.01, sd 0.14; AMReX 0.82-1.04 depending only on mesh), and a ghost-width-matched AMReX build (NUM_GROW 4 = MFC's WENO5 buff_size) LOWERED AMReX's excess 14 % -- the excess metric charges ghost width to the physics denominator on both codes, so it can neither convict nor exonerate MFC's wider halos; what it does show is that two thirds of MFC's excess is the AMR-only families (reflux, regrid, gather, seam, fine halo) and a sixth sits outside every phase bracket
230
+
231
+
**Question.** After ledgers 89-106 the item-2 number was an estimate stitched across days, and the comparison's fairness
232
+
was in question (heavier per-cell physics, wider halos, more variables on MFC's side). Pre-registered
233
+
(notes/ledger_drafts/l107_prereg.md, written with the script): MFC excess 0.8-1.0 s/step, 2.1-2.6x AMReX; relative form
234
+
~1.1-1.2x; per-cell ~3x; today's stock AMReX build reproduces the campaign binary; the NUM_GROW-4 build's excess 10-30 %
235
+
ABOVE stock; falsifier: if grow-4 raises AMReX's excess by less than 10 %, ghost width is not where MFC's extra cost sits.
236
+
237
+
**Protocol.** Ledger 86's twocode.sh, changed only in: the MFC arms carry the shipped defaults (``amr_batched_advance``,
238
+
``amr_bat_pad = 0.10``, ``amr_device_pack``; batched gather off), the binary is c3bc2c51 (the up/mega tip 5986d9b3
239
+
differs from it only in docs and toolchain), the GPU lock is held, and three AMReX binaries run instead of one: the
240
+
campaign binary of ledger 86 (built 2026-09-01), today's rebuild of the working tree at its stock NUM_GROW = 2, and the
241
+
same tree at NUM_GROW = 4 (MFC's ``buff_size`` on this deck: WENO5, inviscid, ``weno_polyn + 2``); MFC's 6 state
242
+
variables vs CNS's 7 were left as they are. Hold 408703 on k004-003, one allocation, 08:58-10:09, three interleaved reps,
243
+
AMReX and MFC pairs differenced from scratch (240-40 and 60-20). Excess = AMR s/step - uniform s/step x (cells advanced
244
+
per step / 400^3); the relative form divides that by the scaled uniform ("ideal"); the per-cell form divides by cells
245
+
advanced per step.
246
+
247
+
| arm (3 reps, mean, sd) | AMR s/step | uniform s/step | cells / base | excess s/step | excess / ideal | us per cell-update |
MFC's AMR step is flat across reps (1.906 / 1.837 / 1.903); its uniform step spreads 9 % (0.262 / 0.221 / 0.237) over a
256
+
40-step difference, and scaled by 3.911 that alone moves the excess by about +/-0.08 -- the reps' 0.880 / 0.975 / 0.977
257
+
are mostly that.
258
+
259
+
**What held and what did not.** Prediction 1 held: 0.944 against 0.8-1.0, 2.63x against 2.1-2.6x; the target (<= 0.72 at
260
+
today's AMReX 0.36) is not met, and the reading is far tighter than ledger 86's (1.51 / 1.21 / 1.26). The campaign
261
+
binary reproduced ledger 86's AMReX excess (0.359 vs 0.388) on the same mesh (5.38x base cells). Prediction 3 FAILED on
262
+
mesh and step: today's stock rebuild refines 4.14x base cells and steps 0.678 s against the campaign binary's 5.38x and
263
+
0.797 s -- consistent with the working tree's uncommitted tagging edit (relative density gradient, dated 2026-09-02,
264
+
after the campaign binary; MFC-like, not MFC-equivalent), though the campaign binary's source cannot be read back, so
265
+
this is inferred. Its excess happens to agree (0.346). Today's mesh is the closer match to MFC's 3.91x, and today's
266
+
NUM_GROW-2 build is byte-identical to the 2026-09-02 ``_reltag`` binary, so the grow-2 / grow-4 pair differs in nothing
267
+
but ghost width. Prediction 4 was wrong in SIGN: NUM_GROW 4 raised AMReX's AMR step 16 % (0.678 -> 0.786) but its uniform
268
+
step 47 % (0.080 -> 0.118), so the ideal rose more than the AMR step and the excess FELL 14 % (0.346 -> 0.298). The
269
+
falsifier fired -- but read narrowly: the experiment varied AMReX's ghost width, never MFC's, and what it establishes is
270
+
that this metric is nearly blind to ghost width (it lands in the denominator on both codes alike, since MFC's uniform
271
+
run carries the same ``buff_size`` as its AMR run). A metric that cannot see a cost can neither convict nor exonerate
272
+
it; the direct test of "MFC pays for its halo width" is an MFC arm at a narrower stencil, not run here. Prediction 2's
273
+
relative form: MFC 1.01 (sd 0.14) against AMReX 0.82 or 1.04 -- indistinguishable, and the same AMReX code moves by 25 %
274
+
on mesh alone, so the form is too mesh-sensitive to headline.
275
+
276
+
**Where MFC's 0.94 sits (240-40 differenced, mean over ranks, per rep, s/step).** Physics: fine ``rhs`` 0.81 + ``coarse``
277
+
0.29-0.31 = 1.09-1.12 against an ideal of 0.86-1.03 from the uniform step, a per-block inflation of 0.09-0.23 (ledger
278
+
86's 0.4-0.55 counted the fine halo inside physics; on that definition today's is 0.15-0.30). The base-grid halo
279
+
bracket ``b:halo`` (0.07-0.09) is called from both the coarse and the fine RHS (4335 calls = 720 + 3615), so it cannot be
280
+
assigned to the coarse row and is not compared to the uniform run's. AMR-only families: reflux 0.15-0.18, regrid
281
+
0.15, gather 0.08-0.09, seam 0.08-0.09, fine halo 0.06-0.07, gfill 0.03, rk 0.03, swap 0.01 -- 0.58-0.62 s/step, two
282
+
thirds of the excess; ghost-fill WORK (fine halo + seam + gather + the b:halo share) is about 0.30 of it, so "not in
283
+
halos" would be false even though halo WIDTH is unmeasured. The bracketed rows sum to 1.67-1.74 of an AMR step of
284
+
1.84-1.91: 0.15-0.17 s/step, a sixth of the excess, sits outside every phase bracket and is unattributed. AMReX's whole
285
+
excess is 0.30-0.36.
286
+
287
+
**How the three forms disagree, and which to read.** Absolute seconds (2.6x) favour the lighter code: the same
288
+
bookkeeping inflates a 0.08 s uniform step less than a 0.24 s one. Relative-to-ideal (1.0-1.2x) hides the absolute
289
+
seconds behind MFC's heavy physics and moves 25 % with the mesh. Per cell advanced (2.9-3.6x) penalises the code that
290
+
refines LESS (MFC 3.9x vs AMReX 4.1-5.4x base cells) for the same fixed costs. The scorecard keeps the absolute form
291
+
because the horizon is stated in seconds per step; the other two are reported beside it so the number cannot be argued
292
+
in either direction.
293
+
294
+
**What it means.** The gap to the 2x target is 0.22 s/step of MFC's 0.94. The excess decomposes as AMR-only families
295
+
0.6 (reflux 0.16 and regrid 0.15 the largest -- the skew wait and the O(P) term earlier ledgers named), per-block RHS
296
+
inflation 0.09-0.23, and 0.15-0.17 unbracketed. Variable count was never the issue (MFC carries fewer). Halo width is
297
+
untested on MFC's side and the metric cannot test it; ghost-fill work is a third of the excess. A second reference
298
+
framework would change none of these numbers; the two things that would are an MFC narrower-stencil arm (halo width)
299
+
and brackets for the missing sixth.
300
+
229
301
## 2026-09-08 (105) — THE 2-NODE RUNG FOUND A CORRECTNESS CLIFF, NOT A SCALING NUMBER: the global box union (every rank's PRE-MERGE bisection leaves, ~1000 per rank) was truncated to amr_max_blocks before the merge, in rank order, so at 16 ranks the last ranks' leaves were dropped at every regrid -- 42% of the level-1 tags fell on cells that never refined (np8: 0%) and the weak-scaled np16 kept 71-80/512-584 boxes of the 128/1024 its doubled domain owns; the accepted arrays now grow to the union and the cap applies to the merged set -- np8 byte-identical; goldens 71/71 on both lanes; rung rerun VALID: np8 -> np16 (weak) = 1.586x per doubling against the 1.20x bar, every cross-node phase 1.6-3.5x, compute flat, InfiniBand confirmed and the tcp lane ruled out
230
302
231
303
**What the rung showed (job 408425, pinned 8644c8b4, np8 on one node vs np16 on two, 40- and 240-step pairs, int=20).**
0 commit comments