You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Copy file name to clipboardExpand all lines: docs/documentation/amr_action_plan.md
+45Lines changed: 45 additions & 0 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -226,6 +226,51 @@ possible while AMR aborts on the target machine at 1 rank, and every increment b
226
226
on a compiler that does not reproduce it. It also means the ladder should add a CCE arm as soon as one
227
227
exists, or the same class of breakage will keep accumulating undetected.
228
228
229
+
## 2026-09-04 (67) — THE BATCHED FINE ADVANCE IS MERGED (default off): GPU gates clean, two controls prove the flag reaches the driver, 0.37 s/step recovered
230
+
231
+
`amr_batched_advance` merged as b0f601cf (4 commits, +470/-205 over 7 source/toolchain files -- 8 files and +471/-205
232
+
counting a doc line -- flag default F). It batches up to 8 same-(level, m, n, p) owned blocks TWO ghost shells apart
233
+
along the last active dimension (`amr_bat_w = amr_bat_ext(sd) + 2*buff_size + 1`; each member carries its own shell, and
234
+
that separation is exactly what makes a member's stencil unable to reach its neighbour's data), so one `s_compute_rhs` and one
235
+
RK kernel cover the batch instead of one launch per block -- attacking dispatch count, which the Step-2 instrument
236
+
identified as where the excess sits (host work between launches, 0.75-0.85 s/step).
237
+
238
+
**Gates, all quoted in task-10-row6-report.md.** Rebased onto up/mega with no conflict and the branch delta byte-for-byte
239
+
unchanged; precheck 7/7. On GPU (amdflang OpenMP offload, k004-008): 67 of 67 non-chemistry AMR goldens pass flag-off
240
+
(the 3 chemistry AMR decks need a separate build variant and are validator-excluded from the flag anyway); the two np=2
241
+
split-tower oracle decks give six `[amr-xa]` family lines each with snd==rcv on every line and counts identical flag-on
242
+
vs flag-off, with `cmp` clean at 54 files per deck because both are exact grids; `MFC_XA_SEED=1` still aborts with ORDER
243
+
ORACLE MISMATCH with the flag on; and the MFC_DEBUG arm ran on BOTH a GPU-debug and a CPU-debug build with no length or
244
+
NaN assert firing.
245
+
246
+
**Two controls, because "identical" is vacuous unless the flag is actually live.** (i) A stretched-grid deck with the flag
247
+
hand-injected aborts with "amr_batched_advance requires a uniform grid" -- this proves the namelist flag reaches the module
248
+
INIT and the guard fires, and says nothing about the driver. (ii) The
249
+
3D batch-forming deck A5DAD70D is byte-identical on its exact 64^3 variant (62/62 files) and differs by 1.066e-14 on its
250
+
original inexact 52^3 grid -- reproducing the CPU measurement exactly. That roundoff difference IS the proof the batched
251
+
path executed (only members 2..n reusing the leader's dx can produce it, and the exact-grid arm rules out the widened
252
+
allocations as the source), and it is the documented design property, not a defect.
253
+
254
+
**Two honest limits on that evidence.** First, the batching proof and the exchange oracle are on DIFFERENT decks: the two
255
+
oracle decks supply the `[amr-xa]` families and the 54-file byte-compare, but nothing in their logs records that any batch
256
+
held more than one member (they own 2-3 blocks spread across two levels), so their "identical on/off" is weak evidence
257
+
about the batching logic itself; A5DAD70D demonstrates multi-member batches at np=2 but carries no `[amr-xa]`
258
+
instrumentation. Recording `amr_bat_n` under `rank_time_wrt` is a one-line fix and is the next thing to add. Second,
259
+
flag-ON coverage is three hand-built decks and no CI: the validator excludes most physics, and no test or example sets
260
+
the flag, so 67/67 is flag-OFF validation, not flag-ON.
261
+
262
+
**What "default off" does and does not mean here.** Two of the four commits change code that runs with the flag OFF: the
263
+
lock-step advance loop now walks the owned list instead of scanning every block, and the flux-capture routine is
264
+
restructured around a member loop that reduces to the old code at batch size 0. So flag-off is a REFACTOR asserted
265
+
bit-identical and gated at 67/67 goldens plus the oracles -- not an untouched path. The residual risk is the owned list
266
+
going stale: it is correct only if every `amr_block_owner` write marks `amr_myblk_dirty`, all five sites now do, and commit
267
+
525d4b75 exists because one of them did not. Review caught that, not a test; no suite case migrates an L0 tile with an
268
+
owner change outside a regrid at np>=2, so that failure mode is invisible to single-rank goldens.
269
+
270
+
**Performance, restated with its caveat.** The pre-registered falsifier passed at 0.294 and 0.440 s/step (mean 0.367,
271
+
12.7%) against a 0.15 bar, but with GPU-aware MPI off in both arms because the only available node cannot run it. The
272
+
rdma-on A/B is queued (405052). The flag stays default-F until that, the NVHPC compile gate, and the CCE gpu-acc lane.
273
+
229
274
## 2026-09-04 (66) — RETRACTION: the AMR restart metadata has NO uninitialized padding; the comparator was misreading the file, and the tier-1 restart byte-compare gate is not blocked
230
275
231
276
Ledger 64 reported that `restart_data/lustre_amr_<step>.dat` carries 6 NaN words and 11 garbage words of uninitialized
0 commit comments