Performance: Avoid rewrite work in passes when no patterns can match - #5315
Open
taalexander wants to merge 3 commits into
Open
Performance: Avoid rewrite work in passes when no patterns can match#5315taalexander wants to merge 3 commits into
taalexander wants to merge 3 commits into
Conversation
Avoid invoking six narrow greedy rewrite passes when their typed pattern roots are absent. A private variadic helper performs one interrupting operation walk, while root-present execution retains the complete existing rewrite path. This reduces root-free MLIR pass time by 66% to 78% across 100k to 1.5M gate inputs in the initial benchmark matrix. Baseline and candidate outputs were byte-identical in every measured root-free case and existing root-present comparisons. Verified with the focused transform and AST lit tests, all 41 optimizer unit tests, full hardware and emulation pipeline comparisons, clang-format, and git diff --check. Signed-off-by: Thomas Alexander <talexander@nvidia.com>
CI Summary (
|
| Job | Result |
|---|---|
binaries |
⏩ skipped |
build_and_test |
✅ success |
changes |
✅ success |
config_devdeps |
✅ success |
config_source_build |
⏩ skipped |
config_wheeldeps |
✅ success |
devdeps |
✅ success |
docker_image |
⏩ skipped |
gen_code_coverage |
⏩ skipped |
metadata |
✅ success |
python_metapackages |
⏩ skipped |
python_wheels |
⏩ skipped |
source_build |
⏩ skipped |
wheeldeps |
✅ success |
⏩ Skipped jobs (7) — intentionally skipped on PR builds; run on merge_group / workflow_dispatch
| Job |
|---|
binaries |
config_source_build |
docker_image |
gen_code_coverage |
python_metapackages |
python_wheels |
source_build |
All sub-jobs (43) — every matrix leg, with links
| Job | Status | Link |
|---|---|---|
| Build and test (amd64, gcc12, openmpi) / Dev environment (Debug) | ✅ success | view |
| Build and test (amd64, gcc12, openmpi) / Dev environment (Python) | ✅ success | view |
| Build and test (amd64, llvm, openmpi) / Dev environment (Debug) | ✅ success | view |
| Build and test (amd64, llvm, openmpi) / Dev environment (Python) | ✅ success | view |
| Build and test (arm64, llvm, openmpi) / Dev environment (Debug) | ✅ success | view |
| Build and test (arm64, llvm, openmpi) / Dev environment (Python) | ✅ success | view |
| CI Summary | ❔ in_progress | view |
| Check for stable CUDA-Q changes | ✅ success | view |
| Configure build (devdeps) | ✅ success | view |
| Configure build (source_build) | ⏩ skipped | view |
| Configure build (wheeldeps) | ✅ success | view |
| Create CUDA Quantum installer | ⏩ skipped | view |
| Create Docker images | ⏩ skipped | view |
| Create Python metapackages | ⏩ skipped | view |
| Create Python wheels | ⏩ skipped | view |
| Gen code coverage | ⏩ skipped | view |
| Load dependencies (amd64, gcc12) / Caching | ✅ success | view |
| Load dependencies (amd64, gcc12) / Finalize | ✅ success | view |
| Load dependencies (amd64, gcc12) / Metadata | ✅ success | view |
| Load dependencies (amd64, llvm) / Caching | ✅ success | view |
| Load dependencies (amd64, llvm) / Finalize | ✅ success | view |
| Load dependencies (amd64, llvm) / Metadata | ✅ success | view |
| Load dependencies (arm64, gcc12) / Caching | ✅ success | view |
| Load dependencies (arm64, gcc12) / Finalize | ✅ success | view |
| Load dependencies (arm64, gcc12) / Metadata | ✅ success | view |
| Load dependencies (arm64, llvm) / Caching | ✅ success | view |
| Load dependencies (arm64, llvm) / Finalize | ✅ success | view |
| Load dependencies (arm64, llvm) / Metadata | ✅ success | view |
| Load source build cache | ⏩ skipped | view |
| Load wheel dependencies (amd64, 12.6) / Caching | ✅ success | view |
| Load wheel dependencies (amd64, 12.6) / Finalize | ✅ success | view |
| Load wheel dependencies (amd64, 12.6) / Metadata | ✅ success | view |
| Load wheel dependencies (amd64, 13.0) / Caching | ✅ success | view |
| Load wheel dependencies (amd64, 13.0) / Finalize | ✅ success | view |
| Load wheel dependencies (amd64, 13.0) / Metadata | ✅ success | view |
| Load wheel dependencies (arm64, 12.6) / Caching | ✅ success | view |
| Load wheel dependencies (arm64, 12.6) / Finalize | ✅ success | view |
| Load wheel dependencies (arm64, 12.6) / Metadata | ✅ success | view |
| Load wheel dependencies (arm64, 13.0) / Caching | ✅ success | view |
| Load wheel dependencies (arm64, 13.0) / Finalize | ✅ success | view |
| Load wheel dependencies (arm64, 13.0) / Metadata | ✅ success | view |
| Prepare cache clean-up | ❔ in_progress | view |
| Retrieve PR info | ✅ success | view |
✅ Required checks (6/6) — declared in .github/required-checks.yml for push
| Required check | Status | Link |
|---|---|---|
| Build and test (amd64, llvm, openmpi) / Dev environment (Debug) | ✅ success | view |
| Build and test (amd64, llvm, openmpi) / Dev environment (Python) | ✅ success | view |
| Build and test (arm64, llvm, openmpi) / Dev environment (Debug) | ✅ success | view |
| Build and test (arm64, llvm, openmpi) / Dev environment (Python) | ✅ success | view |
| Build and test (amd64, gcc12, openmpi) / Dev environment (Debug) | ✅ success | view |
| Build and test (amd64, gcc12, openmpi) / Dev environment (Python) | ✅ success | view |
Signed-off-by: Thomas Alexander <talexander@nvidia.com>
atgeller
approved these changes
Aug 28, 2026
atgeller
left a comment
Collaborator
There was a problem hiding this comment.
One nit, otherwise LGTM.
Describe scaling checks, rewrite-driver choices, avoiding repeated IR work, and timing and tracing. Rename the operation-presence helper for clarity. Signed-off-by: Thomas Alexander <talexander@nvidia.com>
schweitzpgi
reviewed
Aug 31, 2026
| LLVM_DEBUG(llvm::dbgs() << "Before constant prop:\n" << func << '\n'); | ||
|
|
||
| if (failed( | ||
| applyPatternsGreedily(func.getOperation(), std::move(patterns)))) { |
Collaborator
There was a problem hiding this comment.
From the description, I was surprised to see another pre-traversal being added instead of converting these passes to partial conversion passes in favor of greedy rewrites. Is there a reason the latter was not done?
Collaborator
Author
There was a problem hiding this comment.
Some of these passes do use the greedy rewrite behaviour so it was not drop in. Another approach would be something like I did in #5318. I'll take another look.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Six passes currently invoke MLIR's greedy rewrite driver
even when the IR contains none of the operation types their patterns can
match. On large modules, that turns a pass with no relevant work into a full IR
traversal.
This change gives those passes a small typed check. It walks the IR
once, stops at the first relevant operation, and returns early when there is no
possible pattern root. If a matching op is present, the pass runs exactly as it did
before.
I've applied this to:
LiftArrayAllocGlobalizeArrayValuesPyFrontendCleanupGetConcreteMatrixConstantPropagationStatePreparationPerformance
Benchmarking results a fixed number of operations:
At 1.5 million gates, every individual pass spent substantially less time:
LiftArrayAllocGlobalizeArrayValuesPyFrontendCleanupGetConcreteMatrixConstantPropagationStatePreparation