Skip to content

[Perf] Optimize Ascend prepare_wy_repr_bwd with larger kv tiles, bf16 mmad, and fused a2/dg #4600

[Perf] Optimize Ascend prepare_wy_repr_bwd with larger kv tiles, bf16 mmad, and fused a2/dg

[Perf] Optimize Ascend prepare_wy_repr_bwd with larger kv tiles, bf16 mmad, and fused a2/dg #4600

Triggered via pull request September 1, 2026 15:47
Status Success
Total duration 8s
Artifacts

nvidia-h100.yml

on: pull_request
Detect NPU-only changes
4s
Detect NPU-only changes
Test H100 (PyTorch 2.12)  /  test-ops
Test H100 (PyTorch 2.12) / test-ops
Check H100 PyTorch 2.7 Eligibility
0s
Check H100 PyTorch 2.7 Eligibility
Test H100 (PyTorch 2.12)  /  test-models
Test H100 (PyTorch 2.12) / test-models
Test H100 Ops (PyTorch 2.7)  /  test-ops
Test H100 Ops (PyTorch 2.7) / test-ops
Benchmark H100 (PyTorch 2.12)  /  benchmark
Benchmark H100 (PyTorch 2.12) / benchmark
Test H100 Ops (PyTorch 2.7)  /  test-models
Test H100 Ops (PyTorch 2.7) / test-models
Fit to window
Zoom out
Zoom in