triton-ascend-case-reduction-mean-large
大规模reduce最后根轴(mean)行二次切分优化:每个kernel计算多行减少线程块数量、kernel内二次切分避免超UB,grid=40且SUB切分不含尾块时性能最优(16.00us),尾块计算会显著降低性能,适用于非reduce轴中等、reduce轴较大的2D归约场景
Source facts
- Repository
- mindspore-ai/akg
- Last source activity
- March 28, 2026 at 06:51
- Detected SKILL.md language
- Chinese
- Stars
- 259
- Forks
- 48
Install options
The review-first prompt is selected by default. You can switch to a direct command or download a local copy.
Review the source files
Read SKILL.md and any companion files shown by SkillsMP before deciding whether to install.