基于 SOC 职业分类
用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
直接命令不会经过审查 Prompt;运行前请先检查来源。
npx skills add https://github.com/mindspore-ai/akg --skill pypto-case-loss-crossentropy命令会保持在同一行。复制前请横向滚动并检查完整内容。
想先保存到本地?可下载 SkillsMP 当前能够提供的文件。
正在显示 SKILL.md
矩阵乘法矩阵乘法 A[M, K] @ B[K, N] = C[M, N]中,大K维度矩阵乘法(K>>M,N)优化:针对M/N较小但K极大(如M=N=256,K=131072)的场景,Split-K切分K维度并行化、Workspace+Reduce替代全局同步,实现显著性能提升
Triton Ascend hard API restrictions and forbidden syntax. MUST-follow rules that apply to every kernel: forbidden control flow (return/break/continue/lambda/while), tensor slice/index restrictions, scalar conversion rules, BLOCK_SIZE upper bound. Violating any of these produces a compile or runtime error on Ascend.
Triton Ascend 性能优化通用策略: BLOCK_SIZE 选择 (1024-2048 for elementwise, must be <65536), grid configuration (use VEC_CORE_NUM / CUBE_CORE_NUM, 2D/3D grid for matmul / conv / reduce, 1D grid + inner loop for elementwise / pointwise), 256B alignment for memory transfers, autotune block-size patterns, fp16 / fp32 precision conversion. Bind via keywords like matmul, elementwise, reduce, block_size, grid, autotune, alignment, fp16, fp32, tile, interleaved-loop, cube-core, vec-core.
| name | pypto-case-loss-crossentropy |
| description | 模式 D 示例:Loss — CrossEntropyLoss,展示多输入 kernel、两段 tile、softmax+gather+sum、标量输出 |
| category | example |
| version | 1.0.0 |
| metadata | {"backend":"ascend","dsl":"pypto","operator_patterns":"loss,reduction,gather,softmax"} |
def create_cross_entropy_kernel(batch, num_classes):
@pypto.frontend.jit(runtime_options=..., debug_options=...)
def kernel(
predictions: pypto.Tensor((batch, num_classes), pypto.DT_FP32),
targets: pypto.Tensor((batch,), pypto.DT_INT64),
) -> pypto.Tensor((1,), pypto.DT_FP32):
output = pypto.tensor([1], pypto.DT_FP32)
# Phase 1: per-sample softmax + gather
pypto.set_vec_tile_shapes(1024, 16)
log_probs = pypto.log(pypto.softmax(predictions, dim=1))
targets_i32 = pypto.cast(targets, pypto.DT_INT32)
idx = pypto.unsqueeze(targets_i32, 1)
picked = pypto.gather(log_probs, dim=1, index=idx)
neg_picked = pypto.mul(picked, -1.0)
# Phase 2: batch reduction
pypto.set_vec_tile_shapes(2048, 8)
total = pypto.sum(neg_picked, dim=0, keepdim=False)
output[:] = total / batch
return output
return kernel
forward:assert → contiguous → 调 kernel → reshape(1,)
pypto.cast(targets, DT_INT32) — INT64 输入需转 INT32pypto.unsqueeze + pypto.gather — 按 index 取元素pypto.mul(x, -1.0) — 取反的标准写法(规则 R2 的应用)pypto.tensor([1], ...) + output[:] = scalarreshape(-1) 用 1D kernel