Skip to main content

profiling-onnx-models

Stars13
Forks1
UpdatedJune 2, 2026 at 20:12

Use this skill when profiling ONNX model performance on CUDA with OnnxRuntime. Covers ORT session profiling, reading JSON results, identifying compute and memory bottlenecks, debugging memcpy nodes, and measuring GenAI pipeline throughput (tok/s, TTFT).

Installation

Install with Codex or Claude Copy this prompt, paste it into Codex, Claude, or another assistant, and let it review the skill page and install it for you.

SKILL.md
readonly