Skip to main content

profiling-onnx-models

Sterne13
Forks1
Aktualisiert2. Juni 2026 um 20:12

Use this skill when profiling ONNX model performance on CUDA with OnnxRuntime. Covers ORT session profiling, reading JSON results, identifying compute and memory bottlenecks, debugging memcpy nodes, and measuring GenAI pipeline throughput (tok/s, TTFT).

Installation

Mit Codex oder Claude installieren Kopieren Sie diesen Prompt, fügen Sie ihn in Codex, Claude oder einen anderen Assistant ein und lassen Sie die Skill-Seite prüfen und installieren.

SKILL.md
readonly