Skip to main content

inference-workload-profile

Profile live inference traffic into a token/shape distribution artifact that drives profile-matched speculative-decoding draft training -- an analog of Fireworks FireOptimizer's "profile-driven customization" (the documented source of its higher draft hit-rate). Reads an OpenAI-style access JSONL, emits workload-profile.json (input/output length distributions, content-class mix, ISL/OSL bench shapes, and a spec-decode method recommendation), and hands off to inference-spec-decode-train via a hit-rate-matched corpus. The first phase of the adaptive spec-decode loop. Triggers on "profile my workload", "workload profile for spec-decode", "what draft should I train", "match the draft to my traffic", "adaptive speculative decoding", "fireoptimizer equivalent", "profile traffic for a draft model", or any combination of "profile / characterize / sample" with "workload / traffic / requests" and "spec-decode / draft / acceptance / hit-rate".

Zur Installation springen

Quellinformationen

Repository
cfregly/gpu-perf-tune
Letzte Quellaktivität
14. Juni 2026 um 03:33
Erkannte Sprache von SKILL.md
Englisch
Sterne
1
Forks
0

Installationsoptionen

Standardmäßig ist der Prompt ausgewählt, der zuerst die Quelle prüft. Sie können zu einem direkten Befehl wechseln oder eine lokale Kopie herunterladen.

Quelldateien prüfen

Lesen Sie SKILL.md und alle von SkillsMP angezeigten Begleitdateien, bevor Sie sich für eine Installation entscheiden.