Skip to main content

inference-spec-decode-tune

Tune a speculative-decoding DRAFT head's TRAINING hyperparameters (global batch, learning rate, accumulation, warmup) for a target LLM, optimizing the TRUE serving in-engine acceptance length (vLLM spec_decode counters) with a cheap training-acc proxy for triage. Reuses the search ALGORITHMS (hyperband/grid/random natively. Bayesian-TPE via optuna if installed) wired to a GB300/managed K8s pod launcher + an offline EAGLE3/DFlash trainer (distinct from the MLPerf ai_tuning contract, which does not map to draft training). The draft-training analog of inference-tune-sweep (which tunes vLLM SERVING config). Triggers on "tune the draft head", "tune eagle3 / dflash training", "bayesian/hyperband tune the draft", "search global batch and LR for the speculator", "spec-decode hyperparameter sweep", "draft-training tuner", or any combination of "tune / sweep / optimize / search / bayesian / hyperband" with "eagle3 / dflash / draft / speculator / spec-decode".

跳到安装

来源信息

仓库
cfregly/gpu-perf-tune
最近来源活动
2026年6月14日 03:33
检测到的 SKILL.md 语言
英语
星标
1
分支
0

安装方式

默认使用会先检查来源的 Prompt;你也可以切换为直接命令,或下载本地副本。

检查来源文件

决定是否安装前,请先阅读 SKILL.md,以及 SkillsMP 当前展示的配套文件。