Skip to main content

inference-spec-decode-tune

Tune a speculative-decoding DRAFT head's TRAINING hyperparameters (global batch, learning rate, accumulation, warmup) for a target LLM, optimizing the TRUE serving in-engine acceptance length (vLLM spec_decode counters) with a cheap training-acc proxy for triage. Reuses the search ALGORITHMS (hyperband/grid/random natively. Bayesian-TPE via optuna if installed) wired to a GB300/managed K8s pod launcher + an offline EAGLE3/DFlash trainer (distinct from the MLPerf ai_tuning contract, which does not map to draft training). The draft-training analog of inference-tune-sweep (which tunes vLLM SERVING config). Triggers on "tune the draft head", "tune eagle3 / dflash training", "bayesian/hyperband tune the draft", "search global batch and LR for the speculator", "spec-decode hyperparameter sweep", "draft-training tuner", or any combination of "tune / sweep / optimize / search / bayesian / hyperband" with "eagle3 / dflash / draft / speculator / spec-decode".

Aller à l'installation

Informations de source

Dépôt
cfregly/gpu-perf-tune
Dernière activité de la source
14 juin 2026 à 03:33
Langue détectée de SKILL.md
anglais
Étoiles
1
Forks
0

Options d'installation

Le prompt qui vérifie d'abord la source est sélectionné par défaut. Vous pouvez passer à une commande directe ou télécharger une copie locale.

Vérifiez les fichiers source

Lisez SKILL.md et les fichiers associés affichés par SkillsMP avant de décider de l'installer.