Skip to main content
Jeden Skill in Manus ausführen
mit einem Klick

redline

Sterne1
Forks0
Aktualisiert30. Juni 2026 um 21:32

Finds the maximum safe --max-model-len and --max-num-batched-tokens for a vLLM-served model on a fixed GPU, by climbing a doubling ladder one clean reboot at a time and recording the real KV-cache/concurrency signal from each boot rather than guessing. Use when asked to find max context, tune context length, or push a model to its context ceiling.

Installation

Mit Codex oder Claude installieren Kopieren Sie diesen Prompt, fügen Sie ihn in Codex, Claude oder einen anderen Assistant ein und lassen Sie die Skill-Seite prüfen und installieren.

SKILL.md
readonly