Skip to main content

inference-model-optimize

End-to-end orchestrator that takes a NEW model from a bare HuggingFace id to a perf-lake-published, validated CROSS-ENGINE (vLLM + SGLang) champion on B200/GB300. Scaffolds a per-model harness + evidence bundle (run-id == experiment-id), then drives gate-driven phases: deploy baseline -> 4-layer profile (zymtrace L1 / DCGM L3 / ncu L4) -> tune CROSS-ENGINE via the variant A/B (vLLM AND SGLang arms) -> quantize -> validate -> spec-decode train+validate -> bench the multi-workload suite -> champion_select (baseline vs top-X, the obvious production pick) -> publish_to_lake. Pauses at every red gate. Every number defaults to DRAFT (a champion VERDICT needs the multi-workload + accuracy gates + L3 byte-grounding). Triggers on "optimize a new model", "bring up a model end-to-end", "model bring-up pipeline", "find the best perf for <model>", "run the full optimization pipeline", or any combination of "optimize / bring-up / end-to-end / pipeline / champion" with "new model / inference / vllm / sglang".

설치로 이동

소스 정보

저장소
cfregly/gpu-perf-tune
최근 소스 활동
2026년 8월 4일 03:35
감지된 SKILL.md 언어
영어
스타
1
포크
0

설치 방법

기본적으로 소스를 먼저 확인하는 Prompt가 선택됩니다. 직접 명령으로 전환하거나 로컬 사본을 다운로드할 수도 있습니다.

소스 파일 검토

설치 여부를 결정하기 전에 SKILL.md와 SkillsMP에 표시된 보조 파일을 읽어 보세요.