Skip to main content

inference-fleet-leaderboard

Render cross-model fleet leaderboards from the local perf-report campaigns in ONE command (`perftunereport fleet_leaderboard`): a latency tier (aa-1k/10k/100k tok/s/user + TTFT + cost at c=1/c=10), a throughput tier (peak tok/s/GPU per (model,quant,TP) with latency + $/1M-tok), and a decision capstone (the perf Pareto frontier across decode-latency x first-token x throughput, with the perf-dominated set and a hard PERF != quality caveat). Auto-discovers every model's AA + roofline cells in campaigns/*/atlas.jsonl, so re-running refreshes as new campaigns publish. Use to answer "which model do I pick" / "rank the fleet on latency/throughput/cost" / "is model X perf-dominated". Triggers on "fleet leaderboard", "which model should I use", "cross-model comparison", "rank the fleet", "model selection guide", "pareto frontier of models", "perftunereport fleet_leaderboard", or any combination of "fleet / cross-model / which-model / leaderboard / pareto" with "latency / throughput / cost / pick / rank / compare".

설치로 이동

소스 정보

저장소
cfregly/gpu-perf-tune
최근 소스 활동
2026년 6월 14일 03:33
감지된 SKILL.md 언어
영어
스타
1
포크
0

설치 방법

기본적으로 소스를 먼저 확인하는 Prompt가 선택됩니다. 직접 명령으로 전환하거나 로컬 사본을 다운로드할 수도 있습니다.

소스 파일 검토

설치 여부를 결정하기 전에 SKILL.md와 SkillsMP에 표시된 보조 파일을 읽어 보세요.