Skip to main content

inference-optimizer:自托管 LLM 推理测量与调优

为指定模型、硬件和负载规划部署并调优开放权重 LLM 推理。流程先建立实测基线,再通过接口、质量、内存和延迟检查比较配置。

来源信息

仓库
abhiram1809/inference-god-mode
最近来源活动
2026年10月2日 21:47
检测到的 SKILL.md 语言
英语
星标
0
分支
0

必要前提

提供或允许 Agent 检查模型仓库与修订版、GPU 和主机容量、峰值同时生成数、延迟目标、所需 API 能力,以及质量和计算预算限制。复制 Skill 时保留完整 inference-optimizer 目录和其中的 references/。

操作说明

将 README 请求中的占位内容换成自己的配置:

Use inference-optimizer to serve <model repository> on <GPU and memory>.
I need <API dialects>, <peak active generations>, and <latency target>.
Keep the full native context and evaluate 4-bit weights and FP8 KV.
Benchmark the actual workload and report the best verified configuration.

流程会在真实负载上比较基线和候选配置,并保留通过正确性与质量检查的结果。

限制

来源明确排除 Ollama。4-bit 权重与 FP8 KV 是待评估的候选方案,不保证更快或达到质量要求。作者公布的验证覆盖 Skill 结构、引用链接和宿主软链接;部署性能仍须针对所选模型、硬件与运行时测量。

安装方式

默认使用会先检查来源的 Prompt;你也可以切换为直接命令,或下载本地副本。

检查来源文件

决定是否安装前,请先阅读 SKILL.md,以及 SkillsMP 当前展示的配套文件。

文件资源管理器
13 个文件

正在显示 SKILL.md

SKILL.md
来源说明 · 只读预览
name
inference-optimizer
description
Plan, deploy, benchmark, and tune self-hosted open-weight LLM inference on one machine or Kubernetes GPU clusters, including quantization, full-context capacity, OpenAI or Anthropic APIs, cache-aware routing, and prefill/decode disaggregation.
# Inference optimizer Optimize the user's *specific* model, hardware, workload, and client protocol. Work from a measured baseline to a verified best configuration; do not present an untested flag list as an optimization result. ## Establish the target Collect or inspect these facts, then ask only for missing decisions that affect the route: - Exact model repository and revision, component architectures and execution schedule, license/gating, tokenizer and chat template, modalities, reasoning and tool-use requirements. Inspect configuration and the actual model/runtime implementation; classify dense/routed layers, lookup tables, encoders, decoders, recurrent state, and iterative generators before estimating cost. - GPU model, count, memory capacity and bandwidth per GPU, SM/compute-unit and tensor-core resources, clock/power limits, supported dense arithmetic rates by precision, compute capability, interconnect/topology, driver/runtime, host RAM and bandwidth, storage, operating system, and whether the machine is shared. Use read-only hardware commands where available rather than asking for facts the machine can report; verify peak specifications against the exact device's vendor documentation. - Peak simultaneous *active generations*, arrival rate or requests per minute, typical and worst-case input/output lengths, latency target (TTFT and inter-token latency), and throughput target. Number of registered users is not concurrency. Ask for the minimum context window the user would accept **if** the full native window proves infeasible. - Required wire protocols and client features: OpenAI Chat Completions, OpenAI Responses, Anthropic Messages, streaming, tools, structured output, reasoning fields, image/audio support, token counting, and authentication. - Quality tolerance, time and compute budget for local quantization, and whether weights may be downloaded or published. Ask whether preserving reusable KV prefixes in host RAM or SSD is desirable when the workload has repeated prefixes, and capture the available RAM/SSD budget. Read [architecture-specific cost models](references/architecture-cost-models.md) before sizing or calculating throughput. Architecture features can coexist; derive work, traffic, persistent state, and execution frequency per component, then combine them using the real execution dependencies and shared resources. The dense decoder formulas are conditional examples, not universal formulas. Use task-appropriate units for TTS and other pipelines, and apply context/KV assumptions only to components that implement them. Verify that the selected engine supports the complete architecture before preparing a deployment. If workload is vague, provisionally assume one active user and prioritize interactive latency. Default to a **4-bit weight candidate** unless the user chooses another precision; this is a search starting point, not a claim that 4-bit is fastest or accurate enough. For autoregressive text components, assume the user wants the model's full *native* context window, including output-token headroom. Never silently shorten context or apply RoPE scaling to claim a larger window. If full context cannot fit at the requested concurrency, calculate the gap and ask which constraint may change: concurrency, GPU capacity, KV precision, model, or context. Continue independent research while waiting. Prefer **FP8 KV cache** as the first optimized candidate where the model, device, and attention backend support it. Keep a BF16/FP16 KV control and promote FP8 after model-specific quality, full-context, and performance checks. Weight precision and KV precision are independent. Read [FP8 KV selection](references/fp8-kv-cache.md) for kernel fixes, scales, and exceptions. ## Select candidates Read [backend routing](references/backend-routing.md) for the relevant device and [quantization and capacity](references/quantization-capacity.md) before choosing weights. Refresh upstream compatibility documentation, release notes, and the live Hugging Face Hub for the exact model; version-specific support and checkpoint quality change often. When upstream lacks a required model/GPU/quantization combination, search official model recipes, merged changes/nightlies, plugins, and community images or forks. Read [community builds](references/community-builds.md) before recommending or running one; pin its source and image, establish what it changes, and verify its actual kernel path. An EXL3 plugin can be evaluated within vLLM without changing the allowed engine set. - Legacy GPUs: evaluate **llama.cpp** first when current production engines or their required kernels no longer support the device. For older AMD cards such as Radeon RX 560X, prefer a **Vulkan** candidate; verify the exact GPU and driver capabilities, actual offload, and workload fit. Apply this compatibility route before the workload defaults below. Read the legacy section of [backend routing](references/backend-routing.md). - Supported NVIDIA GPUs, multiple concurrent users: start with **vLLM**. Supported NVIDIA GPUs, one active user: start with **SGLang**, while keeping vLLM as the comparator. A hard requirement for all three API dialects may favor a version of vLLM that implements and passes each required route. These are starting hypotheses; benchmark alternatives on the actual workload. - Consider **TensorRT-LLM** when the model architecture, checkpoint, GPU, and current serving API are supported and an engine build is justified by measured gains. Consider **llama.cpp** for GGUF checkpoints, constrained memory, unusual hardware, or a measured single-user advantage. Do not route to Ollama. - Apple Silicon: consider **MLX-LM** or **vLLM-Metal** with verified MLX 4-bit checkpoints; consider llama.cpp/GGUF if API coverage or model support calls for it. Do not assume MLX-LM's basic HTTP server supplies all requested API dialects. - A checkpoint's label is insufficient: verify its quantization metadata, architecture, tokenizer, license, producer, calibration method where available, backend loader, and **actual launched kernel and arithmetic**. On Blackwell, evaluate NVFP4 first **when a genuine compatible checkpoint and kernel exist**. On Ampere/Ada/Hopper, evaluate AWQ/GPTQ with the available optimized kernels, plus other supported 4-bit candidates. NF4 is not interchangeable with AWQ/GPTQ and does not imply fast serving. Distinguish 4-bit storage from 4-bit activations and math; a successful load does not prove a native FP4 path. Treat each required endpoint and behavior as a tested contract. If the preferred engine lacks an endpoint, use another allowed engine or a clearly identified adapter only if that adapter passes end-to-end tests; never describe partial compatibility as full compatibility. Read [API verification](references/api-verification.md). ## Execute the optimization loop Read [optimization loop](references/optimization-loop.md) for sizing, benchmark design, tuning order, and kernel work. In brief: 1. Inventory hardware and exact model. Search all plausible Hugging Face quantized variants; shortlist by actual format, backend support, provenance, and fit. Pin repository revisions and engine versions. 2. Estimate weight, KV-cache, runtime, graph/workspace, and concurrent-request memory **for the chosen engine and kernel path** at the full native context. A GGUF file size or llama.cpp memory result does not establish vLLM/SGLang/TensorRT-LLM fit. Verify with a real load and full-context request; retain OOM headroom. Calculate the [hardware throughput ceilings](references/hardware-throughput-ceilings.md) for each feasible quantization, context, batch, and parallel layout before comparing performance. 3. Run the simplest correct configuration. Exercise every required API feature and a small quality set against an appropriate higher-precision reference. 4. Benchmark representative short, medium, full-context where applicable, single-stream, and target-concurrency cases. Record task-appropriate throughput and latency (TTFT/inter-token latency/tokens/s for text; time to first audio, real-time factor, and generated audio seconds/s for TTS), p50/p95 and sufficiently sampled p99 latency, error rate, peak memory, power if relevant, versions, and exact commands. Report measured throughput as a fraction of its matching theoretical ceiling and calibrated estimate, with the predicted bottleneck and remaining gap. 5. Change one performance dimension at a time: weight/kernel choice, FP8 KV, scheduling/batching and CPU/GPU overlap, prefill, graph capture, prefix caching, parallelism, then [speculative decoding and KV offload](references/speculation-kv-offload.md) where applicable. Search for a compatible native or attached MTP head, EAGLE/other drafter, or DFlash checkpoint for the exact model and engine. Consider CPU/SSD KV tiers only for measured reusable-prefix benefit. Retest correctness and quality after each change. Keep the best Pareto choices for latency, throughput, memory, and quality. 6. If no suitable 4-bit checkpoint exists, propose or execute a local quantization path with calibration and evaluation. If profiling identifies a real kernel bottleneck and the user wants further engineering, use a pinned fresh upstream checkout, make a narrow reproducible change, and compare against the baseline. Do not claim a speedup from synthetic microbenchmarks alone. Use an **architecture-specific roofline estimate** to make performance targets hardware-aware: combine memory bandwidth and bytes moved with the operation count and supported compute rate of each actual component/kernel path. Read [hardware throughput ceilings](references/hardware-throughput-ceilings.md) for formulas, quantization comparisons, calibration, and acceptance rules. Distinguish the optimistic theoretical envelope from an attainable estimate; core count alone or a checkpoint's bit width cannot establish throughput. Identify which configuration approaches its ceiling most closely, and choose the measured winner for the user's objective among configurations that pass quality, API, context, and latency requirements. A higher percentage of a lower ceiling does not necessarily mean a faster configuration. For multi-GPU setups, read [parallelism experiments](references/parallelism-experiments.md) and evaluate supported layouts **sequentially: TP → PP → DP → EP**, with EP only for supported MoE models. Inspect dense versus MoE architecture, per-GPU fit, topology, and valid combinations before generating commands. Present the experiment matrix and ask for missing run-budget or service-interruption constraints; honor existing authorization. Keep a TP=1 control when it fits, but do not settle on it without evaluating useful multi-GPU layouts. Measure real serving traffic before selecting the final configuration; a canned throughput score or successful startup is preliminary evidence. Do not download large checkpoints, run long calibration, rebuild engines, expose a service beyond localhost, or publish weights as an incidental step. Make those costs and side effects explicit before taking them on; follow the user's existing authorization. ## Organizational and multi-machine deployment For organizational fleet serving or Kubernetes deployment, read [cluster serving](references/cluster-serving.md) after establishing a correct engine baseline. Evaluate **llm-d** as an orchestration and routing layer above the allowed engines. Compare replicas with ordinary routing, cache/load-aware routing, and then separate prefill/decode pools when measured interference justifies them. Size both stages from the workload and network measurements; preserve full-context, API, quality, and availability requirements through the complete gateway path. Prepare reproducible manifests or Helm values and an end-to-end benchmark before claiming a production configuration is optimized. ## When startup or inference fails Read [issue-first troubleshooting](references/startup-troubleshooting.md). Capture the root error and exact environment, then search matching upstream GitHub issues and their linked fixes before trying speculative argument changes. Prefer maintainer explanations, reproducible reports, and current code/docs over an LLM's unsupported diagnosis. Preserve the intended model, quantization, context, concurrency, and API contract; distinguish temporary diagnostic settings from the final configuration and retest after fixing the cause. ## Deliverable Give the user a runnable configuration or a concrete next experiment, exact model/checkpoint/revision, backend/version, image digest and fork/plugin commits if used, kernel libraries and selected paths, tested API routes, applicable context and concurrency limits, benchmark table, quality observations, and remaining bottleneck. Include the architecture/component classification and its source evidence, hardware bandwidth/compute sources, per-component and end-to-end theoretical/calibrated estimates in task-appropriate units, measured fractions of those estimates, assumptions/uncertainty, and the configuration closest to its ceiling alongside the measured objective winner. For multi-GPU runs, include the tested TP/PP/DP/EP layouts, rank placement, per-GPU memory, skipped/failed candidates, and the reason for choosing the winner. Mark estimates and untested suggestions as such. If no configuration meets the hard constraints, show the binding constraint and smallest viable alternatives.
在 GitHub 查看