Skip to main content

ai-inference-llm-roi-analysis

Stars12
Forks13
UpdatedJuly 14, 2026 at 08:28

Use when a user asks why a self-hosted LLM's ROI is dropping, why its inference cost is rising, whether a model's serving template / resource pool is right-sized, or whether the deployment mix across several self-hosted models should change (scale down / scale up / reprice / shift traffic). Covers single-model cost-decline attribution and portfolio-level deployment-mix ROI. Also use for Chinese requests like 自部署模型 ROI 为什么下降、成本为什么涨、 当前副本/资源池配置是否合理、要不要缩容/扩容/调价、deepseek-v4-pro 和 GLM-5.1 的 部署比例要不要调整、模型组合整体 ROI 怎么样、哪些模型该承接流量. Triggers on model names like DeepSeek, GLM, Qwen, MiniMax, Kimi and terms gross margin, GPU utilization, 毛利率, 利用率, 单位成本, 缩容护栏.

Installation

Install with Codex or Claude Copy this prompt, paste it into Codex, Claude, or another assistant, and let it review the skill page and install it for you.

File Explorer
3 files
SKILL.md
readonly