Skip to main content

ai-inference-llm-roi-analysis

스타12
포크13
업데이트2026년 7월 14일 08:28

Use when a user asks why a self-hosted LLM's ROI is dropping, why its inference cost is rising, whether a model's serving template / resource pool is right-sized, or whether the deployment mix across several self-hosted models should change (scale down / scale up / reprice / shift traffic). Covers single-model cost-decline attribution and portfolio-level deployment-mix ROI. Also use for Chinese requests like 自部署模型 ROI 为什么下降、成本为什么涨、 当前副本/资源池配置是否合理、要不要缩容/扩容/调价、deepseek-v4-pro 和 GLM-5.1 的 部署比例要不要调整、模型组合整体 ROI 怎么样、哪些模型该承接流量. Triggers on model names like DeepSeek, GLM, Qwen, MiniMax, Kimi and terms gross margin, GPU utilization, 毛利率, 利用率, 单位成本, 缩容护栏.

설치

Codex 또는 Claude로 설치 이 Prompt를 복사해 Codex, Claude 또는 다른 어시스턴트에 붙여 넣으면 Skill 페이지를 검토하고 설치를 진행할 수 있습니다.

파일 탐색기
3 개 파일
SKILL.md
readonly