daocloud-skills
daocloud-skills에는 DaoCloud에서 수집한 skills 13개가 있으며, 저장소 수준 직업 범위와 사이트 내 skill 상세 페이지를 제공합니다.
이 저장소의 skills
Use when operating the dce generated CLI. Discover commands, inspect parameters, check auth state, and execute API operations safely.
Generate a Chinese daily GPU operations snapshot for system resource utilization, idle headroom, and the trade-off between higher utilization and elastic buffer. Use for requests containing 今日系统运维摘要, 系统资源利用率, 空闲余量, GPU利用率, 资源利用效率, or 弹性缓冲. Query only the Crane API endpoint /apis/crane.io/v1alpha1/singlepage/gpu-resource-status through its DCE command.
Query Crane for model cost and revenue over the most recent N complete calendar days and produce a chart-first operations summary. Use for model cost, model revenue, model gross profit, and recent-N-day model finance requests. Default N is 3. Chinese trigger keywords: 模型成本、模型收入、模型毛利、最近N天模型经营数据. Use only Crane-backed business-cockpit queries; never call other DCE modules.
Query recent Token costs for users or tenants through Crane, rank the top M entries by cost, and produce an operational report with a chart. Trigger this skill for requests containing Chinese keywords such as “Token 费用排名”, “费用 Top 用户”, or “最近几天费用最高用户”. N defaults to 3 and M defaults to 5. This skill only ranks costs; it does not split results by model or perform write operations.
Generate a concise leadership-facing AI operations daily summary and business-value analysis from DaoCloud Enterprise DCE / LLM Studio / Hydra data. Use when the user asks for today's AI operations summary, AI usage report, LLM Studio operating metrics, boss/leadership AI daily report, token/API key/model service overview, business value, operating value, risk identification, or wants available DCE CLI data turned into the most important conclusions, especially in table form.
Diagnose OpenClaw request failures or slow requests with DCE Insight tracing, alerts, logs, and pod status. Use when the user asks for OpenClaw request root-cause analysis, R.E.D analysis, error-chain details, or mitigation advice.
Use when a user asks why a self-hosted LLM's ROI is dropping, why its inference cost is rising, whether a model's serving template / resource pool is right-sized, or whether the deployment mix across several self-hosted models should change (scale down / scale up / reprice / shift traffic). Covers single-model cost-decline attribution and portfolio-level deployment-mix ROI. Also use for Chinese requests like 自部署模型 ROI 为什么下降、成本为什么涨、 当前副本/资源池配置是否合理、要不要缩容/扩容/调价、deepseek-v4-pro 和 GLM-5.1 的 部署比例要不要调整、模型组合整体 ROI 怎么样、哪些模型该承接流量. Triggers on model names like DeepSeek, GLM, Qwen, MiniMax, Kimi and terms gross margin, GPU utilization, 毛利率, 利用率, 单位成本, 缩容护栏.
Attribute why LLM gross margin got worse using live DCE and Crane data. Use when the user asks whether today's/this week's LLM gross margin decline is caused by model cost, tenant/customer mix, cache hit-rate changes, token usage, billing, or asks for ranked impact attribution such as "今天毛利变差是模型成本、租户结构还是缓存命中率变化导致的?按影响大小排序". Requires real DCE queries; do not answer with a generic framework.
Analyze whether GPU resources in a DCE environment have converted into real customer usage and billable consumption. Use when a user asks "当前环境的 GPU 资源有没有真正转化成客户用量", "哪些 GPU 只是有资源没客户用", "GPU 算力有没有变现/产生账单", "GPU 资源利用和客户 Token 用量能不能闭合", "GPU customer usage conversion", "GPU monetization", or similar questions about connecting GPU supply to workspaces/customers, LLM/API usage, runtime activity, and Billing Center revenue evidence.
Use when a user asks to plan AI inference token capacity, token budget, supply plans, SLA commitments, QoS targets, or cost boundaries. Covers token modeling, GPU resource mapping, phased rollout, QoS/SLA design, availability, latency, throughput, quota allocation, and pricing strategy. Also use for Chinese requests like 推理服务 SLA 怎么设计、Token 容量规划、年度 Token 预算、 供给计划、成本边界、可用性/延迟/吞吐承诺、SLA 违约补偿. Example prompts include "18 billion Tokens next year", "how to split N tokens into supply/cost/SLA", "AI inference budget planning", and "token quota allocation".
Use when a user asks to inspect, health-check, patrol/巡检, diagnose, or troubleshoot a Kubernetes cluster managed by the DCE/kpanda module. Covers cluster health overview, node status, abnormal Pods, events, cluster unavailability, node NotReady, pending or failed Pods, and Chinese requests like 集群巡检、集群体检、检查集群健康状态、排查集群异常、查看集群状态.
Use when a user asks to analyze GPU pool or GPU cluster capacity, predict bottlenecks under traffic, QPS, or inference load growth, or choose between scaling, throttling, and routing changes. Also use for Chinese requests about GPU 池/集群容量、流量上涨、QPS 增长、推理负载增长、哪个 GPU 池先成为瓶颈、是否扩容/限流/改路由. Example prompts include "如果今晚流量上涨 30%,哪个 GPU 池会先成为瓶颈?", "建议扩容、限流还是改路由?", "which GPU pool will bottleneck first", and "what if traffic increases by X%".
Use when a user asks to diagnose, inspect, troubleshoot, or find the root cause of a specific Pod failure, non-running Pod, or pod-level issue in a Kubernetes cluster managed by the DCE/kpanda module. Also use for Chinese requests like 排查 Pod 故障原因、某个 pod 异常、pod 启动失败、pod 无法运行、 pod 一直重启、CrashLoopBackOff、OOMKilled、ImagePullBackOff、Pending、Evicted、 or Terminating.