Calculate incident impact, detection and recovery durations, classify incidents as L0-L3 from service criticality and impact inputs, compute weighted SLA impact and derived availability, and generate post-incident Markdown reports. Use for SLA calculations,…
云原生 SRE 值班助理和巡检编排技能。用于基础设施早检、Kubernetes 集群健康检查、Ceph 存储巡检、计算集群巡检、网关故障初筛、异常证据汇总和值班日报生成。触发场景:用户要求早检、巡检、检查集群是否正常、检查 Ceph、检查计算集群、排查网关错误或生成基础设施健康报告。
Create, review, tune, and validate production-grade Prometheus alerting rules and Kubernetes PrometheusRule resources. Use when writing or repairing alert rule YAML, designing PromQL alert expressions, recommending defensible thresholds and for durations,…
Guide agents to use Alibaba Cloud CLI correctly and safely. Use when a user asks to run or explain aliyun CLI, Alibaba Cloud CLI, Aliyun CLI, 阿里云 CLI, cloud resource inspection, profile/credential setup, ECS/VPC/SLB/RDS/ACK/OSS/RAM or other Alibaba Cloud…
Ceph 分布式存储集群状态分析技能。获取和分析 Rook-Ceph 集群的健康状态、存储利用率、OSD 性能等。 触发场景:Ceph 集群状态检查、性能问题排查、容量规划、OSD/PG/Monitor/MDS 故障诊断。 支持用户纠正和补充。
多云资源使用率和成本优化分析技能。用于结合 hcloud/KooCLI、aliyun CLI、Prometheus 和账单数据生成华为云与阿里云资源使用情况报告,拉取华为云 ECS/EVS/RDS/DCS/OBS 配置、阿里云 ResourceCenter 全账号资源配置,整合云资源规格、监控使用率和费用,识别闲置、低利用、高风险和可降配资源。适用于:云资源盘点、华为云/阿里云跨云报告、ECS/磁盘/RDS/Redis/SLB/ALB/NLB/PolarDB/Flink/SelectDB/Kafka/OSS/NAT…
Use an installed devopsctl CLI to operate an enterprise DevOps platform. Trigger when the user asks an agent to inspect or manage business spaces, apps, services, deployments, Jenkins publish/build history, login/auth status, or supported DevOps backend API…
Generate editable draw.io diagrams (.drawio, .drawio.svg, .drawio.png) from text, images, Excel, or Markdown. Use for architecture diagrams, flowcharts, sequence diagrams, swimlanes, and Azure/AWS cloud diagrams.