一键导入
evaluation-and-monitoring
Use when measuring coding-agent quality, regression risk, latency, cost, reliability, safety, drift, or production performance.
用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
菜单
Use when measuring coding-agent quality, regression risk, latency, cost, reliability, safety, drift, or production performance.
用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
基于 SOC 职业分类
Use when reviewing code before merge to assess correctness, tests, maintainability, security and user impact.
Use when reviewing plans, recommendations, prioritisation, risk registers, retrospectives or agent outputs for judgement distortion from cognitive bias — not for checking whether an argument's premises validly support its conclusion.
Use when evaluating whether premises in plans, ADRs, reviews, security justifications or agent recommendations validly support their conclusions — for logical fallacies, not judgement distortion from cognitive bias.
Use when a coding agent needs robust problem solving, explicit decomposition, alternative-path exploration, program-aided reasoning, ReAct loops, or self-correction.
Use when generated code, plans, tests, prompts or architecture need critique, repair and verification before completion; especially for coding tasks requiring self-review, test-driven repair or quality gates.
Identifies risks, scores likelihood and impact, selects treatment, assigns owners and review cadence. Use when building or updating a risk register, treating agent autonomy risks, or deciding mitigate, transfer, accept or avoid.
| name | evaluation-and-monitoring |
| description | Use when measuring coding-agent quality, regression risk, latency, cost, reliability, safety, drift, or production performance. |
Use this skill when an agent or generated system must be measured objectively before, during or after deployment. Apply it to coding-agent changes, RAG systems, tool-use workflows, multi-agent systems and production assistants.
Measure agent or system quality with explicit metrics, representative scenarios, and actionable monitoring evidence.
For coding agents, treat tests as first-class evaluations. Add targeted regression tests for every bug fix. For RAG, measure retrieval recall, answer faithfulness and citation accuracy. For tool-use agents, measure tool-call correctness and error recovery.