一键导入
skill-evaluator
Use to evaluate agent skills with should-trigger and should-not-trigger cases, executable validation hooks, and pass/fail reports for review or CI.
用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
菜单
Use to evaluate agent skills with should-trigger and should-not-trigger cases, executable validation hooks, and pass/fail reports for review or CI.
用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
基于 SOC 职业分类
Use for complex software tasks that need interactive main-session planning followed by structured checklist execution.
Use to validate agent skill structure, frontmatter, path/name consistency, required sections, and broken references before skill changes are accepted.
Use to audit the skill ecosystem for capability coverage, overlapping responsibilities, fragmented guidance, and consolidation opportunities.
Use when the user asks to create, improve, copy, port, evaluate, or brainstorm an agent skill. Probes the desired workflow, researches local and upstream skill examples, recommends patterns, then invokes plan-and-execute for implementation.
Trigger keywords: "grill me", "grill-me". Use only when the user explicitly requests grill-me style relentless, one-question-at-a-time interviewing.
Use to pressure-test implementation plans against codebase docs, refine language, capture ADR decisions, and emit machine-readable handoff JSON for downstream workflows.
| name | skill-evaluator |
| description | Use to evaluate agent skills with should-trigger and should-not-trigger cases, executable validation hooks, and pass/fail reports for review or CI. |
skill-validator passes when both static and behavioral validation are needed.skill-validator instead.SKILL.md and extract its description, When to use, non-trigger guidance, validation requirements, and output contract.trigger, do-not-trigger, or ambiguous.When to use guidance.overall_status, target_skill, trigger_cases, executable_checks, and recommendations.overall_status: pass only when trigger cases and executable checks meet expectations.overall_status: needs changes when wording, tests, or commands need revision.