一键导入
skill-test
Validate skill files for structural compliance and behavioral correctness. Three modes: static (linter), spec (behavioral), audit (coverage report).
用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
菜单
Validate skill files for structural compliance and behavioral correctness. Three modes: static (linter), spec (behavioral), audit (coverage report).
用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
基于 SOC 职业分类
Brownfield onboarding — audits existing project artifacts for template format compliance (not just existence), classifies gaps by impact, and produces a numbered migration plan. Run this when joining an in-progress project or upgrading from an older template version. Distinct from /project-stage-detect (which checks what exists) — this checks whether what exists will actually work with the template's skills.
Generate a project progress dashboard from workflow-catalog.yaml. Reads current stage, required steps, artifact evidence, validation gaps, writes production/project-roadmap.md after approval, and mirrors to memory_bank/t2_execution/current_roadmap.md when memory_bank exists.
Memory-bank governance audit — checks whether T0 laws/current state, T1 supporting context, T2 execution mirrors, adapter freshness, and T3 archive indexes align. Read-only by default; may record adapter_state.yaml only after explicit approval.
Constitution Driven Development project governance — establishes, derives, updates, or amends governing principles at any project stage. Reads existing artifacts to derive a constitution, audits alignment, tracks versions, and supports formal amendment workflow. Domain-agnostic, stage-aware unified onboarding entry.
Analyzes what is done and the users query and offers advice on what to do next. Use if user says what should I do next or what do I do now or I'm stuck or I don't know what to do
Generates a contextual onboarding document for a new contributor or agent joining the project. Summarizes project state, architecture, conventions, and current priorities relevant to the specified role or area.
| name | skill-test |
| description | Validate skill files for structural compliance and behavioral correctness. Three modes: static (linter), spec (behavioral), audit (coverage report). |
| argument-hint | static [skill-name | all] | spec [skill-name] | category [skill-name | all] | audit |
| user-invocable | true |
| allowed-tools | Read, Glob, Grep, Write |
/skill-test static [skill-name | all] | spec [skill-name] | category [skill-name | all] | audit; project artifacts referenced below; user decisions and approvals before writes.memory_bank/t3_archive/skill_testing/results/ and update memory_bank/t3_archive/skill_testing/coverage-index.yaml.skill_testing/; with approval writes T3 test evidence under memory_bank/t3_archive/skill_testing/.Detect the skill domain before testing:
A passing skill test must not require deleting game-specific examples.
Validates skills/*/SKILL.md files for structural compliance and
behavioral correctness. No external dependencies — runs entirely within the
existing skill/hook/template architecture.
Four modes:
| Mode | Command | Purpose | Token Cost |
|---|---|---|---|
static | /skill-test static [name|all] | Structural linter — 7 compliance checks per skill | Low (~1k/skill) |
spec | /skill-test spec [name] | Behavioral verifier — evaluates assertions in test spec | Medium (~5k/skill) |
category | /skill-test category [name|all] | Category rubric — checks skill against its category-specific metrics | Low (~2k/skill) |
audit | /skill-test audit | Coverage report — skills, agent specs, last test dates | Low (~3k total) |
Determine mode from the first argument:
static [name] → run 8 structural/parity checks on one skillstatic all → run 8 structural/parity checks on all skills (Glob skills/*/SKILL.md)spec [name] → read skill + test spec, evaluate assertionscategory [name] → run category-specific rubric from skill_testing/quality-rubric.mdcategory all → run category rubric for every skill that has a category: in catalogaudit (or no argument) → read catalog, list all skills and agents, show coverageIf argument is missing or unrecognized, output usage and stop.
For each skill being tested, read its SKILL.md fully and run all 8 checks:
The file must contain all of these in the YAML frontmatter block:
name:description:argument-hint:user-invocable:allowed-tools:FAIL if any are absent.
The skill must have ≥2 numbered phase headings. Look for patterns like:
## Phase N or ## Phase N:## N. (numbered top-level sections)## headings if phases aren't explicitly numberedFAIL if fewer than 2 phase-like headings are found.
The skill must contain at least one of: PASS, FAIL, CONCERNS, APPROVED,
BLOCKED, COMPLETE, READY, COMPLIANT, NON-COMPLIANT
FAIL if none are present.
The skill must contain ask-before-write language. Look for:
"May I write" (canonical form)"before writing" or "approval" near file-write instructions"ask" + "write" in close proximity (within same section)WARN if absent (some read-only skills legitimately skip this).
FAIL if allowed-tools includes Write or Edit but no ask-before-write language is found.
The skill must end with a recommended next action or follow-up path. Look for:
/story-done, /gate-check)WARN if absent.
If frontmatter contains context: fork, the skill should have ≥5 phase headings
(## level or numbered Phase N headers). Fork context is for complex multi-phase
skills; simple skills should not use it.
WARN if context: fork is set but fewer than 5 phases found.
argument-hint must be non-empty. If the skill body mentions multiple modes
(e.g., "Mode A | Mode B"), the hint should reflect them. Cross-reference the
hint against the first phase's "Parse Arguments" section.
WARN if hint is "" or if documented modes don't match hint.
If the skill advertises Product support in frontmatter or Phase 0, it must also contain Product-specific implementation guidance beyond the first 50 lines. Look for at least two of:
product-concept.md, Product CDDs, docs/reference/<stack>/, language specialist, API/CLI/UI/data docs)WARN if Product appears only near the top of the file. FAIL if the skill claims Product support but the body remains game-only and would route a Product project through engine/player/playtest/HUD-only behavior.
Also check that game content is preserved:
For a single skill:
=== Skill Static Check: /[name] ===
Check 1 — Frontmatter Fields: PASS
Check 2 — Multiple Phases: PASS (7 phases found)
Check 3 — Verdict Keywords: PASS (PASS, FAIL, CONCERNS)
Check 4 — Collaborative Protocol: PASS ("May I write" found)
Check 5 — Next-Step Handoff: WARN (no follow-up section found)
Check 6 — Fork Context Complexity: PASS (8 phases, context: fork set)
Check 7 — Argument Hint: PASS
Verdict: WARNINGS (1 warning, 0 failures)
Recommended: Add a "Follow-Up Actions" section at the end of the skill.
For static all, produce a summary table then list any non-compliant skills:
=== Skill Static Check: All 74 Skills ===
Skill | Result | Issues
-----------------------|--------------|-------
gate-check | COMPLIANT |
design-review | COMPLIANT |
story-readiness | WARNINGS | Check 5: no handoff
...
Summary: 48 COMPLIANT, 3 WARNINGS, 1 NON-COMPLIANT
Aggregate Verdict: N WARNINGS / N FAILURES
Static mode displays results by default. If the user wants to preserve the run, ask:
"May I write this static check to
memory_bank/t3_archive/skill_testing/results/static/skill-test-static-[name|all]-[YYYY-MM-DD].md
and update memory_bank/t3_archive/skill_testing/coverage-index.yaml?"
If yes:
memory_bank/t3_archive/skill_testing/results/static/coverage-index.yaml:
last_static: [date]last_static_result: PASS|WARN|FAILlatest_result_path: memory_bank/t3_archive/skill_testing/results/static/skill-test-static-[name|all]-[YYYY-MM-DD].mdFind skill at skills/[name]/SKILL.md.
Look up the spec path from skill_testing/catalog.yaml
— use the spec: field for the matching skill entry.
If either is missing:
skills/."/skill-test audit
to see coverage gaps."Read the skill file and test spec file completely.
For each Test Case in the spec:
For each assertion, evaluate whether the skill's written instructions, if followed correctly given the fixture state, would satisfy it. This is a Claude-evaluated reasoning check, not code execution.
Mark each assertion:
For Protocol Compliance assertions (always present):
=== Skill Spec Test: /[name] ===
Date: [date]
Spec: skill_testing/specs/skills/[category]/[name].md
Case 1: [Happy Path — name]
Fixture: [summary]
Assertions:
[PASS] [assertion text]
[FAIL] [assertion text]
Reason: The skill's Phase 3 says "..." but the fixture state means "..."
Case Verdict: FAIL
Case 2: [Edge Case — name]
...
Case Verdict: PASS
Protocol Compliance:
[PASS] Uses "May I write" before file writes
[PASS] Presents findings before asking approval
[WARN] No explicit next-step handoff at end
Overall Verdict: FAIL (1 case failed, 1 warning)
For any spec that covers a dual-domain skill, evaluate both branches:
design/cdd/game-concept.md exists. The skill should preserve
game terminology, paths, examples, and next steps.design/cdd/product-concept.md exists. The skill should use
Product terminology, paths, examples, and next steps.Mark assertions:
"May I write these results to
memory_bank/t3_archive/skill_testing/results/spec/skill-test-spec-[name]-[YYYY-MM-DD].md
and update memory_bank/t3_archive/skill_testing/coverage-index.yaml?"
If yes:
memory_bank/t3_archive/skill_testing/results/spec/memory_bank/t3_archive/skill_testing/coverage-index.yaml:
last_spec: [date]last_spec_result: PASS|PARTIAL|FAILlatest_result_path: memory_bank/t3_archive/skill_testing/results/spec/skill-test-spec-[name]-[YYYY-MM-DD].mdFind skill at skills/[name]/SKILL.md.
Look up category: field in skill_testing/catalog.yaml.
If skill not found: "Skill '[name]' not found."
If no category: field: "No category assigned for '[name]' in catalog.yaml.
Add category: [name] to the skill entry first."
For category all: collect all skills with a category: field and process each.
category: utility skills are evaluated against U1 (static checks pass) and U2
(gate mode correct if applicable) only — skip to the static mode for U1.
Read skill_testing/quality-rubric.md.
Extract the section matching the skill's category (e.g., ### gate, ### team).
Read the skill's SKILL.md fully.
For each metric in the category's rubric table:
=== Skill Category Check: /[name] ([category]) ===
Metric G1 — Review mode read: PASS
Metric G2 — Full mode directors: FAIL
Gap: Phase 3 spawns only CD-PHASE-GATE; TD-PHASE-GATE, PR-PHASE-GATE, AD-PHASE-GATE absent
Metric G3 — Lean mode: PHASE-GATE only: PASS
Metric G4 — Solo mode: no directors: PASS
Metric G5 — No auto-advance: PASS
Verdict: FAIL (1 failure, 0 warnings)
Fix: Add TD-PHASE-GATE, PR-PHASE-GATE, and AD-PHASE-GATE to the full-mode director
panel in Phase 3.
"May I write this category check to
memory_bank/t3_archive/skill_testing/results/category/skill-test-category-[name]-[YYYY-MM-DD].md
and update memory_bank/t3_archive/skill_testing/coverage-index.yaml
(last_category, last_category_result, latest_result_path) for [name]?"
Read skill_testing/catalog.yaml for the registry and
memory_bank/t3_archive/skill_testing/coverage-index.yaml for test history.
If either is missing, note the missing file and recommend running /constitute
to initialize the memory-bank testing templates.
Glob skills/*/SKILL.md to get the complete list of skills.
Extract skill name from each path (directory name).
Also read the agents: section from skill_testing/catalog.yaml to get the
complete list of agents.
For each skill:
spec: path from catalog, or glob skill_testing/specs/skills/*/[name].md)last_static, last_static_result, last_spec, last_spec_result,
last_category, last_category_result, and latest_result_path from
coverage-index.yaml (or mark as "never" / "—" if not in coverage)category from catalogpriority: field (critical/high/medium/low)For each agent in catalog's agents: section:
spec: path from catalog, or glob skill_testing/specs/agents/*/[name].md)last_spec, last_spec_result, and latest_result_path from
coverage-index.yamltier from catalog=== Skill Test Coverage Audit ===
Date: [date]
SKILLS (74 total)
Specs written: 74 (100%) | Never static tested: 74 | Never category tested: 74
Skill | Cat | Has Spec | Last Static | S.Result | Last Cat | D.Result | Priority
-----------------------|----------|----------|-------------|----------|----------|----------|----------
gate-check | gate | YES | never | — | never | — | critical
design-review | review | YES | never | — | never | — | critical
...
AGENTS (53 total)
Agent specs written: 53 (100%)
Agent | Category | Has Spec | Last Spec | Result
-----------------------|------------|----------|-------------|--------
creative-director | director | YES | never | —
technical-director | director | YES | never | —
...
Top 5 Priority Gaps (skills with no spec, critical/high priority):
(none if all specs are written)
Skill coverage: 74/74 specs (100%)
Agent coverage: 53/53 specs (100%)
Audit mode is read-only by default.
Optional write: ask "May I write this audit report to
memory_bank/t3_archive/skill_testing/results/audit/skill-test-audit-[YYYY-MM-DD].md?"
Only write the audit report if the user approves.
Offer: "Would you like to run /skill-test static all to check structural
compliance across all skills? /skill-test category all to run category rubric
checks? Or /skill-test spec [name] to run a specific behavioral test?"
After any mode completes, offer contextual follow-up:
static [name]: "Run /skill-test spec [name] to validate behavioral
correctness if a test spec exists."static all with failures: "Address NON-COMPLIANT skills first. Run
/skill-test static [name] individually for detailed remediation guidance."spec [name] PASS: "Update memory_bank/t3_archive/skill_testing/coverage-index.yaml
to record this pass date. Consider running /skill-test audit to find the next spec gap."spec [name] FAIL: "Review the failing assertions and update the skill
or the test spec to resolve the mismatch."audit: "Start with the critical-priority gaps. Use the spec template at
skill_testing/templates/skill-test-spec.md to create new specs."