Skip to main content

test-skill

Validate a drafted skill — write its trigger fixtures, update the activation contract table, run the static and real-path activation tests, then behaviourally test it in a cold session against a `--safe-mode` baseline. Reads its inputs from disk — brief or shipped frontmatter — so it runs cold. Encodes the traps that make a broken test look like a broken glob: activation is keyed to the `Read` tool so a Bash `cat` activates nothing, it is a per-session cumulative delta so a silent second match is deduplication not failure, the transcript and the stream-json `commands_changed` record are the only witnesses (`skills-invoked.log` and `--debug-file` cannot see it), and `-SkillGroups` prunes user scope so the standing prune must be restored afterwards. Skills only — rules, subagents and hooks keep the manual procedure in root CLAUDE.md.

来源信息

仓库
wardawgmalvicious/agent-config
最近来源活动
2026年10月1日 18:29
检测到的 SKILL.md 语言
英语
星标
1
分支
0

安装方式

默认使用会先检查来源的 Prompt;你也可以切换为直接命令,或下载本地副本。

检查来源文件

决定是否安装前,请先阅读 SKILL.md,以及 SkillsMP 当前展示的配套文件。

文件资源管理器
2 个文件

正在显示 SKILL.md

SKILL.md
来源说明 · 只读预览
name
test-skill
description
Validate a drafted skill — write its trigger fixtures, update the activation contract table, run the static and real-path activation tests, then behaviourally test it in a cold session against a `--safe-mode` baseline. Reads its inputs from disk — brief or shipped frontmatter — so it runs cold. Encodes the traps that make a broken test look like a broken glob: activation is keyed to the `Read` tool so a Bash `cat` activates nothing, it is a per-session cumulative delta so a silent second match is deduplication not failure, the transcript and the stream-json `commands_changed` record are the only witnesses (`skills-invoked.log` and `--debug-file` cannot see it), and `-SkillGroups` prunes user scope so the standing prune must be restored afterwards. Skills only — rules, subagents and hooks keep the manual procedure in root CLAUDE.md.
when_to_use
Use when asked to test, validate or verify a skill, to check whether a `paths:` glob fires, after editing a `description`, `when_to_use` or `paths:` glob, or as the follow-on to `/author-skill`.
argument-hint
[skill-name]
disable-model-invocation
false
model
inherit
effort
max
# Test a skill Take a **drafted skill** and end at a validated one: its activation contract written and passing, its behaviour checked in a cold session, and the fixtures it ran against provably unmodified. This is the second half of `/author-skill`, which stops at a linted draft and writes no fixtures on purpose. The coupling between them is the **brief on disk**, not session state — so this skill runs cold, in a fresh session, exactly like `/drift-update`. It also runs on a skill that shipped long ago and has since been edited. That case has **no brief** and needs none; step 1 says where its inputs come from instead. Repo-relative paths are relative to the agent-config repo (`C:\Repos\Personal\agent-config`), not the session's cwd. ## What this validates, and what it does not **Skills only** — the same scope as `author-skill`. Rules get exercised incidentally, because the activation harness checks `claude/rules/*.md` alongside skills on the same fixtures, but authoring rule fixtures is not this skill's job. Out of scope, with their procedures elsewhere: | Artifact | Procedure lives in | | --- | --- | | Subagents | `tests/agents/security-reviewer/README.md` | | Enforcement hooks | same, plus the direct unit-test pattern in root `CLAUDE.md` | | Rules, as authored artifacts | root `CLAUDE.md`, "Validating a change" | ## Phase A — the activation contract Skip this phase entirely if the skill has **no `paths:` glob**. An unconditional skill has no activation contract, so there is nothing to fixture and nothing to assert. Say so and go to Phase B; do not manufacture a fixture to make the phase look done. ### 1. Read the trigger contract from disk Three entry paths, and they take their inputs from different files. Establish which one you are on before reading anything. **A newly drafted skill, arriving from `/author-skill`:** ``` docs/handoffs/execute/<skill-name>.md ``` Take from it the `paths:` glob, the named trigger queries, and the scope decisions. **Read it from the file even if you wrote it an hour ago** — the point of the disk contract is that a fresh session with no memory of the authoring run behaves identically. **An existing skill being retested after a `description`, `when_to_use` or `paths:` edit:** there is no brief, and its absence is not a fault. Briefs are deleted when spent, so a skill that shipped has none by design — stopping here would refuse the exact case the description advertises. Take the three inputs from the artifacts that already own them: | Input | Source | | --- | --- | | `paths:` glob | the `SKILL.md` itself — always authoritative, brief or no brief | | scope decisions | the set's `expected_activations.md`, which **is** the committed contract | | trigger queries | the skill's own `description` and `when_to_use` | That last row is sound rather than a fallback: `when_to_use` is defined upstream as "trigger phrases or example requests", so on a skill that carries one the queries are a first-class frontmatter field, not something to invent. Where a skill has no `when_to_use`, draw the queries from the `description` and **say in the report which ones you derived** — a derived query tests the skill against the trigger surface it actually ships, which is the point, but the reader needs to know no brief vouched for it. **An executed drift brief in `docs/audits/<date>/<source>/`, whose execution log deferred behavioural confirmation to a fresh session:** `/drift-update` cannot exercise a skill it just edited, so its log ends with the deferral and this skill is the follow-on. Inputs come from the shipped artifacts exactly as in the row above — it is not an `/author-skill` brief and names no trigger queries — but its *What to change* section names the one thing the edit added, and **that is the discriminating claim for Phase B**: the detail only the skill makes. The brief is a ledger entry and is never deleted (step 10). First run 2026-09-12 on `fabric-semantic-model-ai-instructions` from brief 05. Such a brief is a **content** edit by construction — Kind says so and the `paths:` glob is untouched — so the activation stamp stands and **Phase A needs only its static half**. `skill-status.py` says `retest-activation` when a glob really moved; trust it over re-running. The one thing you may never do is invent expectations the skill was never written to meet. Reading them off the shipped frontmatter is not that; making them up is. ### 2. Pick the fixture set The glob decides it: | Glob targets | Set | | --- | --- | | Fabric item folders (`**/*.DataPipeline/**`, `.Notebook`, `.Eventhouse`, …) | `tests/skills/fabric-triggers/` | | PBIP / report / semantic-model paths | `tests/skills/pbip-triggers/` | The two sets are **disjoint and jointly exhaustive** over the payload's conditional skills. Don't restate the count here — it was duplicated into six files, checked by nothing, and had drifted three different ways by 2026-09-02, one of them two generations stale. Each set's `expected_activations.md` owns its own figure; derive the total when you actually need it: ```powershell ./scripts/test-activation.ps1 -Set fabric -StaticOnly # then -Set pbip ``` A skill whose glob spans both sets is a design smell — raise it rather than splitting fixtures across sets. ### 3. Write or extend the fixtures Add files under that set's `fixtures/` directory, modelled on real exports. Keep them minimal but structurally faithful: the test asserts *which globs match*, so a file needs the right path and enough content to be plausible, not real data. **Mark any fixture built on an unverified shape.** The fabric set's README has a dedicated section for these, and a fixture invented from a guess will happily pass a test that asserts nothing true. **Pin a fixture whose bytes are part of its shape.** This repo's `.gitattributes` commits a CRLF fixture as LF, warning only at `git add`, and `git status` reads clean after. Give the folder `-text` in the fixture's own commit, as `SampleESS.EventSchemaSet` has. Before staging, `git hash-object --path=<f> <f>` and its `--no-filters` form must match; after, `git ls-files --eol <f>` reads `i/crlf` (2026-10-01). ### 4. Update `expected_activations.md` One row per fixture file, naming the skills that must fire. This is the **contract**, and the static check compares globs against it in both directions — a skill that fires and is not listed fails just as loudly as one listed and not firing. **A row that reads *(none)* and should now name your skill is the assertion being changed.** Say that explicitly in the commit message, because the diff on its own looks like a table edit rather than a retired negative assertion. Check whether the set's prose sections still hold too. The fabric set carries a "Fabric item types with no skill at all" section that names examples; a new skill can make one of them stale. ### 5. Static check — always, before spending a session ```powershell ./scripts/test-activation.ps1 -Set fabric -StaticOnly ``` No session, no tokens, no deploy. It compares frontmatter globs against the contract table and exits non-zero on any mismatch. `-Set` is mandatory and takes `pbip` or `fabric`. You cannot skip it by accident — the full run executes it first and refuses to continue past a failure — but run it alone while iterating, because it is the whole feedback loop for steps 3 and 4. **Run the rules pass too.** Rules carry `paths:` globs and load on the same files, so a fixture with no *skill* may still pull a rule. Doing only the skills pass is how the first version of the fabric set reported "activates nothing" for files that load `fabric-git-serialization`. The snippet is in `tests/skills/fabric-triggers/README.md` — same code, `claude/rules/*.md` instead of `skills/*/*/SKILL.md`, keyed on `p.stem`. ### 6. Real-path test — does the harness agree? ```powershell ./scripts/test-activation.ps1 -Set fabric ``` Deploys to a throwaway probe outside the repo, opens **one** cold session, has it Read every fixture, asserts the transcript, and tears down in a `finally`. **One session covers the whole set** — activation is a per-session cumulative delta, so 56 fixtures cost one session rather than 56. That is what makes this affordable enough to actually run. **Skip this step when the activation stamp is already current.** The stamp hashes `paths:` alone, so on a retest after a body or `description` edit `scripts/skill-status.py` still rates activation current and a cold session would re-prove an unchanged fact; the static check in step 5 is the whole regression, and the Phase B probe witnesses activation anyway because a conditional skill is unreachable until a matching file is Read. First applied 2026-09-12 on `fabric-semantic-model-ai-instructions` — activation stamped 2026-09-01, static 16/16, no session spent. The script already refuses the dangerous shapes, so do not re-implement guards around it: it rejects a `ProbeRoot` inside this repo, refuses user scope as a deploy target, and will not reuse a directory lacking its `.activation-probe` marker. Other parameters: `-ProbeRoot`, `-Model` (default `opus[1m]`), `-KeepProbe` to leave the probe for inspection. ## Phase B — behaviour ### 7. Load a platform skill in the probe — never in user scope A platform skill needs this step. So does a **new** skill in `workflow`, `social` or `meta`, which has no junction until the standing form runs once, `./scripts/link-claude.ps1 -SkillGroups workflow,social,meta`, never bare (root `CLAUDE.md`, Commands); the rest of those groups is deployed already, and `.claude/skills/` is read in place. **In a linked worktree, stop instead** for a skill in those three groups: step 8 would test the main checkout's copy, silently. Land it first. Copy the skill's whole group into the payload probe's **project scope**, outside this repo; give the baseline a sibling with no `.claude/`, since `--safe-mode` stops a skill loading, not a `Read` of its file. ```powershell $probe = '<a scratch directory outside this repo>' New-Item -ItemType Directory "$probe/payload/.claude/skills", "$probe/baseline" | Out-Null Get-ChildItem skills/fabric -Directory | ForEach-Object { # or skills/powerbi Copy-Item $_.FullName "$probe/payload/.claude/skills" -Recurse } ``` User scope serves every live session and this touches none; copies make teardown a plain delete, with no junction to follow back into this repo. Re-copy after an edit. Measured 2026-10-01 on `fabric-event-schema-set`: `init` listed the group's 11 unconditional skills, a matching `Read` brought the new one in, then `Skill`. The user-scope deploy this step used to make is in `references/reading-a-failure.md`. ### 8. The cold behavioural session **Cold, always.** Skills hot-reload, but subagents, commands and rules do not, and accumulated context can mask a co-load failure. Establish the baseline first: ```bash claude --safe-mode ``` That starts with the entire payload off — `CLAUDE.md`, skills, hooks, MCP, commands, agents — and is the control condition that separates behaviour the payload produces from behaviour the base model produces. It is a flag you type, never something to wire into `settings.json`. **A close baseline is not a failed skill.** Where the base model already knows the domain, `--safe-mode` reproduces most of a good answer, so "the payload run answered well" measures nothing by itself. Compare a detail that is a claim **only the skill makes**. Measured 2026-09-04 on `pbir-filters`: both runs got `SourceRef.Source` and doubled quotes; only the payload wrote the skill's 20-char hex `name` (baseline: a 32-char GUID). Read the baseline before naming the claim that separates: the sharpest one on paper is often one the baseline volunteers (`land`, 2026-09-17). **`--safe-mode` does not strip the web tools.** A baseline can fetch the Learn pages a skill was drilled from and re-derive it — a second reason it comes back close. Add `--disallowedTools WebFetch,WebSearch` **to both runs**, the payload side too, or a pass cannot separate "the skill delivered it" from "the model fetched the page the skill cites"; or compare on synthesis that sits on **no single page** (2026-09-12). **Those two flags do not turn the web off on this machine, and the gap is one-sided.** `microsoft-learn-mcp` is **user scope**, so a payload probe reaches Learn through it whatever `--disallowedTools` says, while `--safe-mode` strips MCP with the payload and the baseline cannot — leaving the very confound the flags were meant to remove. Add `--strict-mcp-config` to the payload arm (with no `--mcp-config` it drops every server) and check that the `init` record's `tools` count matches the baseline's (`fabric-eventstream`, 2026-09-13). **Prove the baseline actually stripped the payload.** A `--safe-mode` run that silently kept the skill is indistinguishable from one where the base model already knew the answer — both read as "the skill adds nothing". In a `-p` probe the `system`/`init` record names `slash_commands` and `tools`: assert the skill is absent from it and that the count dropped (`fabric-catalog-governance`, 2026-09-12). **Assert on a name only the payload provides.** `code-review` stays listed under `--safe-mode` because the CLI ships a built-in of that name, so a grep for it reads as a leaked payload — and when `code-review` itself is under test, its absence cannot be read off the list at all; assert its sibling payload names instead (measured 2026-09-13, `references/reading-a-failure.md`). Then, in a normal session, run the trigger queries from step 1, and **scope the arms to what the edit touched**: an arm tests only what its invocation path reads. A `description` or `when_to_use` edit wants a model-invoked trigger query and its `Skill` call, which a slash arm never reads; a body edit, a query that reaches the edited section, plus the slash arm if `model:` pins a model, which model-invocation drops; a `model:` edit, the slash arm; a first test or a `name:` edit, both paths. Exercise the refusal modes too — ignoring its own scope guard is a fail; `tests/skills/code-review/README.md` has the four-mode matrix. **Allowlist the tools when a trigger query names a destructive action.** A behavioural probe runs with this machine's credentials — an `az login` survives across tool calls — so "Delete the Marketing domain" or a bulk label write can reach a real tenant, and the skill's own refusal is the only thing in the way. Don't rely on it: pass `--tools Skill`, plus `Read` where a conditional skill needs its matching file, and `--strict-mcp-config`, since `--tools` governs only built-in tools. A deny list goes stale as the CLI adds tools: on 2.1.282 the one this step used to give left `Workflow`, `EnterWorktree`, `CronCreate` and `RemoteTrigger` (`references/reading-a-failure.md`, 2026-09-29). **What is measured is what the skill *says*, not whether it can call an API** — free on a query that asks for an explanation, and a real cost on one that asks the skill to *act*: it stops at the first command it cannot issue, and every later step goes unmeasured on that arm. Phrase one arm as a walkthrough ("the exact commands, start to finish"); `land`, 2026-09-13, is the measurement. A step that decides on a command's output needs it run: `Bash`, writes denied by rule, `auto` mode, never `dontAsk` (`references/reading-a-failure.md`, 2026-10-01). **Give a trigger query enough context to be answerable.** A bare imperative in an empty probe directory routes to file exploration rather than to a skill: "Deploy my data pipeline to production." sent the session hunting for a pipeline on disk and loaded nothing, while the same sentence framed with a dev workspace reached the skill. The first measures the harness's file-hunting instinct, not the trigger surface. Measured 2026-09-12 on `fabric-deployment-pipelines`. **A conditional skill has neither path until a matching file is Read.** The `paths:` glob keeps it out of the startup listing, so its `description` — the whole model-invocation trigger — is never in context, and `/<name>` answers `Unknown command`. Read a matching file first; that injects the listing entry and the model can then invoke it. The slash arm needs that Read **in its own process**: send it, then `/<name>`, as two messages to one `--input-format stream-json` print session. A `--resume` still answered `Unknown command` (`fabric-dataflow`, 2026-09-12); one process expanded it (2026-10-01, in the reference). The four-mode matrix above applies as written only to an *unconditional* skill (`fabric-data-pipeline`, 2026-09-02). **Run the behavioural session outside this repo — for a platform skill.** `.claude/settings.json` here collapses every platform skill description to `name-only`, and the description *is* the trigger — so an in-repo run is a guaranteed false negative that looks exactly like a broken skill. **A project-scope skill must run here**, where the payload is on disk and root `CLAUDE.md` may repeat it: give both arms `--tools Skill`, so neither can read the payload off disk, then add a third arm with `--tools CronList` — **model-invoked, never slash**, or the expansion inlines the body straight past the denial and the arm is inert. Never `--tools ""`: with no tool declared, the model writes its calls as text and invents their results; one read-only tool avoids it. `references/reading-a-failure.md`, 2026-09-13 and 2026-09-29. **Parallel arms read each other as peers**, and only payload arms look: keep `ListAgents` and `SendMessage` out of every arm, as `--tools` does by construction. A skill that reads peers itself (`commit`, `land`, `triage`) keeps `ListAgents`, and its arms launch one at a time (2026-09-29). **Never stop an arm for silence alone.** At max effort an arm streams only `thinking_tokens` until its first block closes, most of a 17-minute baseline run: pace it against the baseline, never the payload arm, and bound spend at launch with `--max-budget-usd` instead (2026-09-29). Confirm the skill actually loaded with `/context` rather than by asking the session — self-report is unreliable, and once omitted an unconditional skill that was certainly present. In a `-p` probe where `/context` is unavailable, use the transcript: a model-invoked skill appears as a `Skill` tool_use, while a slash-invoked one is **inlined as a command expansion** and produces no `Skill` call — so an absent `Skill` record disproves nothing on the slash path. The positive witness is a slash run on a query the NL arm answered itself through `Skill`: carrying the skill's detail with **no** `Skill` call proves the expansion (`references/reading-a-failure.md`). `--output-format stream-json --verbose` is the cheaper route to those records — it carries the `tool_use` blocks, the `init` record and, for a conditional skill, the `commands_changed` record that witnesses the matching `Read` (`references/reading-a-failure.md`) inline, so nothing has to locate a session id under `~/.claude/projects/`. Read the answers rather than grepping them (2026-09-12), and only the first `result`: a peer's message can run in an arm as a second turn (2026-09-29). **Launch a slash probe from PowerShell**, not the Bash tool. MSYS2
在 GitHub 查看
这个 SKILL.md 很大,SkillsMP 这里只预览前一段内容。 在 GitHub 查看