Use when creating new skills, editing existing skills, or verifying skills work before deployment
zh_description
用于writing、技能,支持任务规划、执行、评审和验证。
version
1.0.3
author
seaworld008
source
in-house
source_url
tags
["skills", "authoring", "workflow"]
created_at
2026-04-13
updated_at
2026-07-27
quality
4
complexity
intermediate
Writing Skills
Overview
Writing skills IS Test-Driven Development applied to process documentation.
Personal skills live in your runtime's skills directory (~/.claude/skills/ on Claude Code) — see codex-tools.md or gemini-tools.md for the path on those runtimes. Codex, Copilot CLI, and Gemini CLI all also recognize ~/.agents/skills/ as a cross-runtime alias.
You write test cases (pressure scenarios with subagents), watch them fail (baseline behavior), write the skill (documentation), watch tests pass (agents comply), and refactor (close loopholes).
Core principle: If you didn't watch an agent fail without the skill, you don't know if the skill teaches the right thing.
REQUIRED BACKGROUND: You MUST understand superpowers:test-driven-development before using this skill. That skill defines the fundamental RED-GREEN-REFACTOR cycle. This skill adapts TDD to documentation.
Official guidance: For Anthropic's official skill authoring best practices, see anthropic-best-practices.md. This document provides additional patterns and guidelines that complement the TDD-focused approach in this skill.
What is a Skill?
A skill is a reference guide for proven techniques, patterns, or tools. Skills help future agents find and apply effective approaches.
name: Use letters, numbers, and hyphens only (no parentheses, special chars)
description: Third-person, describes ONLY when to use (NOT what it does)
Start with "Use when..." to focus on triggering conditions
Include specific symptoms, situations, and contexts
NEVER summarize the skill's process or workflow (see SDO section for why)
Keep under 500 characters if possible
---
name: Skill-Name-With-Hyphens
description: Use when [specific triggering conditions and symptoms]
---# Skill Name## Overview
What is this? Core principle in 1-2 sentences.
## When to Use
[Small inline flowchart IF decision non-obvious]
Bullet list with SYMPTOMS and use cases
When NOT to use
## Core Pattern (for techniques/patterns)
Before/after code comparison
## Quick Reference
Table or bullets for scanning common operations
## Implementation
Inline code for simple patterns
Link to file for heavy reference or reusable tools
## Common Mistakes
What goes wrong + fixes
## Real-World Impact (optional)
Concrete results
Skill Discovery Optimization (SDO)
Critical for discovery: Future agents need to FIND your skill
1. Rich Description Field
Purpose: Your agent reads the description to decide which skills to load for a given task. Make it answer: "Should I read this skill right now?"
Format: Start with "Use when..." to focus on triggering conditions
CRITICAL: Description = When to Use, NOT What the Skill Does
The description should ONLY describe triggering conditions. Do NOT summarize the skill's process or workflow in the description.
Why this matters: Testing revealed that when a description summarizes the skill's workflow, an agent may follow the description instead of reading the full skill content. A description saying "code review between tasks" caused an agent to do ONE review, even though the skill's flowchart clearly showed TWO reviews (spec compliance then code quality).
When the description was changed to just "Use when executing implementation plans with independent tasks" (no workflow summary), the agent correctly read the flowchart and followed the two-stage review process.
The trap: Descriptions that summarize workflow create a shortcut agents will take. The skill body becomes documentation agents skip.
# ❌ BAD: Summarizes workflow - agents may follow this instead of reading skilldescription:Usewhenexecutingplans-dispatchessubagentpertaskwithcodereviewbetweentasks# ❌ BAD: Too much process detaildescription:UseforTDD-writetestfirst,watchitfail,writeminimalcode,refactor# ✅ GOOD: Just triggering conditions, no workflow summarydescription:Usewhenexecutingimplementationplanswithindependenttasksinthecurrentsession# ✅ GOOD: Triggering conditions onlydescription:Usewhenimplementinganyfeatureorbugfix,beforewritingimplementationcode
Content:
Use concrete triggers, symptoms, and situations that signal this skill applies
Describe the problem (race conditions, inconsistent behavior) not language-specific symptoms (setTimeout, sleep)
Keep triggers technology-agnostic unless the skill itself is technology-specific
If skill is technology-specific, make that explicit in the trigger
Write in third person (injected into system prompt)
NEVER summarize the skill's process or workflow
# ❌ BAD: Too abstract, vague, doesn't include when to usedescription:Forasynctesting# ❌ BAD: First persondescription:Icanhelpyouwithasynctestswhenthey'reflaky# ❌ BAD: Mentions technology but skill isn't specific to itdescription:UsewhentestsusesetTimeout/sleepandareflaky# ✅ GOOD: Starts with "Use when", describes problem, no workflowdescription:Usewhentestshaveraceconditions,timingdependencies,orpass/failinconsistently# ✅ GOOD: Technology-specific skill with explicit triggerdescription:UsewhenusingReactRouterandhandlingauthenticationredirects
Problem: getting-started and frequently-referenced skills load into EVERY conversation. Every token counts.
Target word counts:
getting-started workflows: <150 words each
Frequently-loaded skills: <200 words total
Other skills: <500 words (still be concise)
Techniques:
Move details to tool help:
# ❌ BAD: Document all flags in SKILL.md
search-conversations supports --text, --both, --after DATE, --before DATE, --limit N
# ✅ GOOD: Reference --help
search-conversations supports multiple modes and filters. Run --helpfor details.
Use cross-references:
# ❌ BAD: Repeat workflow details
When searching, dispatch subagent with template...
[20 lines of repeated instructions]
# ✅ GOOD: Reference other skill
Always use subagents (50-100x context savings). REQUIRED: Use [other-skill-name] for workflow.
Compress examples:
# ❌ BAD: Verbose example (42 words)
your human partner: "How did we handle authentication errors in React Router before?"
You: I'll search past conversations for React Router authentication patterns.
[Dispatch subagent with search query: "React Router authentication error handling 401"]
# ✅ GOOD: Minimal example (20 words)
Partner: "How did we handle auth errors in React Router?"
You: Searching...
[Dispatch subagent → synthesis]
Eliminate redundancy:
Don't repeat what's in cross-referenced skills
Don't explain what's obvious from command
Don't include multiple examples of same pattern
Verification:
wc -w skills/path/SKILL.md
# getting-started workflows: aim for <150 each# Other frequently-loaded: aim for <200 total
Why no @ links:@ syntax force-loads files immediately, consuming 200k+ context before you need them.
Flowchart Usage
digraph when_flowchart {
"Need to show information?" [shape=diamond];
"Decision where I might go wrong?" [shape=diamond];
"Use markdown" [shape=box];
"Small inline flowchart" [shape=box];
"Need to show information?" -> "Decision where I might go wrong?" [label="yes"];
"Decision where I might go wrong?" -> "Small inline flowchart" [label="yes"];
"Decision where I might go wrong?" -> "Use markdown" [label="no"];
}
Use flowcharts ONLY for:
Non-obvious decision points
Process loops where you might stop too early
"When to use A vs B" decisions
Never use flowcharts for:
Reference material → Tables, lists
Code examples → Markdown blocks
Linear instructions → Numbered lists
Labels without semantic meaning (step1, helper2)
See graphviz-conventions.dot in this directory for graphviz style rules.
Visualizing for your human partner: Use render-graphs.js in this directory to render a skill's flowcharts to SVG:
./render-graphs.js ../some-skill # Each diagram separately
./render-graphs.js ../some-skill --combine # All diagrams in one SVG
Code Examples
One excellent example beats many mediocre ones
Choose most relevant language:
Testing techniques → TypeScript/JavaScript
System debugging → Shell/Python
Data processing → Python
Good example:
Complete and runnable
Well-commented explaining WHY
From real scenario
Shows pattern clearly
Ready to adapt (not generic template)
Don't:
Implement in 5+ languages
Create fill-in-the-blank templates
Write contrived examples
You're good at porting - one great example is enough.
File Organization
Self-Contained Skill
defense-in-depth/
SKILL.md # Everything inline
When: All content fits, no heavy reference needed
Skill with Reusable Tool
condition-based-waiting/
SKILL.md # Overview + patterns
example.ts # Working helpers to adapt
Recognition scenarios: Do they recognize when pattern applies?
Application scenarios: Can they use the mental model?
Counter-examples: Do they know when NOT to apply?
Success criteria: Agent correctly identifies when/how to apply pattern
Reference Skills (documentation/APIs)
Examples: API documentation, command references, library guides
Test with:
Retrieval scenarios: Can they find the right information?
Application scenarios: Can they use what they found correctly?
Gap testing: Are common use cases covered?
Success criteria: Agent finds and correctly applies reference information
Common Rationalizations for Skipping Testing
Excuse
Reality
"Skill is obviously clear"
Clear to you ≠ clear to other agents. Test it.
"It's just a reference"
References can have gaps, unclear sections. Test retrieval.
"Testing is overkill"
Untested skills have issues. Always. 15 min testing saves hours.
"I'll test if problems emerge"
Problems = agents can't use skill. Test BEFORE deploying.
"Too tedious to test"
Testing is less tedious than debugging bad skill in production.
"I'm confident it's good"
Overconfidence guarantees issues. Test anyway.
"Academic review is enough"
Reading ≠ using. Test application scenarios.
"No time to test"
Deploying untested skill wastes more time fixing it later.
All of these mean: Test before deploying. No exceptions.
Match the Form to the Failure
Before writing guidance, classify the baseline failure. The form that bulletproofs one failure type measurably backfires on another.
Baseline failure
Right form
Wrong form
Skips/violates a rule under pressure (knows better, does it anyway)
Prohibition + rationalization table + red flags (see Bulletproofing below)
Soft guidance ("prefer...", "consider...")
Complies, but output has the wrong shape (bloated prompt, buried verdict, restated spec)
Positive recipe or contract: state what the output IS — its parts, in order
Prohibition list ("don't restate", "never narrate")
Omits a required element from something they already produce
Structural: REQUIRED field or slot in the template they fill in
Prose reminders near the template
Behavior should depend on a condition
Conditional keyed to an observable predicate ("if the brief exists, reference it")
Unconditional rule + exemption clauses
Why prohibitions backfire on shaping problems: under a competing incentive ("make the prompt self-contained"), agents negotiate with "don't X". In head-to-head wording tests on dispatch-prompt guidance, the prohibition arm produced clearly more of the unwanted content than the recipe arm (fully separated distributions), and trended worse than even the no-guidance control — micro-test your own case rather than assuming, but never reach for the prohibition by default. A recipe leaves nothing to negotiate: the output matches the stated shape or it doesn't.
Rules for whichever form you pick:
No nuance clauses. "Don't X unless it matters" reopens the negotiation — appending a single nuance clause to a winning recipe degraded it from consistent to noisy in the same wording tests. Express a real exception as its own conditional on an observable predicate.
Exemption clauses don't scope. "This limit doesn't apply to code blocks" still suppresses code blocks. If part of the output must be exempt, restructure so the rule can't reach it.
Bulletproofing Skills Against Rationalization
Skills that enforce discipline (like TDD) need to resist rationalization. Agents are smart and will find loopholes when under pressure.
Scope: this toolkit is for discipline failures — an agent that knows the rule and skips it under pressure. For wrong-shaped output or omitted elements, prohibition-based bulletproofing backfires; use the forms in Match the Form to the Failure instead.
Psychology note: Understanding WHY persuasion techniques work helps you apply them systematically. See persuasion-principles.md for research foundation (Cialdini, 2021; Meincke et al., 2025) on authority, commitment, scarcity, social proof, and unity principles.
Close Every Loophole Explicitly
Don't just state the rule - forbid specific workarounds:
No exceptions:
Don't keep it as "reference"
Don't "adapt" it while writing tests
Don't look at it
Delete means delete
</Good>
### Address "Spirit vs Letter" Arguments
Add foundational principle early:
```markdown
**Violating the letter of the rules is violating the spirit of the rules.**
This cuts off entire class of "I'm following the spirit" rationalizations.
Build Rationalization Table
Capture rationalizations from baseline testing (see Testing section below). Every excuse agents make goes in the table:
| Excuse | Reality |
|--------|---------|
| "Too simple to test" | Simple code breaks. Test takes 30 seconds. |
| "I'll test after" | Tests passing immediately prove nothing. |
| "Tests after achieve same goals" | Tests-after = "what does this do?" Tests-first = "what should this do?" |
Create Red Flags List
Make it easy for agents to self-check when rationalizing:
## Red Flags - STOP and Start Over- Code before test
- "I already manually tested it"
- "Tests after achieve the same purpose"
- "It's about spirit not ritual"
- "This is different because..."
**All of these mean: Delete code. Start over with TDD.**
Update SDO for Violation Symptoms
Add to description: symptoms of when you're ABOUT to violate the rule:
Run pressure scenario with subagent WITHOUT the skill. Document exact behavior:
What choices did they make?
What rationalizations did they use (verbatim)?
Which pressures triggered violations?
This is "watch the test fail" - you must see what agents naturally do before writing the skill.
GREEN: Write Minimal Skill
Write skill that addresses those specific rationalizations. Don't add extra content for hypothetical cases.
Run same scenarios WITH skill. Agent should now comply.
REFACTOR: Close Loopholes
Agent found new rationalization? Add explicit counter. Re-test until bulletproof.
Micro-Test Wording Before Full Scenarios
Full pressure-scenario runs are the final gate, but they are slow and expensive per iteration. Verify the wording itself first with micro-tests:
One fresh-context sample per call — a raw API call, or a single-shot subagent if you don't have API access. System prompt = the realistic context the guidance will live in (the full skill or prompt template, not the guidance in isolation); user message = a task that tempts the failure.
Always include a no-guidance control. If the control doesn't exhibit the failure, there is nothing to fix — stop, don't author the guidance.
5+ reps per variant. Single samples lie.
Manually read every flagged match. Score programmatically if you like, but template echoes and quoted counter-examples masquerade as hits; automated counts alone overstate both failure and success.
Variance is a metric. When guidance lands, reps converge on the same shape. Five different interpretations across five reps means the wording isn't binding — tighten the form before adding words.
Micro-tests verify wording; they do not replace pressure scenarios for discipline skills.
helper1, helper2, step3, pattern4
Why bad: Labels should have semantic meaning
STOP: Before Moving to Next Skill
After writing ANY skill, you MUST STOP and complete the deployment process.
Do NOT:
Create multiple skills in batch without testing each
Move to next skill before current one is verified
Skip testing because "batching is more efficient"
The deployment checklist below is MANDATORY for EACH skill.
Deploying untested skills = deploying untested code. It's a violation of quality standards.
Skill Creation Checklist (TDD Adapted)
IMPORTANT: Create a todo for EACH checklist item below.
RED Phase - Write Failing Test:
Create pressure scenarios (3+ combined pressures for discipline skills)
Run scenarios WITHOUT skill - document baseline behavior verbatim
Identify patterns in rationalizations/failures
GREEN Phase - Write Minimal Skill:
Name uses only letters, numbers, hyphens (no parentheses/special chars)
YAML frontmatter with required name and description fields (max 1024 chars; see spec)
Description starts with "Use when..." and includes specific triggers/symptoms
Description written in third person
Keywords throughout for search (errors, symptoms, tools)
Clear overview with core principle
Address specific baseline failures identified in RED
Guidance form matches the failure type (see Match the Form to the Failure)
For behavior-shaping guidance: wording micro-tested against a no-guidance control (5+ reps, every flagged match read manually) — N/A for pure reference skills
Code inline OR link to separate file
One excellent example (not multi-language)
Run scenarios WITH skill - verify agents now comply
REFACTOR Phase - Close Loopholes:
Identify NEW rationalizations from testing
Add explicit counters (if discipline skill)
Build rationalization table from all test iterations
Create red flags list
Re-test until bulletproof
Quality Checks:
Small flowchart only if decision non-obvious
Quick reference table
Common mistakes section
No narrative storytelling
Supporting files only for tools or heavy reference
Deployment:
Commit skill to git and push to your fork (if configured)
Consider contributing back via PR (if broadly useful)