| name | agent-creator |
| description | Comprehensive guide for creating high-quality specialized agents following v2 architecture patterns. Use this skill when users need to design and implement new agents, understand agent architecture, or learn best practices for agent creation. |
| license | MIT |
Agent Creator
Purpose: Teach the principles, patterns, and practices for creating high-quality specialized agents that follow v2 architecture standards.
Critical Use Case: This skill provides structured guidance for creating agents from requirements through deployment, preventing common mistakes and ensuring quality through automated validation.
Differentiation from agent-hr-manager:
- agent-creator (this skill) = Teaching guide, knowledge resource, passive reference 📖
- agent-hr-manager (agent) = Autonomous executor, active creator, can use this skill 👨🏫
Use agent-creator when learning how to create agents. Use agent-hr-manager when you want an agent automatically created.
When to Use This Skill
Use agent-creator when:
- Creating a new specialized agent from scratch
- Learning agent architecture and design patterns
- Understanding quality validation (0-80 rubric)
- Troubleshooting agent quality issues
- Migrating agents to v2 architecture
- Training others on agent creation
Do NOT use for:
- Creating skills (use skill-creator skill instead)
- Quick agent modifications (just edit directly)
- General Claude usage questions
6-Step Agent Creation Workflow
Step 0: Research Existing Patterns (BEFORE DESIGN)
Objective: Understand what already exists before creating something new. This prevents duplicate agents and ensures you leverage proven patterns.
Why this matters: Creating an agent without research leads to:
- Duplicating existing agent functionality
- Missing reusable patterns from similar agents
- Not discovering skills that solve part of the problem
- Reinventing methodology that already exists
Actions:
-
Search for Similar Agents:
ls ~/.claude/agents/ | head -20
grep -l "[domain-keyword]" ~/.claude/agents/*.md 2>/dev/null
-
Review Relevant Agent Examples:
- Read
references/agent-examples.md for quality patterns
- Study agents with high quality scores (60+/80)
- Note phase structures that work for similar domains
-
Check Skill Inventory:
ls ~/.claude/skills/
grep -r "[domain-keyword]" ~/.claude/skills/*/SKILL.md 2>/dev/null | head -10
-
Decision Checkpoint (REQUIRED):
| Question | Answer |
|----------|--------|
| Similar agent exists? | [yes/no - if yes, consider tuning instead] |
| Relevant skills found? | [list skills to integrate] |
| Reusable patterns identified? | [list patterns to follow] |
| Proceed with new agent? | [yes with justification] |
-
Research Novel Domains (if unfamiliar):
- Use WebSearch for domain best practices
- Find authoritative sources and frameworks
- Document key methodologies the agent should follow
Deliverable: Research summary documenting similar agents, skills to integrate, and justification for new agent.
Step 1: Temporal Awareness & Requirements Gathering (CRITICAL)
Objective: Establish current date context and understand what the agent needs to do.
1.1 Establish Temporal Context (REQUIRED)
Why this matters: Legal documents, contracts, compliance reports, and project documentation with incorrect dates create serious risks. The pizza baker contract bug (January 2025 vs November 2025) demonstrated this - wrong dates in legal documents can affect validity and compliance.
Implementation:
## Phase 1: [Phase Name] & Temporal Awareness
**Objective**: [Phase goal]
**Actions**:
1. **Establish Temporal Context** (REQUIRED):
```bash
CURRENT_DATE=$(date '+%Y-%m-%d') # ISO 8601: 2025-11-06
READABLE_DATE=$(date '+%B %d, %Y') # Human: November 06, 2025
TIMESTAMP=$(date '+%Y-%m-%d %H:%M:%S %Z') # Full: 2025-11-06 12:34:56 EET
- Use CURRENT_DATE for document metadata, version numbers
- Use READABLE_DATE for human-readable headers
- Use TIMESTAMP for detailed audit trails
- [Other Phase 1 actions...]
Deliverable: [Concrete output]
**Validation**: The validate_agent.py script checks for temporal awareness pattern in Phase 1.
#### 1.2 Gather Requirements
**Key Questions**:
1. **Problem Definition**: What problem does this agent solve?
2. **Domain Expertise**: What specialized knowledge is needed?
3. **Tool Requirements**: Which tools will it need? (Read, Write, Edit, Bash, Grep, Glob, etc.)
4. **Typical Workflow**: What is the step-by-step process?
5. **Success Metrics**: How do we know it worked?
6. **Edge Cases**: What unusual situations must it handle?
**Techniques**:
- **Example-Based**: Ask for 2-3 concrete usage examples
- **Anti-Pattern Analysis**: What should it NOT do?
- **Boundary Testing**: What are the limits (file size, complexity, scope)?
**Output**: Requirements document or clear mental model before proceeding.
---
### Step 1.5: Skill Discovery & Integration Planning
**Objective**: Identify which existing skills to integrate into the agent and how.
**Why this matters**: This skill moves beyond "prompt engineering" into "cognitive architecture" — ensuring the agent doesn't use a hammer for a screw. Proper skill integration gives agents specialized capabilities without reinventing them.
**Actions**:
1. **Map Requirements to Skill Categories**:
```markdown
| Agent Requirement | Skill Category | Candidate Skills |
|-------------------|----------------|------------------|
| Debugging logic | Reasoning | hypothesis-elimination, self-reflecting-chain |
| Security review | Development | security-analysis-skills, adversarial-reasoning |
| Documentation | Documentation | document-writing-skills |
| Database ops | Integration | chromadb-integration-skills |
| Testing | Development | testing-methodology-skills |
| Error handling | Development | error-handling-skills |
-
Evaluate Each Candidate Skill:
| Skill | Size | Active? | Integrate or Inline? |
|-------|------|---------|---------------------|
| [skill-name] | [lines] | [yes/no] | [integrate/inline/skip] |
Decision Criteria:
- Integrate if: Skill >100 lines, actively maintained, reusable
- Inline if: Simple pattern <20 lines, agent-specific variant needed
- Skip if: Not relevant after review
-
Document Skills Integration:
**Skills Integration**: skill-1, skill-2, skill-3
This goes in the agent's header metadata.
-
Plan Skill Invocation Points:
| Phase | When to Invoke | Skill |
|-------|----------------|-------|
| Phase 2 | Complex decision | integrated-reasoning-v2 |
| Phase 3 | Design validation | adversarial-reasoning |
| Phase 4 | Error recovery | hypothesis-elimination |
-
Check for Handover/Parallelism Needs:
- Will the agent need multi-pattern reasoning? → Add reasoning-handover-protocol
- Will tasks run in parallel? → Add parallel-execution skill
- See
cognitive-skills/INTEGRATION_GUIDE.md for patterns
Deliverable: Skill integration plan with invocation points documented.
Step 2: Architecture Design
Objective: Design the agent's phase structure, tool selection, and quality criteria.
2.1 Determine Agent Complexity
Decision Tree: Simple vs Complex Agent
Simple Agent (3 phases, <200 lines):
- Single domain focus (e.g., PDF manipulation, CSV parsing)
- Linear workflow (no branching)
- Minimal state management
- Examples: pdf-creator-agent, code-formatter
Complex Agent (4-5 phases, 200-250 lines):
- Multiple operation modes (e.g., create, read, update)
- Conditional branching or decision trees
- State tracking across phases
- Examples: legal-agent, ceo-orchestrator, agent-hr-manager
When to use integrated-reasoning-v2: 8+ decision dimensions, strategic importance, >90% confidence required
- 9 patterns available: ToT, BoT, SRC, HE, AR, DR, AT, RTR, NDF
- 11 scoring dimensions for pattern selection
- See
cognitive-skills/INTEGRATION_GUIDE.md for full integration patterns
2.2 Design Phase Structure
Guidelines (from agent-design-patterns.md):
- 3-5 phases optimal (2 too simple, 6+ too complex)
- Each phase has ONE clear objective
- Actions are SPECIFIC, not generic
- Deliverables are CONCRETE artifacts
Phase Structure Template:
## Phase N: [Descriptive Name]
**Objective**: [One sentence describing the goal]
**Actions**:
1. [Specific action with tool: "Use Grep to search for X pattern in Y files"]
2. [Specific action with tool: "Use Edit to modify lines 45-52 in config.yml"]
3. [Specific action with condition: "If errors found, use TodoWrite to track fixes"]
**Deliverable**: [Concrete output: "List of 5 validated regex patterns with test cases"]
Example from kaggle-leak-auditor:
- Phase 1: Static Code Analysis → List of violations
- Phase 2: Runtime Validation → Validation results
- Phase 3: Report Generation → Audit report with recommendations
2.3 Select Tools
Common Tool Combinations:
- File analysis: Read, Grep, Glob
- Code modification: Read, Edit, Write
- Research: WebSearch, WebFetch, Read
- Execution: Bash, TodoWrite, Read
- Complex tasks: Task (invoke other agents)
Tool Selection Criteria:
- Minimal set: Only include tools actually used in phases
- Specific over general: Edit > Write for modifications
- Composed workflows: Grep to find, Read to analyze, Edit to modify
2.4 Define Success Criteria (10-16 items)
Categories:
- Phase Deliverables (3-5 items): "✅ Phase 1 violations list complete with severity scores"
- Quality Gates (2-3 items): "✅ All findings validated with evidence"
- Confidence (1 item): "✅ Confidence level >85% with clear reasoning"
- Documentation (2-3 items): "✅ Report includes examples and references"
- Edge Cases (2-3 items): "✅ Handled missing files gracefully"
- Temporal (1 item): "✅ Document dated with current date"
Format:
## Success Criteria
- ✅ Temporal awareness established in Phase 1
- ✅ Phase 1 deliverable: [specific output]
- ✅ Phase 2 deliverable: [specific output]
- ✅ All files created/modified successfully
- ✅ Quality validation passed with score ≥70/80
- ✅ Confidence level >85% with supporting evidence
- ✅ Edge cases documented and handled
- ✅ Reference documentation created (if using progressive disclosure)
[10-16 total items]
2.5 Design Self-Critique (6-10 questions)
Question Categories:
- Completeness: "Did I check all [domain-specific items]?"
- Confidence: "What is my confidence level? Why?"
- Assumptions: "What assumptions did I make?"
- False Positives: "Could [finding X] be wrong? How?"
- False Negatives: "What might I have missed?"
- Verification: "How can user verify this?"
- Temporal: "Did I use current date correctly?"
Format:
## Self-Critique
1. **Domain Accuracy**: Did I correctly apply [domain] expertise?
2. **Tool Selection**: Did I use optimal tools for each task?
3. **Edge Cases**: Did I handle errors and failures gracefully?
4. **Temporal Accuracy**: Did I establish current date in Phase 1?
5. **Confidence Basis**: What evidence supports my confidence level?
6. **Assumptions**: What assumptions should the user validate?
[6-10 total questions]
2.6 Define Confidence Thresholds
Three-Tier System:
## Confidence Thresholds
- **High (85-95%)**: [Specific conditions: "All criteria met, deliverables complete, tests passed"]
- **Medium (70-84%)**: [Conditions: "Most criteria met, minor issues present, acceptable quality"]
- **Low (<70%)**: [Conditions: "Significant issues, incomplete work - continue working"]
Domain-Specific Examples:
- Code analysis: Based on test coverage, execution traces
- Legal: Based on citation verification, precedent alignment
- Research: Based on source quality, corroboration
- Debugging: Based on reproduction success, log evidence
Step 3: Implementation
Objective: Write the agent definition file following v2 architecture.
3.1 Create Agent Frontmatter
Template:
---
name: agent-name
description: Clear one-sentence description. Use when [specific trigger conditions]. Examples: [concrete user questions].
tools: Read, Write, Edit, Bash, Grep, Glob, TodoWrite
model: claude-sonnet-4-5
color: blue
---
Guidelines: