원클릭으로
agent-harness-construction
Framework for designing quality agents with proper action space and contracts
Codex 또는 Claude로 설치 이 Prompt를 복사해 Codex, Claude 또는 다른 어시스턴트에 붙여 넣으면 Skill 페이지를 검토하고 설치를 진행할 수 있습니다.
메뉴
Framework for designing quality agents with proper action space and contracts
Codex 또는 Claude로 설치 이 Prompt를 복사해 Codex, Claude 또는 다른 어시스턴트에 붙여 넣으면 Skill 페이지를 검토하고 설치를 진행할 수 있습니다.
SOC 직업 분류 기준
Convert PDF/EPUB library to Markdown and generate Obsidian MOC notes
Hook-based compaction suggestions at logical task boundaries
Context window management — track spend, decide when to compact, preserve state
Session-start orientation — loads context, surfaces learnings, confirms registry
Quality and semantic review — catches what automated tools miss
Planner → Architect → Critic deliberation loop — produces a formally validated ADR
| name | agent-harness-construction |
| description | Framework for designing quality agents with proper action space and contracts |
| version | 0.1.0 |
| level | 3 |
| triggers | ["design an agent","agent quality framework","harness construction","agent architecture"] |
| context_files | ["context/project.md","context/decisions.md"] |
| steps | [{"name":"Define Action Space","description":"Specify granularity and tool allowlist"},{"name":"Establish Observation Contract","description":"Define required fields in tool responses"},{"name":"Design Recovery Contract","description":"Specify error handling and retry logic"},{"name":"Budget Context","description":"Calculate context costs and set limits"},{"name":"Select Architecture Pattern","description":"Choose ReAct, function-calling, or Hybrid"},{"name":"Define Success Metrics","description":"Track completion rate, retries, pass@k, cost per task"}] |
Framework for designing quality agents. Defines contracts, budgets, and measurement before implementation.
Without systematic agent design, agents:
Agent harness construction ensures agents are well-specified before deployment.
Defines: What tools can this agent use? At what granularity?
Granularity: Micro (single file/command, high-risk), Medium (edit/read loops, standard dev), Macro (Task/orchestration, complex workflows).
Tool Allowlist Pattern:
tools:
- Read
- Grep
- Glob
disallowedTools:
- Write
- Edit
- Bash
Rule: Start restrictive, expand only when justified. Removing permissions later breaks existing workflows.
Defines: What fields must every tool response include?
Required Fields:
status: SUCCESS | PARTIAL | FAILUREsummary: One-line description of what happenednext_actions: Array of suggested follow-upsartifacts: Paths to files created/modifiedWhy: Enables reliable parsing, orchestrator coordination, and chaining without re-planning. Anti-pattern: raw output with no structure.
Defines: How does agent handle errors?
Strategies: Retry with backoff (transient errors, max 3), escalate to human (ambiguous/security, max 2 auto-attempts), graceful degradation (optional features unavailable), circuit breaker (3 identical failures = stop).
Recovery Contract Template:
recovery:
transient_errors:
max_retries: 3
backoff: [1, 5, 15]
ambiguous_errors:
escalate_after: 2
circuit_breaker:
identical_failures: 3
Defines: Maximum tokens this agent can consume.
Budget Calculation:
Agent prompt: ~2,000 tokens (instructions, examples) Tools: ~500 tokens per tool schema × N tools Working context: File reads, conversation history Output: Agent responses
Example:
Budget Limits:
| Agent Type | Token Budget | Use Case |
|---|---|---|
| Micro (researcher) | 10K-20K | Quick searches, single-file analysis |
| Standard (planner, code-reviewer) | 20K-50K | Multi-file review, planning |
| Complex (orchestrator, architect) | 50K-100K | System-wide analysis, coordination |
If budget exceeded:
Best For: Exploratory tasks, unclear solution paths, research
Pattern:
Pros: Flexible, handles ambiguity, self-correcting Cons: Higher token cost (reasoning overhead), slower
Best For: Well-defined tasks, repetitive operations, production workloads
Pattern:
Pros: Fast, predictable, low token cost Cons: Rigid, fails on ambiguous inputs
Best For: Most agent tasks (recommended default)
Pattern:
Pros: Flexible planning, efficient execution Cons: Slightly higher complexity
Track these metrics for every agent:
Completion Rate:
Retries Per Task:
pass@1 / pass@3:
Cost Per Successful Task:
Example Metrics Dashboard:
Agent: code-reviewer
Period: Last 30 days
Completion Rate: 88% (44/50 tasks)
Retries Per Task: 1.2
pass@1: 72%
pass@3: 92%
Cost Per Task: $0.08 avg
Overpowered agents: Agent has Write, Edit, Bash, Task access when it only needs Read + Grep. Start restrictive.
No observation contract: Tool responses are raw text. Downstream parsing is brittle and breaks on edge cases.
Unlimited retries: Agent retries failed operation 20 times. Use circuit breaker (3 identical failures = stop).
No context budget: Agent consumes 200K tokens on simple task. Budget forces efficiency.
Missing metrics: Can't tell if agent is improving or degrading over time. Track pass@1, cost, completion rate.