一键导入
agent-self-audit
Meta-auditing of Claude agents and their outputs. Use when: reviewing agent quality, validating agent decisions, or improving agent prompts.
用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
菜单
Meta-auditing of Claude agents and their outputs. Use when: reviewing agent quality, validating agent decisions, or improving agent prompts.
用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
基于 SOC 职业分类
Audit the developer experience of a product, SDK, docs site, or SKILL.md by dropping multiple Claude subagents at it with only a tiny task prompt and real tools (WebFetch, Bash, Write). Agents must discover the docs themselves, install deps, ask for credentials if needed, and attempt real execution. The skill captures each agent's trace — tool calls, retries, wall time, errors — and scores on Setup Friction, Speed, Efficiency, Error Recovery, and Doc Quality, then emits an HTML report with an A–F grade and concrete fixes. Use when the user asks to audit agent experience, test a skill, audit docs for agents, check if a SDK is agent-friendly, validate a SKILL.md, measure agent DX, or benchmark how painful onboarding is for an AI agent. Triggers: 'audit agent experience', 'test this skill', 'audit docs for agents', 'is my SDK agent-friendly', 'run a DX audit', 'agent experience test', 'test my docs', 'how do agents do with my product'.
Self-improving browser automation via the auto-research loop. Iteratively runs a browsing task, reads the trace, and improves the navigation skill (strategy.md) until it reliably passes. Supports parallel runs across multiple tasks using sub-agents. Use when you want to build or improve browser automation skills for specific website tasks.
Create polished design artifacts as self-contained HTML: UI mockups, interactive prototypes, wireframes, landing pages, dashboards, app screens, mobile apps, slide decks, and visual explorations. Use whenever the user asks to design, mock up, prototype, wireframe, visualize, or explore an interface, product screen, user flow, content layout, visual artifact, or pitch/deck concept, even if they do not say "design". Also use for setting up, importing, or authoring reusable design systems, UI kits, brand tokens, component libraries, or loadable design-system bundles. The skill guides context gathering, clarifying questions, choosing fidelity, selecting or binding design systems, creating project folders, building one or more HTML deliverables, previewing them, and verifying they load cleanly. It is harness-agnostic for Claude Code, Cursor, Codex Agent, and similar file-capable agents; harness-specific ask, preview, screenshot, and verification tools are resolved from references/.
Use the Browserbase CLI (`browse`) for Browserbase Functions and platform API workflows. Use when the user asks to run `browse`, deploy or invoke functions, manage sessions, projects, contexts, or extensions, fetch a page through the Browserbase Fetch API, search the web through the Browserbase Search API, or scaffold starter templates. Prefer the Browser skill for interactive browsing; use the top-level `browse` driver commands (`browse open`, `browse get`, etc.) only when the user explicitly wants the CLI path.
Sync cookies from local Chrome to a Browserbase persistent context so the browse CLI can access authenticated sites. Use when the user wants to browse as themselves, sync cookies, or log into sites via Browserbase.
Build local constrained-browser agents with a safe_browser tool that owns CDP, enforces a domain allowlist with Fetch interception, and lets a runtime Claude Agent SDK agent complete browsing tasks without raw browser, shell, or CDP access. Use when the user wants an agent to browse or scrape while staying on approved domains, demo blocked off-domain navigation, or generate a safe browser client.
| name | Agent Self Audit |
| description | Meta-auditing of Claude agents and their outputs. Use when: reviewing agent quality, validating agent decisions, or improving agent prompts. |
Meta-auditing framework for Claude agent quality.
| Dimension | Metrics | Target |
|---|---|---|
| Correctness | Bug rate | < 5% |
| Efficiency | Tokens/task | baseline ± 20% |
| Safety | Prompt injection defense | 100% |
| Coherence | Follows instructions | > 90% |
interface AgentEvaluation {
task: string;
intentMatch: 'exact' | 'partial' | 'missed';
codeQuality: {
types: boolean;
security: boolean;
tests: boolean;
};
efficiency: {
tokens: number;
timeMs: number;
};
issues: string[];
}
Audit questions:
Track metrics over time: