用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
直接命令不会经过审查 Prompt;运行前请先检查来源。
npx skills add https://github.com/jmagly/aiwg --skill eval-report命令会保持在同一行。复制前请横向滚动并检查完整内容。
想先保存到本地?可下载 SkillsMP 当前能够提供的文件。
基于 SOC 职业分类
正在显示 SKILL.md
| namespace | aiwg |
| name | eval-report |
| platforms | ["all"] |
| description | Generate an aggregate agent quality report from evaluation results, showing scores, regressions, and recommendations |
Generate a quality report from accumulated evaluation results.
/eval-report
/eval-report --output .aiwg/reports/quality-report.md
/eval-report --compare previous-report.json
/eval-report --mode sdlc --format json
| Option | Default | Description |
|---|---|---|
| --output | stdout | Output file path |
| --compare | none | Previous report to diff against |
| --mode | all | Agent category: sdlc, marketing, forensics, all |
| --format | markdown | Output format: markdown, json |
| --since | none | Only include results after this date (ISO 8601) |
| --threshold | 0.85 | Score below this triggers a warning |
eval-*.json files from .aiwg/reports/Overall health at a glance — total agents tested, aggregate score, regression count.
Pass rates per Roig (2025) failure archetype across all agents.
Agents below the --threshold, with consecutive-failure streaks flagged.
When --compare is provided: agents whose scores dropped since the baseline.
Prioritized action list: which agents to review, which archetypes to harden.
# Agent Quality Report
**Generated**: 2026-04-01T10:30:00Z
**Agents Tested**: 58
**Overall Score**: 87%
**Regressions**: 2
## By Archetype
| Archetype | Pass Rate | Trend |
|-----------|-----------|-------|
| #1 Grounding | 92% | ↑ |
| #2 Substitution | 88% | → |
| #3 Distractor | 78% | ↓ |
| #4 Recovery | 90% | ↑ |
## Agents Needing Attention
| Agent | Score | Consecutive Failures | Issue |
|-------|-------|---------------------|-------|
| data-analyst | 72% | 3 | distractor-test |
| api-designer | 79% | 1 | latency regression (+40%) |
## Recommendations
1. Review `data-analyst` context filtering — failed distractor-test 3 consecutive runs
2. Investigate `api-designer` tool selection — latency regression
3. Increase distractor-test scenarios for marketing agents (78% pass rate below 80% target)
{
"generated": "2026-04-01T10:30:00Z",
"summary": {
"agents_tested": 58,
"overall_score": 0.87,
"regressions": 2
},
"by_archetype": {
"grounding": 0.92,
"substitution": 0.88,
"distractor": 0.78,
"recovery": 0.90
},
"agents_needing_attention": [
{"agent": "data-analyst", "score": 0.72, "consecutive_failures": 3, "issue"
# Standard report to stdout
/eval-report
# Save to file
/eval-report --output .aiwg/reports/quality-$(date +%Y%m%d).md
# Compare against baseline
/eval-report --compare .aiwg/reports/quality-20260301.json
# JSON for CI consumption
/eval-report --format json --threshold 0.80
# SDLC agents only
/eval-report --mode sdlc
/eval-agent - Test individual agents/eval-workflow - Test multi-agent workflowsaiwg lint agents - Static validationGenerate evaluation report: $ARGUMENTS