一键导入
eval-harness
Evaluation framework for measuring agent performance
用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
菜单
Evaluation framework for measuring agent performance
用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
基于 SOC 职业分类
| name | eval-harness |
| description | Evaluation framework for measuring agent performance |
| triggers | ["manual"] |
Purpose: Measure and improve agent performance through structured evaluation
This skill provides a framework for evaluating agent performance across key dimensions.
| Metric | Target |
|---|---|
| First-time success rate | >80% |
| Iterations to solution | <3 |
| Test coverage | >80% |
| Build success | 100% |
# Evaluation Report
## Session: [ID]
## Task: [Description]
### Metrics
| Dimension | Score | Notes |
| :--------- | :---- | :------------------ |
| Accuracy | 9/10 | Minor fix needed |
| Efficiency | 8/10 | 2 iterations |
| Alignment | 10/10 | All constraints met |
| Quality | 9/10 | Good coverage |
### Overall: 36/40 (90%)
### Learnings
- [What went well]
- [What could improve]
Pull request lifecycle domain knowledge — branch strategy detection, PR size classification, confidence-scored review, git-aware context, PR analytics, dependency management, and split/merge/describe operations.
Production readiness audit domains, weighted scoring criteria, and check specifications for the /preflight workflow.
Application scaffolding orchestrator. Creates full-stack applications from requirements, selects tech stack, coordinates agents.
Production deployment workflows, rollback strategies, and CI/CD best practices.
Internationalization and localization patterns for multi-language applications
Mobile UI/UX patterns for iOS and Android. Touch-first, platform-respectful design with React Native/Expo focus.