بنقرة واحدة
eval-harness
Evaluation framework for measuring agent performance
التثبيت باستخدام Codex أو Claude انسخ هذا Prompt والصقه في Codex أو Claude أو مساعد آخر ليراجع صفحة Skill ويثبّتها لك.
القائمة
Evaluation framework for measuring agent performance
التثبيت باستخدام Codex أو Claude انسخ هذا Prompt والصقه في Codex أو Claude أو مساعد آخر ليراجع صفحة Skill ويثبّتها لك.
استنادا إلى تصنيف SOC المهني
Pull request lifecycle domain knowledge — branch strategy detection, PR size classification, confidence-scored review, git-aware context, PR analytics, dependency management, and split/merge/describe operations.
Production readiness audit domains, weighted scoring criteria, and check specifications for the /preflight workflow.
Application scaffolding orchestrator. Creates full-stack applications from requirements, selects tech stack, coordinates agents.
Production deployment workflows, rollback strategies, and CI/CD best practices.
Internationalization and localization patterns for multi-language applications
Mobile UI/UX patterns for iOS and Android. Touch-first, platform-respectful design with React Native/Expo focus.
| name | eval-harness |
| description | Evaluation framework for measuring agent performance |
| triggers | ["manual"] |
Purpose: Measure and improve agent performance through structured evaluation
This skill provides a framework for evaluating agent performance across key dimensions.
| Metric | Target |
|---|---|
| First-time success rate | >80% |
| Iterations to solution | <3 |
| Test coverage | >80% |
| Build success | 100% |
# Evaluation Report
## Session: [ID]
## Task: [Description]
### Metrics
| Dimension | Score | Notes |
| :--------- | :---- | :------------------ |
| Accuracy | 9/10 | Minor fix needed |
| Efficiency | 8/10 | 2 iterations |
| Alignment | 10/10 | All constraints met |
| Quality | 9/10 | Good coverage |
### Overall: 36/40 (90%)
### Learnings
- [What went well]
- [What could improve]