Skip to main content

agentic-evals

Agentic evaluation of skills using multi-trial fixtures, deterministic command assertions, trajectory checks, safety constraints, and evidence-backed readiness scoring. Use when users ask for agentic evals, multi-trial skill evaluation, skill trajectory validation, or readiness scoring for a skill workflow.

소스 정보

저장소
grahama1970/agent-skills
최근 소스 활동
2026년 8월 12일 16:20
감지된 SKILL.md 언어
영어
스타
5
포크
2

설치 방법

기본적으로 소스를 먼저 확인하는 Prompt가 선택됩니다. 직접 명령으로 전환하거나 로컬 사본을 다운로드할 수도 있습니다.

소스 파일 검토

설치 여부를 결정하기 전에 SKILL.md와 SkillsMP에 표시된 보조 파일을 읽어 보세요.

파일 탐색기
10 개 파일

SKILL.md 표시 중

SKILL.md
소스 지침 · 읽기 전용 미리보기
name
agentic-evals
description
Agentic evaluation of skills using multi-trial fixtures, deterministic command assertions, trajectory checks, safety constraints, and evidence-backed readiness scoring. Use when users ask for agentic evals, multi-trial skill evaluation, skill trajectory validation, or readiness scoring for a skill workflow.
triggers
["agentic evals","agentic evaluation","multi trial skill evaluation","skill trajectory validation","readiness scoring","evaluate agent workflow"]
runtime_self_improvement
basic
provides
["agentic-evaluation","multi-trial-evaluation","readiness-scoring","trajectory-validation-pattern"]
composes
["eval-skills"]
complies
["best-practices-skills","best-practices-python"]
taxonomy
["validation","resilience","precision"]
disciplines
["evaluation-quality","agentic-orchestration"]
# agentic-evals Use this skill when a normal one-shot smoke test is too weak and the task needs repeatable, evidence-backed evaluation of a skill or agent workflow. ## Current Scope This initial bundle provides a deterministic fixture runner for command-based cases. It runs each case multiple times, records stdout/stderr/exit status and duration, checks explicit expectations, and emits a machine-readable readiness summary. This proves only the declared fixture behavior. It does not prove semantic correctness, real service integration, LLM-judge quality, or release readiness unless the fixture commands themselves exercise those live paths. ## Usage ```bash ./run.sh run fixtures/agentic_eval.json ./run.sh run fixtures/agentic_eval.json --output /tmp/agentic-evals-report.json ./run.sh audit-skills ../ --output /tmp/agentic-evals-baseline-gap-report.json ./run.sh scaffold-fixture ../some-skill --output ../some-skill/fixtures/agentic_eval.json ./run.sh apply-scaffolds ../ --write --output /tmp/agentic-evals-apply-scaffolds.json ``` ## Fixture Contract ```json { "version": 2, "skill": "example-skill", "trials": 3, "proof_scope": "fixture wiring smoke", "claims": { "proves": "the declared command exits with the expected status", "does_not_prove": "semantic correctness, live service behavior, or full skill readiness" }, "cases": [ { "name": "happy-path", "type": "positive", "command": ["echo", "success"], "expected": { "exit_code": 0, "stdout_contains": ["success"] } } ] } ``` Each case must declare: - `name` - `type`: `positive`, `negative`, or `adversarial` - `command`: a non-empty argv list - `expected.exit_code` Optional expectations: - `expected.stdout_contains` - `expected.stderr_contains` ## Anti-Slop Contract (fail-closed) A skill evaluation fixture is REJECTED at load time (non-zero exit, no run) if it is self-serving deterministic plumbing rather than real-world proof. To pass, a skill fixture MUST: - set `trials` >= 2 (a single trial is not evidence); - include at least one `negative` or `adversarial` case (an all-positive fixture is self-serving); - include at least one **real-world** case: `"real_world": true` whose command exercises a live path (the skill's `run.sh` / a script / live HTTP / a test runner) and does NOT feed itself `fixtures/` stub inputs; - contain no trivial `echo`/constant cases that prove nothing. Rejection message names every violation. This prevents an eval that passes trivially while proving nothing about whether the skill actually works. ### Compliance tier (`"eval_tier": "compliance"`) A fixture that guards a compliance-pipeline stage declares `"eval_tier": "compliance"` and the runner then MANDATES the strong contract on top of the baseline (operator directive 2026-08-12, "this is a compliance pipeline and must be robustly hardened"). Such a fixture is REJECTED unless: - a **strict majority** of cases are `adversarial`/`negative` (more than half, not exactly half — positive controls are the minority); - at least one case is **non-deterministic**: its command samples fresh inputs each run via `--samples`, `--seed`, or a shell `$RANDOM` (a probe *script* name with a fixed key does not count); - every non-deterministic case names `--samples` >= 50, so each stage's coverage is hundreds-to-thousands of assertions per run, targeting ~1000 per stage across its modes. The declaration cannot be quietly relaxed: the compliance pipeline's own fixtures set the tier, so removing it to dodge the gate is itself a regression. `tests/test_compliance_tier_gate.py` pins each rule against its weakening. Two honestly-declared exemptions bypass the gate — never valid for a real skill evaluation: - `"eval_kind": "runner_selftest"` — a fixture that tests this runner itself or is a documentation example. - `"eval_kind": "scaffold"` — the mechanical first-posture fixture emitted by `scaffold-fixture` / `apply-scaffolds`, which the audit still flags as needing real cases. ## Readiness Mapping - `READY`: every case passes every trial. - `USABLE_WITH_GAPS`: at least one trial passes and at least one trial fails. - `NOT_READY`: trials ran but no case fully passed. - `NOT_ESTABLISHED`: no cases were executed. ## Composition `agentic-evals` composes with `eval-skills`: use `eval-skills` for the existing repository fixture schema and broad skill regression checks; use `agentic-evals` when the evaluation needs repeated trials, trajectory-oriented case typing, and readiness-state output. Use `audit-skills` after changing `best-practices-skills` eval rules. The audit does not prove per-skill behavior; it proves the repository's current eval posture by recording which skills already have fixtures, delegate to eval skills, document `eval_not_required`, or still emit `EVAL001`. Use `scaffold-fixture` only as the first mechanical eval posture for a skill. A generated fixture proves wiring only until a human or maintainer adds skill-specific positive, negative, and adversarial cases. Use `apply-scaffolds` to apply that first mechanical posture across all currently scaffoldable `EVAL001` skills. For skills with `sanity.sh` or `run.sh`, it creates an entrypoint-backed fixture. For skills without an entrypoint, it creates a static contract-validation fixture that runs the `best-practices-skills` validator from the skill's `fixtures/` directory. It writes only missing `fixtures/agentic_eval.json` files unless `--force` is passed and emits a JSON receipt. This reduces missing eval posture; it does not establish semantic coverage. ## Regression Fixture Pattern When a live incident exposes an agent-troubleshooting failure, add or strengthen the affected skill's committed `fixtures/agentic_eval.json` instead of leaving the lesson only in chat. The case should name the failure code, exercise the real skill entrypoint, script, or test runner, and assert the recovery wording or receipt fields an agent must see. Example pattern for browser transport incidents: - `type: "adversarial"` for stale sockets, stale tab bindings, lock contention, missing native host dependencies, or provider payload mismatch. - `real_world: true` when the command invokes `run.sh`, a real script, live HTTP, or the production test runner without feeding itself `fixtures/` stubs. - `expected.stderr_contains` or `expected.stdout_contains` should include the stable failure code such as `stale_socket_no_listener`, not only a generic timeout or nonzero exit. This keeps agentic evals tied to the operational mistake future project agents need to recognize.
GitHub에서 보기