Skip to main content

agentic-evals

Agentic evaluation of skills using multi-trial fixtures, deterministic command assertions, trajectory checks, safety constraints, and evidence-backed readiness scoring. Use when users ask for agentic evals, multi-trial skill evaluation, skill trajectory validation, or readiness scoring for a skill workflow.

Ir a la instalación

Datos de origen

Repositorio
grahama1970/agent-skills
Última actividad en el origen
12 de agosto de 2026 a las 16:20
Idioma detectado de SKILL.md
inglés
Estrellas
5
Forks
2

Opciones de instalación

De forma predeterminada está seleccionado el prompt que primero revisa el origen. Puedes cambiar a un comando directo o descargar una copia local.

Revisa los archivos de origen

Lee SKILL.md y los archivos complementarios que muestra SkillsMP antes de decidir si quieres instalarlo.

Explorador de archivos
10 archivos

Mostrando SKILL.md

SKILL.md
Instrucciones de origen · Vista previa de solo lectura
name
agentic-evals
description
Agentic evaluation of skills using multi-trial fixtures, deterministic command assertions, trajectory checks, safety constraints, and evidence-backed readiness scoring. Use when users ask for agentic evals, multi-trial skill evaluation, skill trajectory validation, or readiness scoring for a skill workflow.
triggers
["agentic evals","agentic evaluation","multi trial skill evaluation","skill trajectory validation","readiness scoring","evaluate agent workflow"]
runtime_self_improvement
basic
provides
["agentic-evaluation","multi-trial-evaluation","readiness-scoring","trajectory-validation-pattern"]
composes
["eval-skills"]
complies
["best-practices-skills","best-practices-python"]
taxonomy
["validation","resilience","precision"]
disciplines
["evaluation-quality","agentic-orchestration"]
# agentic-evals Use this skill when a normal one-shot smoke test is too weak and the task needs repeatable, evidence-backed evaluation of a skill or agent workflow. ## Current Scope This initial bundle provides a deterministic fixture runner for command-based cases. It runs each case multiple times, records stdout/stderr/exit status and duration, checks explicit expectations, and emits a machine-readable readiness summary. This proves only the declared fixture behavior. It does not prove semantic correctness, real service integration, LLM-judge quality, or release readiness unless the fixture commands themselves exercise those live paths. ## Usage ```bash ./run.sh run fixtures/agentic_eval.json ./run.sh run fixtures/agentic_eval.json --output /tmp/agentic-evals-report.json ./run.sh audit-skills ../ --output /tmp/agentic-evals-baseline-gap-report.json ./run.sh scaffold-fixture ../some-skill --output ../some-skill/fixtures/agentic_eval.json ./run.sh apply-scaffolds ../ --write --output /tmp/agentic-evals-apply-scaffolds.json ``` ## Fixture Contract ```json { "version": 2, "skill": "example-skill", "trials": 3, "proof_scope": "fixture wiring smoke", "claims": { "proves": "the declared command exits with the expected status", "does_not_prove": "semantic correctness, live service behavior, or full skill readiness" }, "cases": [ { "name": "happy-path", "type": "positive", "command": ["echo", "success"], "expected": { "exit_code": 0, "stdout_contains": ["success"] } } ] } ``` Each case must declare: - `name` - `type`: `positive`, `negative`, or `adversarial` - `command`: a non-empty argv list - `expected.exit_code` Optional expectations: - `expected.stdout_contains` - `expected.stderr_contains` ## Anti-Slop Contract (fail-closed) A skill evaluation fixture is REJECTED at load time (non-zero exit, no run) if it is self-serving deterministic plumbing rather than real-world proof. To pass, a skill fixture MUST: - set `trials` >= 2 (a single trial is not evidence); - include at least one `negative` or `adversarial` case (an all-positive fixture is self-serving); - include at least one **real-world** case: `"real_world": true` whose command exercises a live path (the skill's `run.sh` / a script / live HTTP / a test runner) and does NOT feed itself `fixtures/` stub inputs; - contain no trivial `echo`/constant cases that prove nothing. Rejection message names every violation. This prevents an eval that passes trivially while proving nothing about whether the skill actually works. ### Compliance tier (`"eval_tier": "compliance"`) A fixture that guards a compliance-pipeline stage declares `"eval_tier": "compliance"` and the runner then MANDATES the strong contract on top of the baseline (operator directive 2026-08-12, "this is a compliance pipeline and must be robustly hardened"). Such a fixture is REJECTED unless: - a **strict majority** of cases are `adversarial`/`negative` (more than half, not exactly half — positive controls are the minority); - at least one case is **non-deterministic**: its command samples fresh inputs each run via `--samples`, `--seed`, or a shell `$RANDOM` (a probe *script* name with a fixed key does not count); - every non-deterministic case names `--samples` >= 50, so each stage's coverage is hundreds-to-thousands of assertions per run, targeting ~1000 per stage across its modes. The declaration cannot be quietly relaxed: the compliance pipeline's own fixtures set the tier, so removing it to dodge the gate is itself a regression. `tests/test_compliance_tier_gate.py` pins each rule against its weakening. Two honestly-declared exemptions bypass the gate — never valid for a real skill evaluation: - `"eval_kind": "runner_selftest"` — a fixture that tests this runner itself or is a documentation example. - `"eval_kind": "scaffold"` — the mechanical first-posture fixture emitted by `scaffold-fixture` / `apply-scaffolds`, which the audit still flags as needing real cases. ## Readiness Mapping - `READY`: every case passes every trial. - `USABLE_WITH_GAPS`: at least one trial passes and at least one trial fails. - `NOT_READY`: trials ran but no case fully passed. - `NOT_ESTABLISHED`: no cases were executed. ## Composition `agentic-evals` composes with `eval-skills`: use `eval-skills` for the existing repository fixture schema and broad skill regression checks; use `agentic-evals` when the evaluation needs repeated trials, trajectory-oriented case typing, and readiness-state output. Use `audit-skills` after changing `best-practices-skills` eval rules. The audit does not prove per-skill behavior; it proves the repository's current eval posture by recording which skills already have fixtures, delegate to eval skills, document `eval_not_required`, or still emit `EVAL001`. Use `scaffold-fixture` only as the first mechanical eval posture for a skill. A generated fixture proves wiring only until a human or maintainer adds skill-specific positive, negative, and adversarial cases. Use `apply-scaffolds` to apply that first mechanical posture across all currently scaffoldable `EVAL001` skills. For skills with `sanity.sh` or `run.sh`, it creates an entrypoint-backed fixture. For skills without an entrypoint, it creates a static contract-validation fixture that runs the `best-practices-skills` validator from the skill's `fixtures/` directory. It writes only missing `fixtures/agentic_eval.json` files unless `--force` is passed and emits a JSON receipt. This reduces missing eval posture; it does not establish semantic coverage. ## Regression Fixture Pattern When a live incident exposes an agent-troubleshooting failure, add or strengthen the affected skill's committed `fixtures/agentic_eval.json` instead of leaving the lesson only in chat. The case should name the failure code, exercise the real skill entrypoint, script, or test runner, and assert the recovery wording or receipt fields an agent must see. Example pattern for browser transport incidents: - `type: "adversarial"` for stale sockets, stale tab bindings, lock contention, missing native host dependencies, or provider payload mismatch. - `real_world: true` when the command invokes `run.sh`, a real script, live HTTP, or the production test runner without feeding itself `fixtures/` stubs. - `expected.stderr_contains` or `expected.stdout_contains` should include the stable failure code such as `stale_socket_no_listener`, not only a generic timeout or nonzero exit. This keeps agentic evals tied to the operational mistake future project agents need to recognize.
Ver en GitHub