| name | webmcp-eval-triage |
| description | Triage failing WebArena or WebMCP tasks. Use to classify tool, eval, agent, infrastructure, and drift failures. |
webmcp-eval-triage
Use this skill before changing tools, evals, prompts, or benchmark operations in
response to failures. The first job is to classify the real owner of the
failure.
Required References
Read the references that match the failure:
- failure-rubric.md: fault buckets, compact
critique codes, stale-eval decisions, and WebMCP tool-shape failure modes.
- log-and-artifact-triage.md: logs,
result directories, exact error grouping, long-run drift, and valid-row rules.
Triage Loop
- Read the failing task definition, eval definition, run output, tool logs, and
result artifacts.
- Bucket the failure: tool, stale eval, agent executor, infrastructure, or
long-run drift.
- Check whether the pattern is one task, one site, one tool family, one arm, or
a broad stack failure.
- Use direct runtime evidence when available. A reward score alone is weaker
evidence than tool execution and persisted state.
- Treat invalid evidence as a diagnostic signal. Name the owner, pause only the
smallest poisoned label, slice, site, worker, or stack, and identify what the
bad evidence implies about the system.
- Continue unaffected comparable work only when provenance remains clean. Stop
the whole run only when continuing would corrupt evidence, violate an
explicit user stop, require unavailable authority, or spend uncontrolled
resources.
Output
For each task or cluster, report the bucket, evidence, likely owner, suggested
fix, and whether human sign-off is needed before changing eval expectations.