| name | ax-comparison |
| description | Compare raw and official agent-tool surfaces using replicated AX tasks. |
Compare one controlled task under clearly named conditions, such as public
docs/CLI versus an official Skill, MCP server, or plugin. Hold the fixture,
verifier, resource limits, and prompt intent constant. Randomize or repeat runs
when practical, and record timing, retries, duplicate effects, recovery, and
transcript safety.
Treat differences as agent-usability evidence unless the protocol isolates a
vendor defect. Report non-reproductions and alternative explanations. Do not
generalize from a single successful run.
Attribute each failure to one layer where possible: task/model, harness/session,
tool interface, or external infrastructure. Track human attention and verifier
disagreement alongside completion and wall-clock time.