Run brainsurgery Axon benchmarks correctly. Use when asked to benchmark models, rerun affected rows, smoke-test fidelity, choose codegen2/runtime2/pipeline2 options, monitor active runs, or create log-dir/stream-csv artifacts under repo-root log/.
Generate benchmark run summaries as 4 markdown tables from axon-benchmark logs. Use when asked to "report", "status", or "show tables" for a run directory containing stream CSV/result logs, including completion/ETA, timing counts, generic-vs-materialized quality comparison, and Axon/HF >= 1.0 rows.
Run and debug Axon stage roundtrip tests. Use when asked to roundtrip parser, resolve, normalize, elaborate, flatten, typecheck2, Graph IR, optimize-ast, or optimize-graph outputs, especially with pytest-xdist parallel execution and weak/strong distinctions.