Use this when a benchmark delta hides churn: one config gains some cells and loses others. The unit is a paired cell: same task, same rep, same model/thinking unless the comparison explicitly changes them.
-
Set the analytical roles. If either side is a local model, read
docs/agents/local-model-analysis.md.
State whether each side is a frontier reference, local subject, local
contrast, or same-model config control. Use a frontier model as a capability
reference, not an expected peer; make capability limits and scaffoldable
failure modes the primary output. Completion: the report's question and
model roles are explicit before metrics are interpreted.
-
Lock the comparison and verify delivery. State left config, right config, subset, reps, model, thinking level, and result roots. Confirm both sides have the intended cells, then verify each config's expected prompt, tool, hook, memory, model-setting, or harness surface from session/config artifacts. Classify delivery as delivered, missing, ambiguous, or leaked; keep intention-to-treat results primary. Audit tool-result errors by tool and cause: distinguish nonzero diagnostic commands, malformed arguments, edit mismatches, read failures, and parser/transport failures instead of treating every isError result as a broken tool. Completion: every pair maps to exactly one left and right result.json, every treatment cell has a delivery classification, and tool-error claims include numerators, denominators, and causes.
-
Split net from churn. Compute left-only solves, right-only solves, both solved, neither solved, mean/median partial delta, token/cost/wall/tool deltas, and difficulty/language splits. Predeclare packet triggers for timeout or negative-reward discordance and material partial/f2p/p2p movement, not only binary flips. Keep observed outcomes primary and add an explicit timeout sensitivity view. Completion: the report shows net solve delta, solve-flip counts, timeout discordance, and the reproducible packet-selection rule.
-
Build trajectory packets and stage ledgers. For every selected cell, gather the paired result.json, session JSONL, model.patch, verifier artifacts, changed-file list, patch stats, tool timeline, and config-specific traces. For local-versus-frontier analysis, add successful exact files read, pre-mutation file coverage, frontier file overlap, file-type focus, repeated reads, and validation timing; keep file discovery separate from file reading. Add a stage ledger from initialization through contract representation, seam location, implementation, targeted and regression validation, completion audit, and termination. Corroborate model claims with commands, patch state, and verifier evidence. Completion: each selected cell has a Markdown or JSON packet that can be reviewed without re-running the benchmark, and the first consequential decision divergence—not merely the first different tool call—is named.
-
Decompose grading before assigning a driver. State f2p and p2p passed/total separately on both sides, including missing grading or denominator changes. Then connect each failing test or reward drop to the concrete patch behavior. Before labeling an outcome infrastructure-caused, require an independent infrastructure signature, treatment linkage, paired or neighboring counterfactual evidence, and a disposition; ambiguous or treatment-linked failures remain observed outcomes. Completion: every classified cell names the feature and preservation effects, failed invariant, exact patch behavior, and evidence-backed disposition; uncertainty is explicit.
-
Classify the driver. Use the smallest specific bucket that fits: wrong seam/layer, under-implementation, over-implementation, missing invariant/guard, protocol/interface drift, cross-scope regression, validation gap, resource exhaustion, or likely variance. Completion: every selected cell has one primary bucket, optional secondary bucket, and evidence bullets.
-
Compare winning and losing patterns. Do not infer skill guidance from losses alone. Run the same packet method on right-only wins, then compare recurring patterns in wins vs losses. Semantic embeddings may prioritize review only: separate prompt-only from outcome/trajectory inputs, exclude same-task reps, and require direct trajectory evidence before promoting a mechanism. Completion: the synthesis separates “keep” patterns from “prevent” patterns and labels embedding findings exploratory.
-
Translate to skill-design hypotheses. Apply writing-for-agents
discipline: propose checkable process changes, not vague advice — each with a
trigger, an action, and a completion criterion. For a local model, build the scaffoldability ledger required by docs/agents/local-model-analysis.md and separate serving, harness, execution-control, repository-understanding, and core-capability failures. Completion: each proposed guidance change has a trigger, an action, a completion criterion, observed cells it could have changed, known counterexamples, and a minimal same-model A/B; single-case proposals remain hypotheses.
-
Publish as an evidence-first report. Deliver per
report-delivery: a self-contained
HTML report served on the Tailnet. Show the complete task × rep outcome table and total trajectory count before any filtered cohort or selected packets; label packet examples as rep-specific rather than task-wide evidence. Include the packet links, bucket table, concrete task examples, and a short conclusion. Separate direct session/harness evidence, patch/verifier evidence, statistical direction, exploratory structure, and interpretation/confidence. For a local model, frame the hero and conclusion around capability shape, frontier gaps, and support experiments rather than winner or product-selection language unless the user asks for a ranking. Completion: the URL works and readers can identify the full denominator, each filtered cohort, and what was observed versus inferred without reconstructing the analysis.