| name | seesaw |
| description | Weighs competing findings or conclusions from a review/comparison output before it's finalized — checks whether the stated conclusion is actually balanced against the evidence on both sides. Triggers after `check` (comparison) or `ship` (ideation) produces a verdict. |
seesaw — balance check
Process
- Isolate the verdict. Find the conclusion/verdict the output actually asserts (e.g. "method A outperforms method B", "this is the most promising direction").
- Weigh both sides explicitly. List the evidence supporting the verdict and, separately, the evidence or conditions that would weaken or reverse it (different eval setup, smaller scale test, cherry-picked benchmark, narrower assumptions).
- Check the tilt. If the supporting side is thin (single benchmark, single paper, favorable-only setup) relative to the counter-evidence, the verdict is overweighted — flag it and soften the claim or add the missing caveat.
- Re-balance the output. Rewrite the verdict to state the conditions under which it holds, not as an unconditional universal claim.
Anti-Rationalization
- "The comparison table speaks for itself" — a table without an explicit statement of what conditions the conclusion depends on invites overgeneralization by the reader.
- "It's obviously better, no need to caveat" — obviousness is not evidence; state the actual scope of the claim.
- "One counterexample doesn't matter" — one counterexample from a fair comparison does matter; note it rather than omitting it.
Evidence
Balance is achieved when: the final verdict states the conditions/scope it applies under, and any known counter-evidence or weaker-comparison caveat is explicitly surfaced rather than omitted.
Red Flags
- A verdict stated as universal ("X is better") when it only holds under one benchmark/setup
- Supporting evidence listed but no counter-evidence considered at all
- Silently dropping a comparison data point that contradicts the stated verdict