Skip to main content

multi-agent-debate

For a contested or high-stakes answer, have several agents propose answers then critique each other across a couple of rounds until they converge, instead of trusting one model or a blind vote. Use on hard design calls, ambiguous root-cause debates, or claims where reasoning quality matters. Trigger with /multi-agent-debate or "debate this", "have agents argue it out", "stress-test this answer".

Source facts

Repository
Zavelinski/multi-agent-debate
Last source activity
June 30, 2026 at 03:09
Detected SKILL.md language
English
Stars
0
Forks
0

Install options

The review-first prompt is selected by default. You can switch to a direct command or download a local copy.

Review the source files

Read SKILL.md and any companion files shown by SkillsMP before deciding whether to install.

Showing SKILL.md

SKILL.md
Source instructions · Read-only preview
name
multi-agent-debate
description
For a contested or high-stakes answer, have several agents propose answers then critique each other across a couple of rounds until they converge, instead of trusting one model or a blind vote. Use on hard design calls, ambiguous root-cause debates, or claims where reasoning quality matters. Trigger with /multi-agent-debate or "debate this", "have agents argue it out", "stress-test this answer".
version
0.1.0
user-invocable
true
metadata
{"emoji":"⚔️"}
# multi-agent-debate Several agents answer independently, then read each other's answers and critique/revise across a round or two, converging on a better-reasoned result. Debate surfaces flaws a single pass (or a blind majority vote) misses. ## Why this exists (evidence) - Multiagent debate (Du et al., arXiv:2305.14325): multiple LLM instances proposing and debating over rounds improves factuality and reasoning vs single-agent and vs self-consistency on several benchmarks; the critique step catches errors that independent sampling alone leaves in. - It differs from self-consistency: there, samples never interact and you majority-vote; here, agents SEE and challenge each other, so a well-argued correction can flip the group. ## When to use - High-stakes design/architecture decisions with multiple defensible options. - Disputed root-cause analysis (two plausible diagnoses). - Claims where the QUALITY of the argument matters, not just the modal answer. - NOT for cheap/clear tasks: debate is the most expensive technique here (N agents x R rounds). ## The method 1. **Propose:** N agents (2-4) independently answer the question with reasoning. 2. **Debate round:** each agent reads the others' answers and either defends, concedes, or revises, citing specific flaws. 1-2 rounds is usually enough. 3. **Converge:** stop when they agree or positions stabilize. Synthesize the converged answer + note any unresolved dissent. 4. **Verify the winner:** debate improves reasoning, not ground truth, confirm critical facts/behavior (compose with cite-guard / adversarial-verify). ## How to run it - Workflow: a propose stage (parallel N agents, structured answers), then a debate stage where each agent gets the others' outputs and revises, then a synthesis stage. Optionally give agents distinct lenses (correctness / security / simplicity) so the debate is diverse, not echo. - Route proposers to a cheaper model; keep the synthesis/judge sharp (model-router). ## Composes with - `self-consistency`: cheaper first pass; escalate to debate only when the vote is split or the call is high-stakes. - `run-cost`: N x R is the priciest pattern here; budget before launching. - `orchestrate`: debate the design at the GATE before building. ## Honest limits - Most expensive technique in this collection (multiplies cost by agents x rounds). Reserve for decisions worth it. - Agents can converge on a confident shared error (groupthink); diverse lenses + a verification step mitigate, not eliminate. - Gains are benchmark-specific; measure your own.
View on GitHub