Skip to main content

self-consistency

For a hard reasoning step or a high-stakes decision, sample several independent solutions and take the consistent answer instead of trusting one shot. Use on tricky bugs, ambiguous design calls, math/logic, or any answer where being wrong is expensive. Trigger with /self-consistency or "sample multiple", "are you sure", "vote on this", "think this through N ways".

Ir para a instalação

Informações da origem

Repositório
Zavelinski/self-consistency
Última atividade na origem
30 de junho de 2026 às 02:51
Idioma detectado do SKILL.md
inglês
Estrelas
0
Forks
0

Opções de instalação

Por padrão, está selecionado o prompt que primeiro revisa a origem. Você pode mudar para um comando direto ou baixar uma cópia local.

Revise os arquivos de origem

Leia o SKILL.md e os arquivos complementares exibidos pelo SkillsMP antes de decidir se vai instalar.

Exibindo SKILL.md

SKILL.md
Instruções da origem · Visualização somente leitura
name
self-consistency
description
For a hard reasoning step or a high-stakes decision, sample several independent solutions and take the consistent answer instead of trusting one shot. Use on tricky bugs, ambiguous design calls, math/logic, or any answer where being wrong is expensive. Trigger with /self-consistency or "sample multiple", "are you sure", "vote on this", "think this through N ways".
version
0.1.0
user-invocable
true
metadata
{"emoji":"🗳️"}
# self-consistency Don't trust a single greedy chain of thought on a hard problem. Generate several INDEPENDENT solutions and take the one they agree on. Agreement across diverse reasoning paths is a strong correctness signal. ## Why this exists (evidence) - Self-consistency (sample diverse reasoning paths, marginalize to the most consistent answer) beats single greedy chain-of-thought by large margins: +17.9pp on GSM8K, +11.0 SVAMP, +12.2 AQuA, +6.4 StrategyQA over the same model. Source: Wang et al., arXiv:2203.11171. - The gain is free of retraining: it is pure test-time compute spent where being wrong is costly. ## When to use - A bug whose root cause is unclear (sample several diagnoses). - A design/architecture call with multiple defensible options. - Math, logic, estimates, migrations, or anything where a confident-but-wrong answer is expensive. - NOT for cheap/obvious steps: it multiplies cost. Use it where the stakes justify the spend. ## The method 1. Restate the problem and the answer format (so samples are comparable). 2. Generate N INDEPENDENT attempts (3-5 typical), each reasoning from scratch and from a different angle. Do not let them see each other. 3. Extract the final answer from each. 4. Take the answer that recurs most (majority / consistency vote). Tie or wide spread = low confidence -> get more samples or escalate to the user. 5. Report the vote: the chosen answer, the count (e.g. 4/5), and the dissent (what the minority said and why it might be right). ## How to run it - Inline: write N short independent analyses yourself, then vote. - Scaled: use a Workflow to fan out N agents on the same prompt with `parallel(...)`, collect their structured answers, and majority-vote in code. Route the sampling to a cheaper model (see model-router); keep the final adjudication sharp. - For code: sample N candidate patches, then prefer the one that passes tests AND recurs across samples. ## Composes with - `model-router`: sample cheap, adjudicate sharp. - `adversarial-verify` / `testsmith`: verify the winning answer, don't just trust the vote. - `run-cost`: N samples = N x cost; budget before fanning out. ## Honest limits - It multiplies token cost by N. Only worth it when an error is more expensive than the extra samples. - Consensus is a heuristic, not proof: all samples can share the same wrong bias. Pair with verification for anything critical. - The cited gains are on reasoning benchmarks; code-task transfer is plausible but measure your own.
Ver no GitHub