Structured ideation before specification writing. Transform vague ideas into actionable feature proposals through guided brainstorming — grounded in 2024–2026 research on AI-assisted ideation.
在撰寫規格前進行結構化發想。以 2024–2026 年 AI 輔助發想研究為基礎,透過引導式腦力激盪,將模糊構想轉化為可執行的功能提案。
Implements: XSPEC-296 Brainstorm Quality Standard (BQS v1) — brainstorm v4, layered on
XSPEC-247 brainstorm v3 (Multi-Persona Ensemble + Multi-Critic Convergence).
What changed in v3 | v3 的核心改動: Divergence is no longer a single AI voice racing to a count — it is a persona ensemble (each role reasons via chain-of-thought, in isolation) crossed with diversity lenses. Convergence is no longer one AI scorer plus one devil's-advocate — it is a multi-critic panel plus a (Devil's Advocate + Steelman). This directly targets the strongest finding in the literature: multiple personas beat a single pass, and a single LLM critic is weak and prone to sycophancy.
hard-role rebuttal
v3 把發散從「單一 AI 衝數量」改為persona 集成(每個角色以思維鏈獨立推理)× 多樣性透鏡;把收斂從「單一 AI 評分 + 單一反駁」改為多評審面板 + 硬角色反駁(Devil's Advocate + Steelman)。直接對應文獻最強結論:多 persona 勝過單一 pass,單一 LLM 評審既弱又易諂媚。
What changed in v4 | v4 的核心改動: v3 supplied strong mechanisms but no decidable pass/fail quality gate. v4 layers a time-sequenced quality contract on top — the Brainstorm Quality Standard (BQS v1) — without removing any v3 behaviour. BQS is a four-layer × timeline structure: Layer 0 declares explore/exploit intent; Layer 1 (process, leading, visible during divergence) runs dimensions D1–D4; Layer 2 (product, leading, applied to the Recommended Set only after convergence — ideas with Agg. Score ≥ 3.5, uncapped) runs D5–D8 plus a Seeds column and a contested zone; Layer 3 is an ungoverned Judgment Override. First principle: decisions use leading signals; calibration uses lagging signals — these are ordered in time, not a right/wrong trade-off. The hard sequence gate forbids D5–D8 from being revealed or scored during divergence. See "BQS v1 — Quality Contract" below.
First principle | 第一原理: A brainstorm's output is a hypothesis, not an answer. Whether an idea is good is unknowable at ideation time, so quality can only be judged on leading signals (process + epistemic integrity); the standard's legitimacy is calibrated by lagging signals (later outcomes). Decisions use leading; calibration uses lagging; they are ordered in time, not a right/wrong choice. (This corrects any "use only leading, never lagging" absolutism.)
腦力激盪的產物是假說,不是答案。一個點子好不好在發想當下不可知,所以品質只能用 leading 訊號(過程+認識論完整性)判;標準的正當性靠 lagging 訊號(事後結果)校準。決策用 leading、校準用 lagging,兩者是時間前後,不是對錯取捨。(此處修正任何「只用 leading 不用 lagging」的絕對說法。)
BQS is a four-layer × timeline contract. It is additive to v3 — every v3 flag and mechanism is preserved (see "Backward Compatibility").
At the start, declare the explore/exploit ratio and the bet type (incremental vs barbell long-tail). This layer modulates dimension weights: under an exploit-leaning session, a low D2 (divergence coverage) score is correct and is NOT penalised.
開場宣告本次 explore/exploit 配比 與賭注類型(漸進 vs barbell 長尾)。此層調節維度權重:偏 exploit 的工作階段,D2(發散覆蓋)低分是正確的、不扣分。
Hard sequence gate | 硬序列閘: D5–D8 are forbidden from being revealed or scored during the divergence period. The CONVERGE critics MUST NOT be invoked before the last persona has produced its set.
D5–D8 禁止在發散期揭示或評分;CONVERGE 的 critic 不得在最後一個 persona 產完前被呼叫。
Layer 2 — Product (post-convergence leading, applied to the Recommended Set only) | 第 2 層 — 產物(收斂後 leading,僅施於推薦集)
Evaluative dimensions are confined to after convergence × the Recommended Set only (ideas with Agg. Score ≥ 3.5, uncapped in count) — never as a full gate on every divergence idea (that would retroactively pollute divergence, kill evidence-free future ideas, and create form-filling theatre). | 評判維度限縮在收斂後 × 僅推薦集(Agg. Score ≥ 3.5 的想法,不限筆數),絕不在發散期對全部想法當硬閘(否則回溯污染發散、扼殺無證據的未來想法、製造填表劇場)。
Dim
Oracle
Oracle
D5 Grounding
A [current state / external fact] claim with no file:line/source → fail; a [future / hypothesis] claim needs no grounding, marked [hypothesis]. An external-fact claim is a cross-tier floor (applies even at the creative tier).
A selected idea that does not answer "whose problem / do we actually have it / cost of not doing it" → fail; at least one "not worth doing" elimination required. Attach a lagging registry field: which signal will later validate this judgement?
Two-state: [falsifiable now] or [need to do X first to define falsification]; the latter is routed to a next-step feeding D8 and does NOT count as fail
二態:[現可陳述證偽] 或 [需先做 X 才能定義證偽];後者轉 next-step 餵 D8,不算 fail
D8 Actionability
No next-step decision (including "defer") → fail
無 next-step 裁決(含「暫不做」)→fail
+ Seeds column
When killing an idea, force-store "what was wrong, what real problem it points to"; if ≥1 idea was killed then Seeds ≥1 (only checked for non-empty)
殺想法時強制存「錯在哪、指向什麼真問題」;被殺 ≥1 則 Seeds ≥1(只檢查非空)
+ Contested zone
Ideas whose critic variance exceeds a threshold are surfaced in a split view and not eliminated by mean ranking (protects barbell long-tail)
A reserved space no oracle enters: a human's intuition to "keep / kill" needs only a one-line reason and overrides the aggregate score — it passes through no dimension. This admits the standard is incomplete and avoids "brainstorming to pass dimensions".
Meta stop rule | Meta 停止規則: stop when the layer's dimensions are all green and, after one more round, the Recommended Set membership is unchanged (set membership, not internal ranking). Hard cap: 2 rounds. This replaces the vague "decision didn't flip" criterion. | 該層維度全綠 且 再跑一輪後 推薦集成員不變(看集合成員、不看內部排序)→停。硬上限 2 輪。取代模糊的「不翻決策」。
Judge ≠ generator | 判官≠產生者: D2/D4/D5/D7 judgements need an independent viewpoint; single-context self-evaluation may only be [degraded], never pass. | D2/D4/D5/D7 判定須獨立視角;單 context 自評只能 [degraded],不得 pass。
Calibration loop (consumes lagging) | 校準回路(吃 lagging): v3's three Session Self-Evaluation metrics are folded in as the lagging end — Adoption Rate = D6 lagging validation, Diversity = D2/D3 lagging observation, Cognitive Load = a cost constraint. Two parallel evaluation systems are forbidden — there is one loop, not a separate self-eval. | v3 的 Adoption Rate=D6 滯後驗證、Diversity=D2/D3 滯後觀測、Cognitive Load=成本約束;三者收編為此回路 lagging 端,禁兩套平行評估。
BQS self-evolution (lightweight) | BQS 自我演化(輕量): versioned + an evidence last-reviewed that flags when overdue; "periodically brainstorm BQS itself" is optional. | 版本化 + 證據清單 last-reviewed 逾期 flag;「定期對 BQS 自身 brainstorm」列可選。
Minimal sufficiency | 最小充分原則: apply only the fewest dimensions the tier requires + confining Layer 2 to the Recommended Set (Agg. Score ≥ 3.5) instead of every divergence idea — this is the gate for cognitive economy; no separate cost dimension is added. | 套用該 tier 所需的最少維度 + 把第 2 層限縮在推薦集(Agg. Score ≥ 3.5),而非全部發散想法——認知經濟性的守門,不另設成本維度。
Tiering (bound to v3 objective triggers, not self-assessment) | 分級(綁 v3 既有客觀觸發,非自評)
Tier (from Mode Selection triggers)
BQS dimensions applied
套用維度
creative / --quick
D1–D3 + D5 external-fact floor
D1–D3 + D5 外部事實地板
default
D1–D5 + D8
D1–D5 + D8
strategic / architecture / business
all layers
全層
Tiers are selected by the objective Mode Selection triggers (word count, flags, existence of a spec) — not by the initiator's self-assessment. | 分級由客觀模式選擇觸發表(字數、旗標、是否有規格)決定,而非發起者自評。
Phase 0: PRE-FLIGHT | 防止 AI 錨定
Why this phase exists: Independent ideas written before the AI generates anything consistently produce more diverse results. In AI-assisted contexts this matters more, not less: research on design fixation shows that AI output — being fluent and high-fidelity — deepens fixation rather than relieving it (Wadinambiarachchi et al., CHI 2024).
本階段存在的原因: 在 AI 生成任何內容之前先寫下自己的想法,能持續產出更多樣的結果。在 AI 情境下這更重要:設計固著研究顯示,流暢、高擬真的 AI 輸出反而加深固著(Wadinambiarachchi 等,CHI 2024)。
Before the AI generates any content, the user completes three items:
在 AI 生成任何內容之前,使用者完成三件事:
Item
Prompt
說明
1
One-sentence problem description
一句話描述問題
2
Three initial ideas (any format, any quality)
3 個初始想法(任意形式、不限品質)
3
"Solution types I do NOT want" (N/A allowed)
「我最不想要的解法類型」(可填 N/A)
After the user submits, the AI reads all three inputs and proceeds to FRAME. The AI's first DIVERGE output MUST explore directions the user did not mention and MUST NOT duplicate the user's three ideas.
Anti-seed guardrail (new in v3): Do NOT accept or generate a "like X but for Y" framing as the seed (e.g. "Slack but for doctors"). Such analogical seeds lock the LLM into one solution space and measurably reduce idea variety. Capture the underlying problem, not a product analogy.
反種子 guardrail(v3 新增): 不要用「像 X 但給 Y」的框架當種子(如「給醫生用的 Slack」)。這類產品類比種子會把 LLM 鎖進單一解空間、明顯降低想法多樣性。請捕捉底層問題,而非產品類比。
BQS Layer 0 — Intent declaration (new in v4): Before FRAME, state the explore/exploit ratio and bet type (incremental vs barbell long-tail) for this session. This modulates D2 weighting downstream — under an exploit-leaning session, low divergence coverage is correct and is not penalised. Default when unstated: explore-leaning.
Flag:--skip-preflight bypasses this phase with a one-line warning:
⚠ Skipping Pre-flight may cause AI anchoring
Phase 1: FRAME | 定義問題
Define the problem space clearly before generating ideas.
在產生想法之前,先清楚定義問題空間。
Step
Action
步驟
1
Clarify the problem with 5 Whys
用 5 Whys 釐清問題根因
2
Reframe as "How Might We" (HMW) questions
重構為 HMW 問題
3
Identify stakeholders and constraints
識別利害關係人與限制條件
4
Gather context from codebase (if applicable)
從程式碼庫蒐集脈絡(如適用)
Phase 2: DIVERGE | 發散思考(v3:persona 集成 + 多樣性透鏡)
Core mechanism in v3: a persona ensemble, each persona reasoning via chain-of-thought in isolation, crossed with diversity lenses. Meincke, Mollick & Terwiesch (2024) found that chain-of-thought + personas produces the highest idea diversity of any prompting strategy — close to human groups. A raw idea count is a weak proxy; structurally forcing distinct viewpoints is the real lever.
v3 核心機制:persona 集成——每個 persona 以思維鏈在隔離狀態下推理——再乘上多樣性透鏡。Meincke、Mollick、Terwiesch(2024)發現「思維鏈 + persona」的想法多樣性高於所有受測提示策略,接近人類團體。光衝數量是弱代理;結構性逼出不同視角才是真正槓桿。
Step 2a — Persona ensemble | persona 集成
Generate ideas through a default ensemble of personas. Each persona reasons step by step (chain-of-thought) and produces 2–4 ideas from its own lens only. The user may add, drop, or rename personas via --personas.
透過預設 persona 組生成想法。每個 persona 逐步推理(思維鏈),只從自己的視角產出 2–4 個想法。使用者可用 --personas 增減或改名。
Default persona
Lens it argues from
視角
Domain expert
What does best-practice in this domain demand?
領域最佳實務要求什麼?
Skeptic / risk
Where does this break? What fails first?
哪裡會壞?什麼先失敗?
Cross-domain analogist
How do biology / other fields solve an analogous problem?
生物/他領域如何解類似問題?
Cost / constraint
What is the cheapest, smallest thing that works?
最便宜最小可行解是什麼?
End-user advocate
What does the actual user feel and need?
真實使用者的感受與需求?
Branch isolation: In baseline mode, generate each persona's ideas without showing it the other personas' output — this prevents intra-session anchoring. Present all personas' ideas together only after every persona has produced its set. (In the Enhanced tier, run personas as parallel isolated agents — see "Enhanced Tier" below.)
分支隔離: baseline 模式下,生成每個 persona 的想法時不讓它看到其他 persona 的輸出,以防止 session 內錨定。等所有 persona 都產完才一起呈現。(Enhanced 層以平行隔離 agent 跑——見下方「Enhanced Tier」。)
Step 2b — Diversity lenses | 多樣性透鏡
Apply at least one lens across the ensemble to push past the "obvious answer zone." Connecting disparate concepts measurably increases originality (Mehrotra, Parab & Gulwani, 2024).
在 persona 組上至少套用一個透鏡,以突破「顯而易見答案區」。連結異域概念能可量測地提升原創性(Mehrotra、Parab、Gulwani,2024)。
Lens
Prompt pattern
透鏡
Analogical / cross-domain
"Find a system in [biology / logistics / games] that solves an analogous problem. What can we borrow?"
類比/跨域:借用他領域結構
Assumption reversal
"List what everyone assumes must be true, then invert each one."
假設反轉:列出共識假設並逐一反轉
Morphological matrix
"Build a 3-axis matrix (e.g. User × Trigger × Constraint); fill rare combinations."
形態矩陣:系統性填補罕見組合
Force --lens analogical|reversal|morphological to make a specific lens the primary one.
Step 2c — Continue nudge (auxiliary) | 繼續發散提示(輔助)
The "best ideas appear in the second half" pattern (Nijstad) is a human-group finding and is not confirmed for LLMs (which tend to plateau / exhaust). So a fixed idea-count gate is demoted to an auxiliary nudge: if fewer than ~8 distinct ideas exist across the ensemble, prompt "Continue — add a persona or lens you haven't used." Diversity (distinct lenses covered), not raw count, is the gate.
「好點子在後半」(Nijstad)是人類群體現象,未在 LLM 證實(LLM 多為高原/枯竭)。故固定數量門檻降為輔助提示:若全組少於約 8 個相異想法,提示「繼續——加一個還沒用過的 persona 或透鏡」。真正的門檻是多樣性(覆蓋了幾個不同視角),而非數量。
BQS Layer 1 applies here (D1–D4, leading, visible) | 此處生效 BQS 第 1 層(D1–D4,leading,全程可見): during divergence, only D1 (frame purity), D2 (divergence coverage, weighted by Layer 0), D3 (cross-session diversity), D4 (evaluation de-bias) are scored and shown. Hard sequence gate: the evaluative dimensions D5–D8 must NOT be revealed or scored before the last persona has produced its set — the CONVERGE critics are not invoked until divergence is complete.
Need multiple perspectives (works well as personas)
需要多角度(很適合當 persona)
Phase 3: CONVERGE | 收斂(v3:多評審面板 + 硬角色反駁)
Core mechanism in v3: a multi-critic panel replaces the single weighted scorer. A single LLM is a weak, biased evaluator (Li et al., 2025: LLMs are strong at generation/refinement but weak at evaluation — keep the human as final arbiter). Three independent critic lenses score each idea; their scores are aggregated.
Run three independent critics, each scoring every idea 1–5 on its own lens. Aggregate (mean) to reduce single-critic bias. Each critic uses the weighted formula below.
A soft "please critique this" instruction yields mostly agreement (sycophancy). v3 assigns hard roles: for each idea in the Recommended Set (Agg. Score ≥ 3.5, uncapped), run a Devil's Advocate ("Your job is to argue this idea WILL fail") and a Steelman ("State the strongest charitable version of the counterargument"). Together they stress-test resilience rather than merely poke.
Each counterargument must take the form: "This idea will fail in [specific context] because [specific reason]." Vague concerns ("this might be hard") are rejected.
Flag:--no-rebuttal skips this step; report section marked "Rebuttal: skipped".
BQS D4 — judge ≠ generator | BQS D4 — 判官≠產生者: D4 passes only when the critics/Devil's Advocate run in an independent context (the --enhanced isolated-agent host). On a baseline single context, the panel runs as same-context self-evaluation, which is marked [degraded] and must NOT be marked pass — it is honest but not independent. Do not silently treat a baseline run as a D4 pass.
Step 3c: BQS Layer 2 — product gate on the Recommended Set | BQS 第 2 層——對推薦集的產物閘
After convergence, apply D5–D8 to the Recommended Set only (ideas with Agg. Score ≥ 3.5, uncapped — never to every divergence idea). If no idea reaches 3.5, keep the single highest-scoring idea and mark it [below threshold — shown for reference] so the report is never empty. For each idea in the Recommended Set:
D5 Grounding: split each claim into [current state / external fact] (needs a file:line or source, else fail) vs [future / hypothesis] (no grounding needed, mark [hypothesis]). An external-fact claim is a cross-tier floor — it must be grounded even at the creative tier. | 將每個主張分流為 [現狀/外部事實](需 file:line 或來源,否則 fail)vs [未來/假說](免接地、標 [假說])。外部事實宣稱為跨級地板——creative 級也須接地。
D6 Net benefit: each idea in the Recommended Set must answer "whose problem / do we actually have it / cost of not doing it"; at least one idea must be eliminated as "not worth doing"; attach a lagging registry field (which later signal validates this). | 推薦集中每個想法須答「解誰問題/我們真有嗎/不做的代價」;至少一個淘汰為「不值得做」;掛 lagging 登記欄(事後哪個訊號驗證)。
D7 Falsifiability (two-state): mark [falsifiable now]or[need to do X first to define falsification]. The latter is routed to a next-step that feeds D8 — it does NOT count as fail. | 標 [現可陳述證偽]或[需先做 X 才能定義證偽]。後者轉 next-step 餵 D8——不算 fail。
D8 Actionability: every surviving idea needs a next-step decision (including an explicit "defer / not now"); none → fail. | 每個存活想法需 next-step 裁決(含明確「暫不做」);無 → fail。
Meta stop rule (BQS structural rule 1) | Meta 停止規則: stop iterating when the applied dimensions are all green and one more round leaves the Recommended Set membership unchanged (set membership, not internal ordering). Hard cap: 2 rounds.
當套用的維度全綠 且 再跑一輪後 推薦集成員不變(看集合成員、非內部排序)→停。硬上限 2 輪。
Phase 4: OUTPUT | 輸出提案
Produce a Brainstorm Report ready for /requirement or /sdd. Each surviving idea is marked ✓ Passed rebuttal with a one-line summary of the user's response, its originating persona/lens, and its aggregated critic score.
# Brainstorm Report: [Topic]## Problem Statement
[Refined problem + root cause from FRAME]
## HMW Questions1. How might we ...?
## Ideas Generated
| # | Idea | Persona | Lens | Critic-Feas | Critic-Impact | Critic-Align | Agg. Score |
|---|------|---------|------|-------------|---------------|--------------|-----------|
| 1 | ... | Skeptic | Reversal | 4.0 | 4.5 | 4.0 | 4.2 |
## Recommendations (BQS Layer 2 applied, Agg. Score ≥ 3.5 — uncapped)1.**[Idea]** (Agg. X.X) ✓ Passed rebuttal — [Why] — Persona: [..] — [User rebuttal response]
- D5 grounding: [current-state claims with file:line | future claims marked [hypothesis]]
- D6 net benefit: [whose problem / do we have it / cost of not doing] — lagging signal: [..]
- D7 falsifiability: [falsifiable now | need to do X first → next-step]
- D8 next-step: [action | defer]
2.**[Idea]** (Agg. X.X) ✓ Passed rebuttal — ...
<!-- List every idea with Agg. Score ≥ 3.5, sorted descending — do not cap at 3. If none reach 3.5, list only the single highest-scoring idea marked "[below threshold — shown for reference]". -->
## Contested Zone (high critic-variance ideas)
[Ideas whose critic variance exceeded threshold — surfaced, NOT eliminated by mean ranking (barbell long-tail)]
## Seeds (from killed ideas)
[For each killed idea: what was wrong, what real problem it points to. Non-empty if ≥1 idea was killed.]
## Judgment Override (Layer 3, ungoverned)
[Human keep/kill decisions that override the aggregate score, each with a one-line reason. Optional.]
## Diversity Note
[How many distinct lenses/personas the surviving ideas span — flag if all from one cluster]
## Discarded Ideas (with reasons)
| Idea | Reason |
## Next Steps- [ ] Proceed to `/requirement` with top idea
- [ ] Proceed to `/sdd` if requirements are clear
BQS output additions (v4): the Seeds section (rule: non-empty when ≥1 idea was killed), the Contested Zone (high-variance ideas not eliminated by mean ranking), and the Judgment Override channel (human decision overriding the aggregate score) are required by BQS Layer 2/3. The Recommendations block records the D5–D8 status per idea, for every idea in the Recommended Set (uncapped).
Using a single LLM for ideation reduces the diversity of ideas across users, even when each individual feels more creative (Anderson, Shah & Kreminski, 2024; corroborated by the widely-cited Doshi & Hauser, Science Advances 2024). Guard against it:
Never seed with a competitor or product analogy ("like X but for Y"). | 絕不用競品/產品類比當種子。
Vary the lens, not just the wording — reword ≠ diversify. | 變的是透鏡而非措辭——換句話不等於多樣化。
If the surviving Recommended Set all came from one persona/lens, flag it and run one more lens before OUTPUT. | 若推薦集全來自同一 persona/透鏡,標示並在輸出前再跑一個透鏡。
Enhanced Tier — Parallel Personas | 強化層——平行 persona
Multi-agent ideation (independent agents conversing/contributing) outperforms a single agent on perceived quality and novelty (Quan et al., 2025, MultiColleagues). Where the host supports parallel subagents (e.g. Claude Code's Agent/Workflow tools), --enhanced runs each persona — and each critic — as a parallel, context-isolated agent, then merges and de-duplicates the results.
多 agent 發想(獨立 agent 互相對話/貢獻)在感知品質與新穎度上勝過單 agent(Quan 等,2025,MultiColleagues)。在支援平行子代理的宿主(如 Claude Code 的 Agent/Workflow 工具),--enhanced 會把每個 persona 與每個評審當作平行、context 隔離的 agent 跑,再合併去重。
Graceful degradation: This tier is optional. On hosts without subagents, --enhanced silently falls back to baseline (single-context simulated personas). The skill remains scope: universal.
This is the lagging end of the single BQS quality loop — NOT a second, parallel evaluation system. The leading decision is made by BQS Layers 0–2 during the session; these three metrics calibrate the standard afterwards. Do not run a separate "self-evaluation" alongside BQS.