- name
- multi-model-validation
- description
- Runs the same task across multiple AI models in parallel and aggregates verdicts. Use when the user wants a second opinion, multi-expert validation, or consensus from Grok, Gemini, GPT-5, or Kimi.
- user-invocable
- false
# Multi-Model Validation
**Version:** 3.3.0
**Purpose:** Patterns for running multiple AI models in parallel via Claudish proxy with **context-aware preferences**, dynamic model discovery, session-based workspaces, and performance statistics
**Status:** Production Ready
## Overview
Multi-model validation is the practice of running multiple AI models (Grok, Gemini, GPT-5, DeepSeek, etc.) in parallel to validate code, designs, or implementations from different perspectives. This achieves:
- **3-5x speedup** via parallel execution (15 minutes → 5 minutes)
- **Consensus-based prioritization** (issues flagged by all models are CRITICAL)
- **Diverse perspectives** (different models catch different issues)
- **Cost transparency** (know before you spend)
- **Free model discovery** (NEW v3.0) - find high-quality free models from trusted providers
- **Performance tracking** - identify slow/failing models for future exclusion
- **Data-driven recommendations** - optimize model shortlist based on historical performance
**Key Innovations:**
1. **Context-Aware Preferences** (NEW v3.3.0) - Automatically use saved model preferences per task type (debug/research/coding/review) from `.claude/multimodel-team.json`
2. **Dynamic Model Discovery** (v3.0) - Read the live catalog (`list_models`) for current available models (live, 24h cache)
3. **Session-Based Workspaces** (v3.0) - Each validation session gets a unique directory to prevent conflicts
4. **4-Message Pattern** - Ensures true parallel execution by using only Agent tool calls in a single message
5. **Pattern 7-8** - Statistics collection and data-driven model recommendations
This skill is extracted from the `/review` command and generalized for use in any multi-model workflow.
---
## ⚠️ MANDATORY: Learn and Reuse User Preferences
> **Model preferences are learned per context and reused automatically.**
>
> - First time a context is used → ASK user → SAVE to that context
> - Next time same context → VALIDATE the saved IDs against `list_models`, then use
> the survivors automatically (no asking). "No asking" applies to the *selection*,
> never to the catalog check — saved IDs go stale and must be re-checked every run.
> - User explicitly says "change models" or "different models" → ASK and UPDATE
```bash
# FIRST STEP - Read preferences file
cat .claude/multimodel-team.json 2>/dev/null
```
**Flow:**
```
1. Detect context from task keywords
- "debug", "error", "bug", "fix" → debug
- "research", "analyze", "investigate" → research
- "implement", "build", "create", "code" → coding
- "review", "audit", "check" → review
2. Check if contextPreferences[context] exists and is non-empty
IF EXISTS (has models saved):
→ Call: list_models (claudish MCP) and KEEP ONLY the saved IDs it still lists
Saved preferences are user policy, not a catalog snapshot — they go stale
silently. This applies to defaultModels and contextPreferences alike; see
claudish:claudish-usage → "Every field of the preferences file is untrusted"
→ Name every dropped ID in your reply
→ DO NOT ask the user to re-pick while at least one saved ID survives
→ If NOTHING survives, say so and offer live alternatives
IF EMPTY/MISSING (first time for this context):
→ Call: list_models (claudish MCP — current models, pricing, capabilities)
→ Ask user to select models (AskUserQuestion)
→ Save to contextPreferences[context]
→ Proceed with validation
3. User override triggers (explicit request to change):
- "use different models"
- "change models"
- "update model preferences"
→ Ask user to select new models
→ Update contextPreferences[context]
```
**Example - Learning Flow:**
```
# First debug task ever:
Task: "Debug this authentication error"
→ Context: debug
→ contextPreferences.debug is empty
→ ASK: "Which models for debug tasks?"
→ User selects: grok, glm, minimax
→ SAVE to contextPreferences.debug
→ Run with those models
# Second debug task:
Task: "Debug the API timeout"
→ Context: debug
→ contextPreferences.debug = ["grok", "glm", "minimax"]
→ USE directly (no asking)
→ Run with saved models
# User wants to change:
Task: "Debug this error, use different models"
→ Detected: "different models" override trigger
→ ASK: "Which models for debug tasks?"
→ User selects: gemini, LATEST_GPT_MODEL
→ UPDATE contextPreferences.debug
→ Run with new models
```
---
## Related Skills
> **CRITICAL: Tracking Protocol Required**
>
> Before using any patterns in this skill, ensure you have completed the
> pre-launch setup from `multimodel:model-tracking-protocol`.
>
> Launching models without tracking setup = INCOMPLETE validation.
**Cross-References:**
- **multimodel:model-tracking-protocol** - MANDATORY tracking templates and protocols (NEW in v0.6.0)
- Pre-launch checklist (8 required items)
- Tracking table templates
- Failure documentation format
- Results presentation template
- **multimodel:quality-gates** - Approval gates and severity classification
- **multimodel:task-orchestration** - Progress tracking during execution
- **multimodel:error-recovery** - Handling failures and retries
**Skill Integration:**
This skill (`multi-model-validation`) defines **execution patterns** (how to run models in parallel).
The `model-tracking-protocol` skill defines **tracking infrastructure** (how to collect and present results).
**Use both together:**
```yaml
skills: multimodel:multi-model-validation, multimodel:model-tracking-protocol
```
---
## Core Patterns
### Pattern 0: Session Setup and Model Discovery (NEW v3.0)
**Purpose:** Create isolated session workspace and discover available models dynamically.
**Why Session-Based Workspaces:**
Using a fixed directory like `ai-docs/reviews/` causes problems:
- ❌ Multiple sessions overwrite each other's files
- ❌ Stale data from previous sessions pollutes results
- ❌ Hard to track which files belong to which session
Instead, create a **unique session directory** for each validation:
```bash
# Generate unique session ID
TARGET_SLUG=$(echo "${TASK_NAME:-review}" | tr '[:upper:] ' '[:lower:]-' | sed 's/[^a-z0-9-]//g' | head -c20)
SESSION_ID="review-${TARGET_SLUG}-$(date +%Y%m%d-%H%M%S)-$(head -c 4 /dev/urandom | xxd -p)"
SESSION_DIR="ai-docs/sessions/${SESSION_ID}"
# Create session workspace
mkdir -p "$SESSION_DIR"
echo "Session: $SESSION_ID"
echo "Directory: $SESSION_DIR"
# Example output:
# Session: review-auth-impl-20251212-143052-a3f2
# Directory: ai-docs/sessions/review-auth-impl-20251212-143052-a3f2
```
**Benefits:**
- ✅ Each session is isolated (no cross-contamination)
- ✅ Traceable - can associate files with a specific session
- ✅ Session ID can be used for tracking in statistics
- ✅ Parallel sessions don't conflict
- ✅ Aligned with the `dev:dev` session pattern
- ✅ Committed to git for audit trail (unlike `/tmp/`)
> **⚠️ Do NOT use `/tmp/` for session directories.** Files in `/tmp/` are not
> traceable, not committable, and parallel runs will overwrite each other.
---
**Dynamic Model Discovery:**
**NEVER hardcode model lists.** Models change frequently — new ones appear, old ones deprecate, pricing updates. Instead, read the live catalog (`list_models`) for current available models:
Call the `list_models` MCP tool (claudish). It returns the current recommended
set — model IDs, pricing, context window, capabilities, and the `provider@model`
access prefixes — served from claudish’s catalog with a 24-hour cache.
For every live variant in one family, call `search_models` with the family name.
**Recommended Free Models for Code Review:**
| Model | Provider | Context | Capabilities | Why Good |
|-------|----------|---------|--------------|----------|
| `qwen/LATEST_FREE_CODING_MODEL` | Qwen | 262K | Tools ✓ | Coding-specialized, large context |
| `mistralai/LATEST_FREE_CODING_MODEL` | Mistral | 262K | Tools ✓ | Dev-focused, excellent for code |
| `qwen/LATEST_FREE_REASONING_MODEL` | Qwen | 131K | Tools ✓ Reasoning ✓ | Massive 235B model, reasoning |
**Model Selection Flow (Learn and Reuse):**
```
1. Read Preferences File
→ cat .claude/multimodel-team.json
→ If file NOT exists → create empty one
2. Detect Task Context
→ Parse task for keywords (case-insensitive):
- "debug", "error", "bug", "fix", "trace", "issue" → debug
- "research", "investigate", "analyze", "explore", "find" → research
- "implement", "build", "create", "code", "develop", "feature" → coding
- "review", "audit", "check", "validate", "verify" → review
→ If no keywords match → context = "default"
3. Check for Override Triggers in User Message
→ "use different models", "change models", "update preferences"
→ If found → force_ask = true
4. Load or Learn Models
→ models = contextPreferences[context]
IF models exist AND NOT force_ask:
→ USE models directly (no asking)
→ Go to step 6
IF models empty OR force_ask:
→ Read: the live catalog (list_models)
→ AskUserQuestion with multiSelect
→ Save user selection to contextPreferences[context]
→ Go to step 6
5. Save Updated Preferences
→ Write .claude/multimodel-team.json
→ Update lastUpdated timestamp
6. Execute with Models
→ Launch parallel validation
→ No further confirmation needed
```
**Context Keywords:**
| Context | Keywords |
|---------|----------|
| debug | debug, error, bug, fix, trace, issue |
| research | research, investigate, analyze, explore, find |
| coding | implement, build, create, code, develop, feature |
| review | review, audit, check, validate, verify |
**Override Triggers (force re-selection):**
- "use different models"
- "change models"
- "update model preferences"
- "select new models"
### Routing is Claudish's
**Send the `id` from `list_models`. Never build an address.** Claudish owns backend
selection, credentials and fallback; this repo implements none of it.
A prefix/backend/key table used to sit here. It is deleted — it had drifted to the wrong
separator (`/` where claudish uses `@`), listed alias env vars as if canonical, covered a
third of the providers, and marked models "collision-free" that had since gained a direct
provider. A second copy in `claudish-usage` had drifted differently, which is the point:
restating claudish's routing anywhere in this repo guarantees two versions of the truth and
no way to tell which is stale.
If a model will not route, that is a claudish bug — report it with `report_error`. Do not
work around it by choosing a different prefix here.
**Interactive Model Selection (AskUserQuestion with multiSelect):**
**CRITICAL:** Use AskUserQuestion tool with `multiSelect: true` to let users choose models interactively. This provides a better UX than just showing recommendations.
```typescript
// Use AskUserQuestion to let user select models
AskUserQuestion({
questions: [{
question: "Which external models should validate your code? (Internal Claude reviewer always included)",
header: "Models",
multiSelect: true,
options: [
// Top paid (from the live catalog (list_models) + historical data)
{
label: "grok ⚡",
description: "$0.85/1M | Quality: 87% | Avg: 42s | Fast + accurate"
},
{
label: "gemini",
description: "$7.00/1M | Quality: 91% | Avg: 55s | High accuracy"
},
// Free models — filter the list_models result by pricing
{
label: "qwen/LATEST_FREE_CODING_MODEL 🆓",
description: "FREE | Quality: 82% | 262K context | Coding-specialized"
},
{
label: "mistralai/LATEST_FREE_CODING_MODEL 🆓",
description: "FREE | 262K context | Dev-focused, new model"
}
]
}]
})
```
**Remember Selection for Session:**
Store the user's model selection in the session directory so it persists throughout the validation:
```bash
# After user selects models, save to session
save_session_models() {
local session_dir="$1"
shift
local models=("$@")
# Always include internal reviewer
echo "claude-embedded" > "$session_dir/selected-models.txt"
# Add user-selected models
for model in "${models[@]}"; do
echo "$model" >> "$session_dir/selected-models.txt"
done
echo "Session models saved to $session_dir/selected-models.txt"
}
# Load session models for subsequent operations
load_session_models() {
local session_dir="$1"
cat "$session_dir/selected-models.txt"
}
# Usage:
# After AskUserQuestion returns selected models
save_session_models "$SESSION_DIR" "grok" "qwen/LATEST_FREE_CODING_MODEL"
# Later in the session, retrieve the selection
MODELS=$(load_session_models "$SESSION_DIR")
```
**Session Model Memory Structure:**
```
$SESSION_DIR/
├── selected-models.txt # User's model selection (persists for session)
├── claude-review.md # Internal review
├── grok-review.md # External review (if selected)
├── qwen-coder-review.md # External review (if selected)
└── consolidated-review.md # Final consolidated review
```
**Why Remember the Selection:**
1. **Re-runs**: If validation needs to be re-run, use same models
2. **Consistency**: All phases of validation use identical model set
3. **Audit trail**: Know which models produced which results
4. **Cost tracking**: Accurate cost attribution per session
**Always Include Internal Reviewer:**
```
BEST PRACTICE: Always run internal Claude reviewer alongside external models.
Why?
✓ FREE (embedded Claude, no API costs)
✓ Fast baseline (usually fastest)
✓ Provides comparison point
✓ Works even if ALL external models fail
✓ Consistent behavior (same model every time)
The internal reviewer should NEVER be optional - it's your safety net.
In the code-review panels `dev` dispatches — where claudish is optional — it is not a
claudish slot. Launch it as its own
`Agent(subagent_type: "dev:reviewer", run_in_background: false, …)` in the SAME message
as the `team` call, carrying the contract lines — `TARGET: BRANCH`, `FOCUS:`, its own
`OUTPUT:` path, and `MODELS:` naming the externals. Never seat it in that team's `models`
list. Its return is in your hands: list its OUTPUT file on REVIEWS: whether or not it
carries a `**Verdict**:` line — the aggregator counts a file with no verdict as
no-verdict, never as one more approval. (`/team` itself, where claudish is a hard
dependency, seats it as the `internal` slot instead — Pattern 3 states the rule.)
```
---
### Pattern 1: The 4-Message Pattern (MANDATORY)
This pattern is **CRITICAL** for achieving true parallel execution with multiple AI models.
**Why This Pattern Exists:**
Claude Code executes tools **sequentially by default** when different tool types are mixed in the same message. To achieve true parallelism, you MUST:
1. Use ONLY one tool type per message
2. Ensure all Agent calls are in a single message
3. Separate preparation (Bash) from execution (Task) from presentation
**The Pattern:**
```
Message 1: Preparation (Bash Only)
- Create workspace directories
Voir sur GitHub