Skip to main content

multi-model-validation

Runs the same task across multiple AI models in parallel and aggregates verdicts. Use when the user wants a second opinion, multi-expert validation, or consensus from Grok, Gemini, GPT-5, or Kimi.

Informations de source

Dépôt
MadAppGang/magus
Dernière activité de la source
15 septembre 2026 à 02:45
Langue détectée de SKILL.md
anglais
Étoiles
10
Forks
4

Options d'installation

Le prompt qui vérifie d'abord la source est sélectionné par défaut. Vous pouvez passer à une commande directe ou télécharger une copie locale.

Vérifiez les fichiers source

Lisez SKILL.md et les fichiers associés affichés par SkillsMP avant de décider de l'installer.

Affichage de SKILL.md

SKILL.md
Instructions source · Aperçu en lecture seule
name
multi-model-validation
description
Runs the same task across multiple AI models in parallel and aggregates verdicts. Use when the user wants a second opinion, multi-expert validation, or consensus from Grok, Gemini, GPT-5, or Kimi.
user-invocable
false
# Multi-Model Validation **Version:** 3.3.0 **Purpose:** Patterns for running multiple AI models in parallel via Claudish proxy with **context-aware preferences**, dynamic model discovery, session-based workspaces, and performance statistics **Status:** Production Ready ## Overview Multi-model validation is the practice of running multiple AI models (Grok, Gemini, GPT-5, DeepSeek, etc.) in parallel to validate code, designs, or implementations from different perspectives. This achieves: - **3-5x speedup** via parallel execution (15 minutes → 5 minutes) - **Consensus-based prioritization** (issues flagged by all models are CRITICAL) - **Diverse perspectives** (different models catch different issues) - **Cost transparency** (know before you spend) - **Free model discovery** (NEW v3.0) - find high-quality free models from trusted providers - **Performance tracking** - identify slow/failing models for future exclusion - **Data-driven recommendations** - optimize model shortlist based on historical performance **Key Innovations:** 1. **Context-Aware Preferences** (NEW v3.3.0) - Automatically use saved model preferences per task type (debug/research/coding/review) from `.claude/multimodel-team.json` 2. **Dynamic Model Discovery** (v3.0) - Read the live catalog (`list_models`) for current available models (live, 24h cache) 3. **Session-Based Workspaces** (v3.0) - Each validation session gets a unique directory to prevent conflicts 4. **4-Message Pattern** - Ensures true parallel execution by using only Agent tool calls in a single message 5. **Pattern 7-8** - Statistics collection and data-driven model recommendations This skill is extracted from the `/review` command and generalized for use in any multi-model workflow. --- ## ⚠️ MANDATORY: Learn and Reuse User Preferences > **Model preferences are learned per context and reused automatically.** > > - First time a context is used → ASK user → SAVE to that context > - Next time same context → VALIDATE the saved IDs against `list_models`, then use > the survivors automatically (no asking). "No asking" applies to the *selection*, > never to the catalog check — saved IDs go stale and must be re-checked every run. > - User explicitly says "change models" or "different models" → ASK and UPDATE ```bash # FIRST STEP - Read preferences file cat .claude/multimodel-team.json 2>/dev/null ``` **Flow:** ``` 1. Detect context from task keywords - "debug", "error", "bug", "fix" → debug - "research", "analyze", "investigate" → research - "implement", "build", "create", "code" → coding - "review", "audit", "check" → review 2. Check if contextPreferences[context] exists and is non-empty IF EXISTS (has models saved): → Call: list_models (claudish MCP) and KEEP ONLY the saved IDs it still lists Saved preferences are user policy, not a catalog snapshot — they go stale silently. This applies to defaultModels and contextPreferences alike; see claudish:claudish-usage → "Every field of the preferences file is untrusted" → Name every dropped ID in your reply → DO NOT ask the user to re-pick while at least one saved ID survives → If NOTHING survives, say so and offer live alternatives IF EMPTY/MISSING (first time for this context): → Call: list_models (claudish MCP — current models, pricing, capabilities) → Ask user to select models (AskUserQuestion) → Save to contextPreferences[context] → Proceed with validation 3. User override triggers (explicit request to change): - "use different models" - "change models" - "update model preferences" → Ask user to select new models → Update contextPreferences[context] ``` **Example - Learning Flow:** ``` # First debug task ever: Task: "Debug this authentication error" → Context: debug → contextPreferences.debug is empty → ASK: "Which models for debug tasks?" → User selects: grok, glm, minimax → SAVE to contextPreferences.debug → Run with those models # Second debug task: Task: "Debug the API timeout" → Context: debug → contextPreferences.debug = ["grok", "glm", "minimax"] → USE directly (no asking) → Run with saved models # User wants to change: Task: "Debug this error, use different models" → Detected: "different models" override trigger → ASK: "Which models for debug tasks?" → User selects: gemini, LATEST_GPT_MODEL → UPDATE contextPreferences.debug → Run with new models ``` --- ## Related Skills > **CRITICAL: Tracking Protocol Required** > > Before using any patterns in this skill, ensure you have completed the > pre-launch setup from `multimodel:model-tracking-protocol`. > > Launching models without tracking setup = INCOMPLETE validation. **Cross-References:** - **multimodel:model-tracking-protocol** - MANDATORY tracking templates and protocols (NEW in v0.6.0) - Pre-launch checklist (8 required items) - Tracking table templates - Failure documentation format - Results presentation template - **multimodel:quality-gates** - Approval gates and severity classification - **multimodel:task-orchestration** - Progress tracking during execution - **multimodel:error-recovery** - Handling failures and retries **Skill Integration:** This skill (`multi-model-validation`) defines **execution patterns** (how to run models in parallel). The `model-tracking-protocol` skill defines **tracking infrastructure** (how to collect and present results). **Use both together:** ```yaml skills: multimodel:multi-model-validation, multimodel:model-tracking-protocol ``` --- ## Core Patterns ### Pattern 0: Session Setup and Model Discovery (NEW v3.0) **Purpose:** Create isolated session workspace and discover available models dynamically. **Why Session-Based Workspaces:** Using a fixed directory like `ai-docs/reviews/` causes problems: - ❌ Multiple sessions overwrite each other's files - ❌ Stale data from previous sessions pollutes results - ❌ Hard to track which files belong to which session Instead, create a **unique session directory** for each validation: ```bash # Generate unique session ID TARGET_SLUG=$(echo "${TASK_NAME:-review}" | tr '[:upper:] ' '[:lower:]-' | sed 's/[^a-z0-9-]//g' | head -c20) SESSION_ID="review-${TARGET_SLUG}-$(date +%Y%m%d-%H%M%S)-$(head -c 4 /dev/urandom | xxd -p)" SESSION_DIR="ai-docs/sessions/${SESSION_ID}" # Create session workspace mkdir -p "$SESSION_DIR" echo "Session: $SESSION_ID" echo "Directory: $SESSION_DIR" # Example output: # Session: review-auth-impl-20251212-143052-a3f2 # Directory: ai-docs/sessions/review-auth-impl-20251212-143052-a3f2 ``` **Benefits:** - ✅ Each session is isolated (no cross-contamination) - ✅ Traceable - can associate files with a specific session - ✅ Session ID can be used for tracking in statistics - ✅ Parallel sessions don't conflict - ✅ Aligned with the `dev:dev` session pattern - ✅ Committed to git for audit trail (unlike `/tmp/`) > **⚠️ Do NOT use `/tmp/` for session directories.** Files in `/tmp/` are not > traceable, not committable, and parallel runs will overwrite each other. --- **Dynamic Model Discovery:** **NEVER hardcode model lists.** Models change frequently — new ones appear, old ones deprecate, pricing updates. Instead, read the live catalog (`list_models`) for current available models: Call the `list_models` MCP tool (claudish). It returns the current recommended set — model IDs, pricing, context window, capabilities, and the `provider@model` access prefixes — served from claudish’s catalog with a 24-hour cache. For every live variant in one family, call `search_models` with the family name. **Recommended Free Models for Code Review:** | Model | Provider | Context | Capabilities | Why Good | |-------|----------|---------|--------------|----------| | `qwen/LATEST_FREE_CODING_MODEL` | Qwen | 262K | Tools ✓ | Coding-specialized, large context | | `mistralai/LATEST_FREE_CODING_MODEL` | Mistral | 262K | Tools ✓ | Dev-focused, excellent for code | | `qwen/LATEST_FREE_REASONING_MODEL` | Qwen | 131K | Tools ✓ Reasoning ✓ | Massive 235B model, reasoning | **Model Selection Flow (Learn and Reuse):** ``` 1. Read Preferences File → cat .claude/multimodel-team.json → If file NOT exists → create empty one 2. Detect Task Context → Parse task for keywords (case-insensitive): - "debug", "error", "bug", "fix", "trace", "issue" → debug - "research", "investigate", "analyze", "explore", "find" → research - "implement", "build", "create", "code", "develop", "feature" → coding - "review", "audit", "check", "validate", "verify" → review → If no keywords match → context = "default" 3. Check for Override Triggers in User Message → "use different models", "change models", "update preferences" → If found → force_ask = true 4. Load or Learn Models → models = contextPreferences[context] IF models exist AND NOT force_ask: → USE models directly (no asking) → Go to step 6 IF models empty OR force_ask: → Read: the live catalog (list_models) → AskUserQuestion with multiSelect → Save user selection to contextPreferences[context] → Go to step 6 5. Save Updated Preferences → Write .claude/multimodel-team.json → Update lastUpdated timestamp 6. Execute with Models → Launch parallel validation → No further confirmation needed ``` **Context Keywords:** | Context | Keywords | |---------|----------| | debug | debug, error, bug, fix, trace, issue | | research | research, investigate, analyze, explore, find | | coding | implement, build, create, code, develop, feature | | review | review, audit, check, validate, verify | **Override Triggers (force re-selection):** - "use different models" - "change models" - "update model preferences" - "select new models" ### Routing is Claudish's **Send the `id` from `list_models`. Never build an address.** Claudish owns backend selection, credentials and fallback; this repo implements none of it. A prefix/backend/key table used to sit here. It is deleted — it had drifted to the wrong separator (`/` where claudish uses `@`), listed alias env vars as if canonical, covered a third of the providers, and marked models "collision-free" that had since gained a direct provider. A second copy in `claudish-usage` had drifted differently, which is the point: restating claudish's routing anywhere in this repo guarantees two versions of the truth and no way to tell which is stale. If a model will not route, that is a claudish bug — report it with `report_error`. Do not work around it by choosing a different prefix here. **Interactive Model Selection (AskUserQuestion with multiSelect):** **CRITICAL:** Use AskUserQuestion tool with `multiSelect: true` to let users choose models interactively. This provides a better UX than just showing recommendations. ```typescript // Use AskUserQuestion to let user select models AskUserQuestion({ questions: [{ question: "Which external models should validate your code? (Internal Claude reviewer always included)", header: "Models", multiSelect: true, options: [ // Top paid (from the live catalog (list_models) + historical data) { label: "grok ⚡", description: "$0.85/1M | Quality: 87% | Avg: 42s | Fast + accurate" }, { label: "gemini", description: "$7.00/1M | Quality: 91% | Avg: 55s | High accuracy" }, // Free models — filter the list_models result by pricing { label: "qwen/LATEST_FREE_CODING_MODEL 🆓", description: "FREE | Quality: 82% | 262K context | Coding-specialized" }, { label: "mistralai/LATEST_FREE_CODING_MODEL 🆓", description: "FREE | 262K context | Dev-focused, new model" } ] }] }) ``` **Remember Selection for Session:** Store the user's model selection in the session directory so it persists throughout the validation: ```bash # After user selects models, save to session save_session_models() { local session_dir="$1" shift local models=("$@") # Always include internal reviewer echo "claude-embedded" > "$session_dir/selected-models.txt" # Add user-selected models for model in "${models[@]}"; do echo "$model" >> "$session_dir/selected-models.txt" done echo "Session models saved to $session_dir/selected-models.txt" } # Load session models for subsequent operations load_session_models() { local session_dir="$1" cat "$session_dir/selected-models.txt" } # Usage: # After AskUserQuestion returns selected models save_session_models "$SESSION_DIR" "grok" "qwen/LATEST_FREE_CODING_MODEL" # Later in the session, retrieve the selection MODELS=$(load_session_models "$SESSION_DIR") ``` **Session Model Memory Structure:** ``` $SESSION_DIR/ ├── selected-models.txt # User's model selection (persists for session) ├── claude-review.md # Internal review ├── grok-review.md # External review (if selected) ├── qwen-coder-review.md # External review (if selected) └── consolidated-review.md # Final consolidated review ``` **Why Remember the Selection:** 1. **Re-runs**: If validation needs to be re-run, use same models 2. **Consistency**: All phases of validation use identical model set 3. **Audit trail**: Know which models produced which results 4. **Cost tracking**: Accurate cost attribution per session **Always Include Internal Reviewer:** ``` BEST PRACTICE: Always run internal Claude reviewer alongside external models. Why? ✓ FREE (embedded Claude, no API costs) ✓ Fast baseline (usually fastest) ✓ Provides comparison point ✓ Works even if ALL external models fail ✓ Consistent behavior (same model every time) The internal reviewer should NEVER be optional - it's your safety net. In the code-review panels `dev` dispatches — where claudish is optional — it is not a claudish slot. Launch it as its own `Agent(subagent_type: "dev:reviewer", run_in_background: false, …)` in the SAME message as the `team` call, carrying the contract lines — `TARGET: BRANCH`, `FOCUS:`, its own `OUTPUT:` path, and `MODELS:` naming the externals. Never seat it in that team's `models` list. Its return is in your hands: list its OUTPUT file on REVIEWS: whether or not it carries a `**Verdict**:` line — the aggregator counts a file with no verdict as no-verdict, never as one more approval. (`/team` itself, where claudish is a hard dependency, seats it as the `internal` slot instead — Pattern 3 states the rule.) ``` --- ### Pattern 1: The 4-Message Pattern (MANDATORY) This pattern is **CRITICAL** for achieving true parallel execution with multiple AI models. **Why This Pattern Exists:** Claude Code executes tools **sequentially by default** when different tool types are mixed in the same message. To achieve true parallelism, you MUST: 1. Use ONLY one tool type per message 2. Ensure all Agent calls are in a single message 3. Separate preparation (Bash) from execution (Task) from presentation **The Pattern:** ``` Message 1: Preparation (Bash Only) - Create workspace directories
Voir sur GitHub
Ce SKILL.md est tres volumineux, SkillsMP affiche donc ici seulement la premiere section. Voir sur GitHub