Skip to main content

multi-model-validation

Runs the same task across multiple AI models in parallel and aggregates verdicts. Use when the user wants a second opinion, multi-expert validation, or consensus from Grok, Gemini, GPT-5, or Kimi.

Datos de origen

Repositorio
MadAppGang/magus
Última actividad en el origen
15 de septiembre de 2026 a las 02:45
Idioma detectado de SKILL.md
inglés
Estrellas
10
Forks
4

Opciones de instalación

De forma predeterminada está seleccionado el prompt que primero revisa el origen. Puedes cambiar a un comando directo o descargar una copia local.

Revisa los archivos de origen

Lee SKILL.md y los archivos complementarios que muestra SkillsMP antes de decidir si quieres instalarlo.

Mostrando SKILL.md

SKILL.md
Instrucciones de origen · Vista previa de solo lectura
name
multi-model-validation
description
Runs the same task across multiple AI models in parallel and aggregates verdicts. Use when the user wants a second opinion, multi-expert validation, or consensus from Grok, Gemini, GPT-5, or Kimi.
user-invocable
false
# Multi-Model Validation **Version:** 3.3.0 **Purpose:** Patterns for running multiple AI models in parallel via Claudish proxy with **context-aware preferences**, dynamic model discovery, session-based workspaces, and performance statistics **Status:** Production Ready ## Overview Multi-model validation is the practice of running multiple AI models (Grok, Gemini, GPT-5, DeepSeek, etc.) in parallel to validate code, designs, or implementations from different perspectives. This achieves: - **3-5x speedup** via parallel execution (15 minutes → 5 minutes) - **Consensus-based prioritization** (issues flagged by all models are CRITICAL) - **Diverse perspectives** (different models catch different issues) - **Cost transparency** (know before you spend) - **Free model discovery** (NEW v3.0) - find high-quality free models from trusted providers - **Performance tracking** - identify slow/failing models for future exclusion - **Data-driven recommendations** - optimize model shortlist based on historical performance **Key Innovations:** 1. **Context-Aware Preferences** (NEW v3.3.0) - Automatically use saved model preferences per task type (debug/research/coding/review) from `.claude/multimodel-team.json` 2. **Dynamic Model Discovery** (v3.0) - Read the live catalog (`list_models`) for current available models (live, 24h cache) 3. **Session-Based Workspaces** (v3.0) - Each validation session gets a unique directory to prevent conflicts 4. **4-Message Pattern** - Ensures true parallel execution by using only Agent tool calls in a single message 5. **Pattern 7-8** - Statistics collection and data-driven model recommendations This skill is extracted from the `/review` command and generalized for use in any multi-model workflow. --- ## ⚠️ MANDATORY: Learn and Reuse User Preferences > **Model preferences are learned per context and reused automatically.** > > - First time a context is used → ASK user → SAVE to that context > - Next time same context → VALIDATE the saved IDs against `list_models`, then use > the survivors automatically (no asking). "No asking" applies to the *selection*, > never to the catalog check — saved IDs go stale and must be re-checked every run. > - User explicitly says "change models" or "different models" → ASK and UPDATE ```bash # FIRST STEP - Read preferences file cat .claude/multimodel-team.json 2>/dev/null ``` **Flow:** ``` 1. Detect context from task keywords - "debug", "error", "bug", "fix" → debug - "research", "analyze", "investigate" → research - "implement", "build", "create", "code" → coding - "review", "audit", "check" → review 2. Check if contextPreferences[context] exists and is non-empty IF EXISTS (has models saved): → Call: list_models (claudish MCP) and KEEP ONLY the saved IDs it still lists Saved preferences are user policy, not a catalog snapshot — they go stale silently. This applies to defaultModels and contextPreferences alike; see claudish:claudish-usage → "Every field of the preferences file is untrusted" → Name every dropped ID in your reply → DO NOT ask the user to re-pick while at least one saved ID survives → If NOTHING survives, say so and offer live alternatives IF EMPTY/MISSING (first time for this context): → Call: list_models (claudish MCP — current models, pricing, capabilities) → Ask user to select models (AskUserQuestion) → Save to contextPreferences[context] → Proceed with validation 3. User override triggers (explicit request to change): - "use different models" - "change models" - "update model preferences" → Ask user to select new models → Update contextPreferences[context] ``` **Example - Learning Flow:** ``` # First debug task ever: Task: "Debug this authentication error" → Context: debug → contextPreferences.debug is empty → ASK: "Which models for debug tasks?" → User selects: grok, glm, minimax → SAVE to contextPreferences.debug → Run with those models # Second debug task: Task: "Debug the API timeout" → Context: debug → contextPreferences.debug = ["grok", "glm", "minimax"] → USE directly (no asking) → Run with saved models # User wants to change: Task: "Debug this error, use different models" → Detected: "different models" override trigger → ASK: "Which models for debug tasks?" → User selects: gemini, LATEST_GPT_MODEL → UPDATE contextPreferences.debug → Run with new models ``` --- ## Related Skills > **CRITICAL: Tracking Protocol Required** > > Before using any patterns in this skill, ensure you have completed the > pre-launch setup from `multimodel:model-tracking-protocol`. > > Launching models without tracking setup = INCOMPLETE validation. **Cross-References:** - **multimodel:model-tracking-protocol** - MANDATORY tracking templates and protocols (NEW in v0.6.0) - Pre-launch checklist (8 required items) - Tracking table templates - Failure documentation format - Results presentation template - **multimodel:quality-gates** - Approval gates and severity classification - **multimodel:task-orchestration** - Progress tracking during execution - **multimodel:error-recovery** - Handling failures and retries **Skill Integration:** This skill (`multi-model-validation`) defines **execution patterns** (how to run models in parallel). The `model-tracking-protocol` skill defines **tracking infrastructure** (how to collect and present results). **Use both together:** ```yaml skills: multimodel:multi-model-validation, multimodel:model-tracking-protocol ``` --- ## Core Patterns ### Pattern 0: Session Setup and Model Discovery (NEW v3.0) **Purpose:** Create isolated session workspace and discover available models dynamically. **Why Session-Based Workspaces:** Using a fixed directory like `ai-docs/reviews/` causes problems: - ❌ Multiple sessions overwrite each other's files - ❌ Stale data from previous sessions pollutes results - ❌ Hard to track which files belong to which session Instead, create a **unique session directory** for each validation: ```bash # Generate unique session ID TARGET_SLUG=$(echo "${TASK_NAME:-review}" | tr '[:upper:] ' '[:lower:]-' | sed 's/[^a-z0-9-]//g' | head -c20) SESSION_ID="review-${TARGET_SLUG}-$(date +%Y%m%d-%H%M%S)-$(head -c 4 /dev/urandom | xxd -p)" SESSION_DIR="ai-docs/sessions/${SESSION_ID}" # Create session workspace mkdir -p "$SESSION_DIR" echo "Session: $SESSION_ID" echo "Directory: $SESSION_DIR" # Example output: # Session: review-auth-impl-20251212-143052-a3f2 # Directory: ai-docs/sessions/review-auth-impl-20251212-143052-a3f2 ``` **Benefits:** - ✅ Each session is isolated (no cross-contamination) - ✅ Traceable - can associate files with a specific session - ✅ Session ID can be used for tracking in statistics - ✅ Parallel sessions don't conflict - ✅ Aligned with the `dev:dev` session pattern - ✅ Committed to git for audit trail (unlike `/tmp/`) > **⚠️ Do NOT use `/tmp/` for session directories.** Files in `/tmp/` are not > traceable, not committable, and parallel runs will overwrite each other. --- **Dynamic Model Discovery:** **NEVER hardcode model lists.** Models change frequently — new ones appear, old ones deprecate, pricing updates. Instead, read the live catalog (`list_models`) for current available models: Call the `list_models` MCP tool (claudish). It returns the current recommended set — model IDs, pricing, context window, capabilities, and the `provider@model` access prefixes — served from claudish’s catalog with a 24-hour cache. For every live variant in one family, call `search_models` with the family name. **Recommended Free Models for Code Review:** | Model | Provider | Context | Capabilities | Why Good | |-------|----------|---------|--------------|----------| | `qwen/LATEST_FREE_CODING_MODEL` | Qwen | 262K | Tools ✓ | Coding-specialized, large context | | `mistralai/LATEST_FREE_CODING_MODEL` | Mistral | 262K | Tools ✓ | Dev-focused, excellent for code | | `qwen/LATEST_FREE_REASONING_MODEL` | Qwen | 131K | Tools ✓ Reasoning ✓ | Massive 235B model, reasoning | **Model Selection Flow (Learn and Reuse):** ``` 1. Read Preferences File → cat .claude/multimodel-team.json → If file NOT exists → create empty one 2. Detect Task Context → Parse task for keywords (case-insensitive): - "debug", "error", "bug", "fix", "trace", "issue" → debug - "research", "investigate", "analyze", "explore", "find" → research - "implement", "build", "create", "code", "develop", "feature" → coding - "review", "audit", "check", "validate", "verify" → review → If no keywords match → context = "default" 3. Check for Override Triggers in User Message → "use different models", "change models", "update preferences" → If found → force_ask = true 4. Load or Learn Models → models = contextPreferences[context] IF models exist AND NOT force_ask: → USE models directly (no asking) → Go to step 6 IF models empty OR force_ask: → Read: the live catalog (list_models) → AskUserQuestion with multiSelect → Save user selection to contextPreferences[context] → Go to step 6 5. Save Updated Preferences → Write .claude/multimodel-team.json → Update lastUpdated timestamp 6. Execute with Models → Launch parallel validation → No further confirmation needed ``` **Context Keywords:** | Context | Keywords | |---------|----------| | debug | debug, error, bug, fix, trace, issue | | research | research, investigate, analyze, explore, find | | coding | implement, build, create, code, develop, feature | | review | review, audit, check, validate, verify | **Override Triggers (force re-selection):** - "use different models" - "change models" - "update model preferences" - "select new models" ### Routing is Claudish's **Send the `id` from `list_models`. Never build an address.** Claudish owns backend selection, credentials and fallback; this repo implements none of it. A prefix/backend/key table used to sit here. It is deleted — it had drifted to the wrong separator (`/` where claudish uses `@`), listed alias env vars as if canonical, covered a third of the providers, and marked models "collision-free" that had since gained a direct provider. A second copy in `claudish-usage` had drifted differently, which is the point: restating claudish's routing anywhere in this repo guarantees two versions of the truth and no way to tell which is stale. If a model will not route, that is a claudish bug — report it with `report_error`. Do not work around it by choosing a different prefix here. **Interactive Model Selection (AskUserQuestion with multiSelect):** **CRITICAL:** Use AskUserQuestion tool with `multiSelect: true` to let users choose models interactively. This provides a better UX than just showing recommendations. ```typescript // Use AskUserQuestion to let user select models AskUserQuestion({ questions: [{ question: "Which external models should validate your code? (Internal Claude reviewer always included)", header: "Models", multiSelect: true, options: [ // Top paid (from the live catalog (list_models) + historical data) { label: "grok ⚡", description: "$0.85/1M | Quality: 87% | Avg: 42s | Fast + accurate" }, { label: "gemini", description: "$7.00/1M | Quality: 91% | Avg: 55s | High accuracy" }, // Free models — filter the list_models result by pricing { label: "qwen/LATEST_FREE_CODING_MODEL 🆓", description: "FREE | Quality: 82% | 262K context | Coding-specialized" }, { label: "mistralai/LATEST_FREE_CODING_MODEL 🆓", description: "FREE | 262K context | Dev-focused, new model" } ] }] }) ``` **Remember Selection for Session:** Store the user's model selection in the session directory so it persists throughout the validation: ```bash # After user selects models, save to session save_session_models() { local session_dir="$1" shift local models=("$@") # Always include internal reviewer echo "claude-embedded" > "$session_dir/selected-models.txt" # Add user-selected models for model in "${models[@]}"; do echo "$model" >> "$session_dir/selected-models.txt" done echo "Session models saved to $session_dir/selected-models.txt" } # Load session models for subsequent operations load_session_models() { local session_dir="$1" cat "$session_dir/selected-models.txt" } # Usage: # After AskUserQuestion returns selected models save_session_models "$SESSION_DIR" "grok" "qwen/LATEST_FREE_CODING_MODEL" # Later in the session, retrieve the selection MODELS=$(load_session_models "$SESSION_DIR") ``` **Session Model Memory Structure:** ``` $SESSION_DIR/ ├── selected-models.txt # User's model selection (persists for session) ├── claude-review.md # Internal review ├── grok-review.md # External review (if selected) ├── qwen-coder-review.md # External review (if selected) └── consolidated-review.md # Final consolidated review ``` **Why Remember the Selection:** 1. **Re-runs**: If validation needs to be re-run, use same models 2. **Consistency**: All phases of validation use identical model set 3. **Audit trail**: Know which models produced which results 4. **Cost tracking**: Accurate cost attribution per session **Always Include Internal Reviewer:** ``` BEST PRACTICE: Always run internal Claude reviewer alongside external models. Why? ✓ FREE (embedded Claude, no API costs) ✓ Fast baseline (usually fastest) ✓ Provides comparison point ✓ Works even if ALL external models fail ✓ Consistent behavior (same model every time) The internal reviewer should NEVER be optional - it's your safety net. In the code-review panels `dev` dispatches — where claudish is optional — it is not a claudish slot. Launch it as its own `Agent(subagent_type: "dev:reviewer", run_in_background: false, …)` in the SAME message as the `team` call, carrying the contract lines — `TARGET: BRANCH`, `FOCUS:`, its own `OUTPUT:` path, and `MODELS:` naming the externals. Never seat it in that team's `models` list. Its return is in your hands: list its OUTPUT file on REVIEWS: whether or not it carries a `**Verdict**:` line — the aggregator counts a file with no verdict as no-verdict, never as one more approval. (`/team` itself, where claudish is a hard dependency, seats it as the `internal` slot instead — Pattern 3 states the rule.) ``` --- ### Pattern 1: The 4-Message Pattern (MANDATORY) This pattern is **CRITICAL** for achieving true parallel execution with multiple AI models. **Why This Pattern Exists:** Claude Code executes tools **sequentially by default** when different tool types are mixed in the same message. To achieve true parallelism, you MUST: 1. Use ONLY one tool type per message 2. Ensure all Agent calls are in a single message 3. Separate preparation (Bash) from execution (Task) from presentation **The Pattern:** ``` Message 1: Preparation (Bash Only) - Create workspace directories
Ver en GitHub
Este SKILL.md es muy grande, por eso SkillsMP muestra aqui solo la primera seccion. Ver en GitHub