Skip to main content

claudish-usage

Runs models through the claudish MCP tools — team, create_session, run_prompt — and resolves model IDs against the live catalog. Use when the user mentions claudish, OpenRouter, or external AI models.

Informations de source

Dépôt
MadAppGang/magus
Dernière activité de la source
15 septembre 2026 à 02:45
Langue détectée de SKILL.md
anglais
Étoiles
10
Forks
4

Options d'installation

Le prompt qui vérifie d'abord la source est sélectionné par défaut. Vous pouvez passer à une commande directe ou télécharger une copie locale.

Vérifiez les fichiers source

Lisez SKILL.md et les fichiers associés affichés par SkillsMP avant de décider de l'installer.

Affichage de SKILL.md

SKILL.md
Instructions source · Aperçu en lecture seule
name
claudish-usage
description
Runs models through the claudish MCP tools — team, create_session, run_prompt — and resolves model IDs against the live catalog. Use when the user mentions claudish, OpenRouter, or external AI models.
user-invocable
false
# Claudish Usage Skill **Purpose:** How agents run models — external and native — through the claudish MCP tools **Status:** Production Ready ## Everything runs through the MCP tools There is no CLI tier and no "direct usage" tier. **Every model run goes through a claudish MCP tool** — one model or a panel, external or native, inside `/team` and `/delegate` or anywhere else. `Bash("claudish ...")` is not a fallback, not a shortcut for a small task, and not acceptable "just this once". | You need | Tool | |---|---| | A panel of models on one prompt, in parallel | `team` | | One model, long task, its own tools and working directory | `create_session` | | One model, one completion, no tools and no session | `run_prompt` | | What models exist right now | `list_models`, `search_models` | | Capability comparison between models | `compare_models` | | Report a failure to claudish's developers | `report_error` | Two properties are why: MCP sessions run in their own process, so a model's transcript never enters your context window; and every tool returns a **structured per-slot result** with status and errors, so a failure is data you can act on rather than text you have to parse. Both are lost the moment you shell out. The only CLI that survives is a set of three read-only diagnostics — see "Diagnostics: the only CLI left". They investigate the runtime. They never run a task. ## Quick Start ### Step 1: Install the runtime (once) ```bash npm install -g claudish # or: bun add -g claudish ``` This installs the **runtime**, not a way to run tasks. The plugin's `.mcp.json` starts the MCP server by launching this binary with `--mcp`, so it must be on `PATH` or the tools below do not exist at all. Restart the session after installing so the server registers. ### Step 2: Find out what models exist right now ``` list_models() → current recommended set, pricing, capabilities, access lines search_models({ query: "kimi" }) → every live variant in one family ``` Never skip this. Model IDs from memory are dead IDs — see "Model Alias Resolution". ### Step 3: Run something ``` // One completion, no tools, no session lifecycle run_prompt(model="grok", prompt="Review this diff for security issues") // A real Claude Code session with tools, in a working directory create_session(model="grok", prompt=TASK_PROMPT, timeout_seconds=300, agent="dev:developer", work_dir=WORK_DIR) → watch channel events → get_output(session_id) // A whole panel. `run` STARTS it and returns immediately — it does not wait. team(mode="run", path=SESSION_DIR, models=["internal", "grok", "gemini"], input_file=SESSION_DIR + "/input.md", require_pattern="```vote", agent="dev:researcher") → poll team(mode="status", path=SESSION_DIR) until settled → read response-NN.md ``` ## Model Alias Resolution All commands that use external models (/team, /delegate, /dev:fix, etc.) MUST resolve model names through this three-step chain before calling claudish. **The catalog is live, never committed.** Model IDs come from the `list_models` / `search_models` MCP tools, which claudish serves from its own Firebase-backed catalog with a 24-hour cache. There is no model-aliases file in this repo, and you must not resolve model IDs from memory — training data carries dead IDs. ### Three-Step Resolution Chain ``` Step 1: INTERPRET (Claude Code LLM) User says anything → Claude infers what family/capability they mean "use Elon's model" → xAI family "the Google one" → Google family "kimi3" → Moonshot family, major version 3 "latest gpt" → OpenAI family, newest version Step 2: RESOLVE (live catalog lookup — list_models / search_models) Family/intent → an ID that EXISTS in the live catalog right now xAI family → whatever grok-* the catalog currently lists "kimi3" → kimi-k3 (matched in the catalog, not guessed) "latest gpt" → highest gpt-* version the catalog lists Step 3: ROUTE (Claudish) Live model ID → correct provider API endpoint ``` ### Resolving a model name 1. Call `list_models` first — it is cheap, cached, and returns the current recommended set with pricing, capabilities and access prefixes. 2. If the user's request isn't covered there, call `search_models` with the family name (e.g. `search_models("kimi")`) to see every live variant. 3. Check `.claude/multimodel-team.json` → `customAliases` for a user-defined shorthand. A custom alias always wins on key conflict — but if it maps to an ID the catalog no longer lists, say so instead of using it silently. 4. `"internal"` and `"default"` select the host Claude tier, and `"opus"`/`"sonnet"`/ `"haiku"` a specific one. They ARE sent to claudish and run through its native passthrough on the user's own subscription — no API key, no provider prefix, no translation. They are not catalog IDs, so `list_models` will not list them and a catalog check must not reject them. Requires `claudish >= 8.0.0` — native names have been runnable since 7.65.0, but the `team` contract in this skill needs 8.0.0. ### Use the resolver — do not do this by hand `resolve-models.ts` performs the whole check and prints the disclosure. It ships **in this plugin**, beside this skill, so `${CLAUDE_PLUGIN_ROOT}` resolves to it on every channel claudish publishes to. Call `list_models` first, then hand it the IDs: ```bash bun "${CLAUDE_PLUGIN_ROOT}/scripts/resolve-models.ts" \ --catalog "<comma-separated ids from list_models>" [--context review] [--json] ``` It verifies every model-bearing field, drops dead IDs individually, computes provenance, and emits a receipt — **print that receipt verbatim.** Exit `3` means nothing survived; `0` means proceed with what it selected. Doing this in your head is what the rest of this section explains. That is the fallback for paths the resolver does not cover, and for consumers that depend on claudish without `multimodel` installed beside it. Measured over 30 benchmark runs, prose alone produced the disclosure at best 14/15 times; the resolver produces it every time, because it is code. ### Verify every field of the preferences file **Every field of the preferences file is untrusted.** `customAliases` is not the only place a dead ID hides, and in practice it is the least likely — a file found in the wild had `customAliases: {}` and six decommissioned IDs sitting in `defaultModels`. **Verify every ID you take from this file against the live catalog, whichever field it came from:** | Field | Verify? | |---|---| | `defaultModels` | yes | | `contextPreferences[*]` | yes | | `customAliases` values | yes | No field is exempt. Drop each ID the catalog does not list, name the dropped IDs in your reply, and **carry on with the survivors.** **A dead entry invalidates that entry, never the request.** Resolving is the next step, not a fallback: - A stale `customAliases` mapping means *the alias* is wrong. If the user named a version, resolve that intent against the catalog and use what you find — `kimi3` with a dead `kimi3 → kimi-k2.5` alias still resolves to `kimi-k3` when the catalog lists it. - Dead entries in `defaultModels` or `contextPreferences[*]` mean *those entries* are wrong. Run with whatever survives. Returning "no models" is correct only when the catalog genuinely offers nothing that satisfies the request. Refusing a run while a live model sits in the catalog is the same failure as using a dead one — it just fails in the other direction. ### Report what the check found, not what the file claims When you report your model choice, **state the result of the catalog check**: > `3 of 7 saved model IDs are no longer in the catalog: grok-4.20-beta, gpt-5.4, kimi-k2.5` Report it every run, including when nothing was dropped — `all 5 saved IDs are still live` is the same disclosure with a different value. That count is derived from the comparison you just performed, so it **cannot be silently wrong**. It is the disclosure that matters: the user's question is "are my models alive and what did you actually use", not "what date is in my file". ### File age is secondary, and `lastUpdated` cannot carry it If you state an age, take it from the **file's modification time** and say so: `preferences file modified 12 days ago (filesystem mtime)`. - **Never present `lastUpdated` as the file's age.** It is declared metadata and is not maintained by every write path — a file has been seen reporting March while its own `history[0].date` said July. Quoting it as an age is false precision. - If `lastUpdated` and the newest `history[].date` disagree, report `freshness metadata inconsistent` and name both. Do not pick the newer one. - If no trustworthy source exists, `freshness unknown` is a complete answer. - `mtime` has its own limits — a checkout or copy resets it — which is exactly why the source is always labelled. **Age never gates.** It never rejects a model (a 157-day-old file whose IDs are all live is fine — use it) and never approves one (a file written today can be entirely dead). Catalog membership is what decides; age is context for the human. ### Version intent is a hard constraint, not a hint When the user names a version — `kimi3`, `gpt-5.6`, `sonnet 5` — that version is a **requirement**. Resolve it against the catalog and use what you find. - If the exact version exists → use it. - If it does not exist → **say so and show the live alternatives.** Ask which one they want. - **NEVER** fall back to a lower version because its name is closer as a string. `kimi3` resolving to `kimi-k3` is a bug, not a near-miss: string distance cannot tell a version bump from a typo, and silently downgrading a model is worse than erroring. ### Interpreting User Intent (Step 1) | User says | Resolve by | Notes | |---|---|---| | "grok" | `search_models("grok")` | Take the current flagship, not a remembered ID | | "Elon's AI" / "xAI model" | `search_models("grok")` | Company association | | "Google's model" | `search_models("gemini")` | Company association | | "the cheap one" | `list_models` → Quick picks → Budget | Cost intent | | "something fast for coding" | `list_models` → Fast variants | Capability intent | | "biggest context" | `list_models` → Quick picks → Large context | Capability intent | | "kimi3" | `search_models("kimi")`, require major v3 | Version is a constraint | | "LATEST_MINIMAX_MODEL" | Verify it's in the catalog, then pass through | Already a full ID | When uncertain, show the live candidates and ask the user to pick. Listing real options is always better than guessing one. ### Identity vs routing address Every catalog record carries the model's **identity** and several **addresses** for reaching it. Only the identity is the model. Addresses live in sibling fields, so it is easy to copy the wrong one out of the same record: | In the record | Example | What it is | |---|---|---| | `id` | `kimi-k3` | **the identity — this is the model** | | `openrouterId` | `moonshotai/kimi-k3` | an address: route via OpenRouter | | Access line | `kimi@kimi-k3` · `kc@kimi-k3` | addresses: same model, different accounts | **Bare means no `@` AND no `/`.** `moonshotai/kimi-k3` is not a bare ID — the vendor slug is a route, not part of the name. Both prefix forms pin the request to one provider and bypass the subscription-aware backend selection and fallback that passing `id` gives you. `z-ai/glm-5.2` is as wrong as `gc@glm-5.2`, for the same reason. - If the user names an address (`cx@LATEST_GPT_MODEL`), pass it through **verbatim**. - Otherwise pass `id`, and let claudish pick the backend. - Never assemble an address yourself, and never substitute one field for another — "the catalog reports it" is not a licence to send it, because the catalog reports every address too. ### Responsibility Boundaries | Responsibility | Owner | |---|---| | Understanding user intent → family/capability | **Claude Code** (LLM heuristic) | | Which model IDs exist right now | **Claudish** (live catalog, 24h cache) | | User custom aliases | **Magus** (`.claude/multimodel-team.json` `customAliases`) | | Model ID → API endpoint | **Claudish** (provider routing) | | API keys, backend fallbacks | **Claudish** | ### Rules - ALWAYS resolve against the live catalog; NEVER from memory or a committed file - NEVER invent a model ID — if nothing matches, show live options and ask - NEVER silently downgrade to an older version than the user asked for - ALWAYS send the catalog's `id`. NEVER send an address (`vendor/model`, `provider@model`) where a model belongs — not even one the catalog reports, since it reports `openrouterId` and every Access route alongside `id`. An address goes through only when the user named it themselves - User `customAliases` override, but flag any that the catalog no longer lists ## MCP Tool Reference Parameter names below are verified against claudish 8.0.0. Use them exactly. ### `team` — a panel of models, started in one call **`run` starts the panel and returns immediately. It does not wait for the models, and their answers are not in its response.** You start the run, poll `status` until it settles, then read each slot's answer off disk. Requires **claudish >= 8.0.0**. ``` team(mode, path, models, judges, input, input_file, require_pattern, min_output_bytes, agent, claude_flags, slot) ``` | Parameter | What it does | |---|---| | `mode` | `"run"` · `"status"` · `"cancel"` · `"judge"` · `"run-and-judge"` | | `path` | Session directory. **Must be within the current working directory.** | | `models` | The panel. Native names are **ordinary entries** — see below. | | `judges` | Models that judge the collected responses (`judge` / `run-and-judge`) | | `input` | The prompt every slot receives, inline | | `input_file` | The same prompt, read from a file. **Prefer this.** Passing both is a hard error. | | `require_pattern` | Regex the response MUST match, or the slot is reported FAILED | | `min_output_bytes` | Floor below which a response counts as empty | | `agent` | Subagent every child runs as. Applies to EVERY child — there is no per-model form. | | `claude_flags` | Any other Claude Code flags, space-separated. An `--agent` here loses to `agent`. | | `slot` | `cancel` only — cancel one slot instead of the whole run | **There is no `timeout` any more, and passing one is silently ignored.** The schema does not set `additionalProperties: false`, so a leftover `timeout=180` raises no error — it just does nothing, while the call still reads as though it set a deadline. Nothing terminates a slot on a timer; see "Deciding whether a quiet slot is stuck" below. **Prefer `input_file` over `input`.** A panel prompt is routinely 100+ lines, and an inline `input` echoes the whole thing verbatim in the user's terminal, burying every other argument in the call. Write the prompt to `<path>/input.md` first, then name that file. The path must be inside the working directory. **Native names are ordinary slots.** `internal` / `default` select the host Claude tier; `opus` / `sonnet` / `haiku` / `claude-*` select a specific one. They go in the same `models` array as Grok and Gemini and run on the user's Claude subscription through the native passthrough — no API key, no translation. Because they go through the tool, they are covered by `require_pattern` like every other slot. **`require_pattern` is the success oracle, not the exit code.** A slot that exits 0 having never produced the shape the prompt demanded is reported FAILED (state EMPTY, reason `shape_mismatch`). Set it whenever the prompt mandates an output shape. Omitting it is how a vote-less response gets counted as a vote. #### The three-step lifecycle **Every caller of `mode="run"` follows these three steps.** There is no shortcut that skips step 2 — a workflow that reads results straight out of the `run` response reads nothing. **Step 1 — start the run.** ``` team(mode="run", path=SESSION_DIR, models=[...resolved models, native names included...], input_file=`${SESSION_DIR}/input.md`, require_pattern=<regex for the shape the prompt mandates>, agent=RESOLVED_AGENT) ``` It returns a slot map, not results: ```json { "started": true, "team_session_id": "team-20260827-0015", "session_path": "/abs/path/to/SESSION_DIR", "slots": { "gpt-5.6-sol": "01", "grok-4.6": "02", "internal": "03" }, "next": { "status": "...", "cancel": "...", "judge": "..." } } ``` `slots` maps each display model name to its anonymised slot id. Keep it — that id addresses everything else on disk for that model: | Path | Contents | |---|---|
Voir sur GitHub
Ce SKILL.md est tres volumineux, SkillsMP affiche donc ici seulement la premiere section. Voir sur GitHub