- name
- claudish-usage
- description
- Runs models through the claudish MCP tools — team, create_session, run_prompt — and resolves model IDs against the live catalog. Use when the user mentions claudish, OpenRouter, or external AI models.
- user-invocable
- false
# Claudish Usage Skill
**Purpose:** How agents run models — external and native — through the claudish MCP tools
**Status:** Production Ready
## Everything runs through the MCP tools
There is no CLI tier and no "direct usage" tier. **Every model run goes through a claudish
MCP tool** — one model or a panel, external or native, inside `/team` and `/delegate` or
anywhere else. `Bash("claudish ...")` is not a fallback, not a shortcut for a small task,
and not acceptable "just this once".
| You need | Tool |
|---|---|
| A panel of models on one prompt, in parallel | `team` |
| One model, long task, its own tools and working directory | `create_session` |
| One model, one completion, no tools and no session | `run_prompt` |
| What models exist right now | `list_models`, `search_models` |
| Capability comparison between models | `compare_models` |
| Report a failure to claudish's developers | `report_error` |
Two properties are why: MCP sessions run in their own process, so a model's transcript
never enters your context window; and every tool returns a **structured per-slot result**
with status and errors, so a failure is data you can act on rather than text you have to
parse. Both are lost the moment you shell out.
The only CLI that survives is a set of three read-only diagnostics — see
"Diagnostics: the only CLI left". They investigate the runtime. They never run a task.
## Quick Start
### Step 1: Install the runtime (once)
```bash
npm install -g claudish # or: bun add -g claudish
```
This installs the **runtime**, not a way to run tasks. The plugin's `.mcp.json` starts the
MCP server by launching this binary with `--mcp`, so it must be on `PATH` or the tools
below do not exist at all. Restart the session after installing so the server registers.
### Step 2: Find out what models exist right now
```
list_models() → current recommended set, pricing, capabilities, access lines
search_models({ query: "kimi" }) → every live variant in one family
```
Never skip this. Model IDs from memory are dead IDs — see "Model Alias Resolution".
### Step 3: Run something
```
// One completion, no tools, no session lifecycle
run_prompt(model="grok", prompt="Review this diff for security issues")
// A real Claude Code session with tools, in a working directory
create_session(model="grok", prompt=TASK_PROMPT, timeout_seconds=300,
agent="dev:developer", work_dir=WORK_DIR)
→ watch channel events → get_output(session_id)
// A whole panel. `run` STARTS it and returns immediately — it does not wait.
team(mode="run", path=SESSION_DIR, models=["internal", "grok", "gemini"],
input_file=SESSION_DIR + "/input.md", require_pattern="```vote", agent="dev:researcher")
→ poll team(mode="status", path=SESSION_DIR) until settled → read response-NN.md
```
## Model Alias Resolution
All commands that use external models (/team, /delegate, /dev:fix, etc.) MUST resolve model names through this three-step chain before calling claudish.
**The catalog is live, never committed.** Model IDs come from the `list_models` /
`search_models` MCP tools, which claudish serves from its own Firebase-backed
catalog with a 24-hour cache. There is no model-aliases file in this repo, and
you must not resolve model IDs from memory — training data carries dead IDs.
### Three-Step Resolution Chain
```
Step 1: INTERPRET (Claude Code LLM)
User says anything → Claude infers what family/capability they mean
"use Elon's model" → xAI family
"the Google one" → Google family
"kimi3" → Moonshot family, major version 3
"latest gpt" → OpenAI family, newest version
Step 2: RESOLVE (live catalog lookup — list_models / search_models)
Family/intent → an ID that EXISTS in the live catalog right now
xAI family → whatever grok-* the catalog currently lists
"kimi3" → kimi-k3 (matched in the catalog, not guessed)
"latest gpt" → highest gpt-* version the catalog lists
Step 3: ROUTE (Claudish)
Live model ID → correct provider API endpoint
```
### Resolving a model name
1. Call `list_models` first — it is cheap, cached, and returns the current
recommended set with pricing, capabilities and access prefixes.
2. If the user's request isn't covered there, call `search_models` with the
family name (e.g. `search_models("kimi")`) to see every live variant.
3. Check `.claude/multimodel-team.json` → `customAliases` for a user-defined
shorthand. A custom alias always wins on key conflict — but if it maps to an
ID the catalog no longer lists, say so instead of using it silently.
4. `"internal"` and `"default"` select the host Claude tier, and `"opus"`/`"sonnet"`/
`"haiku"` a specific one. They ARE sent to claudish and run through its native
passthrough on the user's own subscription — no API key, no provider prefix, no
translation. They are not catalog IDs, so `list_models` will not list them and a
catalog check must not reject them. Requires `claudish >= 8.0.0` — native names have
been runnable since 7.65.0, but the `team` contract in this skill needs 8.0.0.
### Use the resolver — do not do this by hand
`resolve-models.ts` performs the whole check and prints the disclosure. It ships **in this
plugin**, beside this skill, so `${CLAUDE_PLUGIN_ROOT}` resolves to it on every channel
claudish publishes to. Call `list_models` first, then hand it the IDs:
```bash
bun "${CLAUDE_PLUGIN_ROOT}/scripts/resolve-models.ts" \
--catalog "<comma-separated ids from list_models>" [--context review] [--json]
```
It verifies every model-bearing field, drops dead IDs individually, computes
provenance, and emits a receipt — **print that receipt verbatim.** Exit `3` means
nothing survived; `0` means proceed with what it selected.
Doing this in your head is what the rest of this section explains. That is the fallback for
paths the resolver does not cover, and for consumers that depend on claudish without
`multimodel` installed beside it. Measured over 30 benchmark runs, prose alone produced
the disclosure at best 14/15 times; the resolver produces it every time, because it is code.
### Verify every field of the preferences file
**Every field of the preferences file is untrusted.** `customAliases` is not the only place
a dead ID hides, and in practice it is the least likely — a file found in the wild had
`customAliases: {}` and six decommissioned IDs sitting in `defaultModels`. **Verify every
ID you take from this file against the live catalog, whichever field it came from:**
| Field | Verify? |
|---|---|
| `defaultModels` | yes |
| `contextPreferences[*]` | yes |
| `customAliases` values | yes |
No field is exempt. Drop each ID the catalog does not list, name the dropped IDs
in your reply, and **carry on with the survivors.**
**A dead entry invalidates that entry, never the request.** Resolving is the next
step, not a fallback:
- A stale `customAliases` mapping means *the alias* is wrong. If the user named a
version, resolve that intent against the catalog and use what you find —
`kimi3` with a dead `kimi3 → kimi-k2.5` alias still resolves to `kimi-k3` when
the catalog lists it.
- Dead entries in `defaultModels` or `contextPreferences[*]` mean *those entries*
are wrong. Run with whatever survives.
Returning "no models" is correct only when the catalog genuinely offers nothing
that satisfies the request. Refusing a run while a live model sits in the catalog
is the same failure as using a dead one — it just fails in the other direction.
### Report what the check found, not what the file claims
When you report your model choice, **state the result of the catalog check**:
> `3 of 7 saved model IDs are no longer in the catalog: grok-4.20-beta, gpt-5.4, kimi-k2.5`
Report it every run, including when nothing was dropped — `all 5 saved IDs are
still live` is the same disclosure with a different value.
That count is derived from the comparison you just performed, so it **cannot be
silently wrong**. It is the disclosure that matters: the user's question is "are
my models alive and what did you actually use", not "what date is in my file".
### File age is secondary, and `lastUpdated` cannot carry it
If you state an age, take it from the **file's modification time** and say so:
`preferences file modified 12 days ago (filesystem mtime)`.
- **Never present `lastUpdated` as the file's age.** It is declared metadata and
is not maintained by every write path — a file has been seen reporting March
while its own `history[0].date` said July. Quoting it as an age is false
precision.
- If `lastUpdated` and the newest `history[].date` disagree, report
`freshness metadata inconsistent` and name both. Do not pick the newer one.
- If no trustworthy source exists, `freshness unknown` is a complete answer.
- `mtime` has its own limits — a checkout or copy resets it — which is exactly
why the source is always labelled.
**Age never gates.** It never rejects a model (a 157-day-old file whose IDs are
all live is fine — use it) and never approves one (a file written today can be
entirely dead). Catalog membership is what decides; age is context for the human.
### Version intent is a hard constraint, not a hint
When the user names a version — `kimi3`, `gpt-5.6`, `sonnet 5` — that version is
a **requirement**. Resolve it against the catalog and use what you find.
- If the exact version exists → use it.
- If it does not exist → **say so and show the live alternatives.** Ask which
one they want.
- **NEVER** fall back to a lower version because its name is closer as a string.
`kimi3` resolving to `kimi-k3` is a bug, not a near-miss: string distance
cannot tell a version bump from a typo, and silently downgrading a model is
worse than erroring.
### Interpreting User Intent (Step 1)
| User says | Resolve by | Notes |
|---|---|---|
| "grok" | `search_models("grok")` | Take the current flagship, not a remembered ID |
| "Elon's AI" / "xAI model" | `search_models("grok")` | Company association |
| "Google's model" | `search_models("gemini")` | Company association |
| "the cheap one" | `list_models` → Quick picks → Budget | Cost intent |
| "something fast for coding" | `list_models` → Fast variants | Capability intent |
| "biggest context" | `list_models` → Quick picks → Large context | Capability intent |
| "kimi3" | `search_models("kimi")`, require major v3 | Version is a constraint |
| "LATEST_MINIMAX_MODEL" | Verify it's in the catalog, then pass through | Already a full ID |
When uncertain, show the live candidates and ask the user to pick. Listing real
options is always better than guessing one.
### Identity vs routing address
Every catalog record carries the model's **identity** and several **addresses** for
reaching it. Only the identity is the model. Addresses live in sibling fields, so it is
easy to copy the wrong one out of the same record:
| In the record | Example | What it is |
|---|---|---|
| `id` | `kimi-k3` | **the identity — this is the model** |
| `openrouterId` | `moonshotai/kimi-k3` | an address: route via OpenRouter |
| Access line | `kimi@kimi-k3` · `kc@kimi-k3` | addresses: same model, different accounts |
**Bare means no `@` AND no `/`.** `moonshotai/kimi-k3` is not a bare ID — the vendor
slug is a route, not part of the name. Both prefix forms pin the request to one
provider and bypass the subscription-aware backend selection and fallback that passing
`id` gives you. `z-ai/glm-5.2` is as wrong as `gc@glm-5.2`, for the same reason.
- If the user names an address (`cx@LATEST_GPT_MODEL`), pass it through **verbatim**.
- Otherwise pass `id`, and let claudish pick the backend.
- Never assemble an address yourself, and never substitute one field for another —
"the catalog reports it" is not a licence to send it, because the catalog reports
every address too.
### Responsibility Boundaries
| Responsibility | Owner |
|---|---|
| Understanding user intent → family/capability | **Claude Code** (LLM heuristic) |
| Which model IDs exist right now | **Claudish** (live catalog, 24h cache) |
| User custom aliases | **Magus** (`.claude/multimodel-team.json` `customAliases`) |
| Model ID → API endpoint | **Claudish** (provider routing) |
| API keys, backend fallbacks | **Claudish** |
### Rules
- ALWAYS resolve against the live catalog; NEVER from memory or a committed file
- NEVER invent a model ID — if nothing matches, show live options and ask
- NEVER silently downgrade to an older version than the user asked for
- ALWAYS send the catalog's `id`. NEVER send an address (`vendor/model`,
`provider@model`) where a model belongs — not even one the catalog reports, since it
reports `openrouterId` and every Access route alongside `id`. An address goes through
only when the user named it themselves
- User `customAliases` override, but flag any that the catalog no longer lists
## MCP Tool Reference
Parameter names below are verified against claudish 8.0.0. Use them exactly.
### `team` — a panel of models, started in one call
**`run` starts the panel and returns immediately. It does not wait for the models, and
their answers are not in its response.** You start the run, poll `status` until it settles,
then read each slot's answer off disk. Requires **claudish >= 8.0.0**.
```
team(mode, path, models, judges, input, input_file,
require_pattern, min_output_bytes, agent, claude_flags, slot)
```
| Parameter | What it does |
|---|---|
| `mode` | `"run"` · `"status"` · `"cancel"` · `"judge"` · `"run-and-judge"` |
| `path` | Session directory. **Must be within the current working directory.** |
| `models` | The panel. Native names are **ordinary entries** — see below. |
| `judges` | Models that judge the collected responses (`judge` / `run-and-judge`) |
| `input` | The prompt every slot receives, inline |
| `input_file` | The same prompt, read from a file. **Prefer this.** Passing both is a hard error. |
| `require_pattern` | Regex the response MUST match, or the slot is reported FAILED |
| `min_output_bytes` | Floor below which a response counts as empty |
| `agent` | Subagent every child runs as. Applies to EVERY child — there is no per-model form. |
| `claude_flags` | Any other Claude Code flags, space-separated. An `--agent` here loses to `agent`. |
| `slot` | `cancel` only — cancel one slot instead of the whole run |
**There is no `timeout` any more, and passing one is silently ignored.** The schema does
not set `additionalProperties: false`, so a leftover `timeout=180` raises no error — it
just does nothing, while the call still reads as though it set a deadline. Nothing
terminates a slot on a timer; see "Deciding whether a quiet slot is stuck" below.
**Prefer `input_file` over `input`.** A panel prompt is routinely 100+ lines, and an inline
`input` echoes the whole thing verbatim in the user's terminal, burying every other
argument in the call. Write the prompt to `<path>/input.md` first, then name that file. The
path must be inside the working directory.
**Native names are ordinary slots.** `internal` / `default` select the host Claude tier;
`opus` / `sonnet` / `haiku` / `claude-*` select a specific one. They go in the same `models`
array as Grok and Gemini and run on the user's Claude subscription through the native
passthrough — no API key, no translation. Because they go through the tool, they are
covered by `require_pattern` like every other slot.
**`require_pattern` is the success oracle, not the exit code.** A slot that exits 0 having
never produced the shape the prompt demanded is reported FAILED (state EMPTY, reason
`shape_mismatch`). Set it whenever the prompt mandates an output shape. Omitting it is how
a vote-less response gets counted as a vote.
#### The three-step lifecycle
**Every caller of `mode="run"` follows these three steps.** There is no shortcut that skips
step 2 — a workflow that reads results straight out of the `run` response reads nothing.
**Step 1 — start the run.**
```
team(mode="run", path=SESSION_DIR,
models=[...resolved models, native names included...],
input_file=`${SESSION_DIR}/input.md`,
require_pattern=<regex for the shape the prompt mandates>,
agent=RESOLVED_AGENT)
```
It returns a slot map, not results:
```json
{
"started": true,
"team_session_id": "team-20260827-0015",
"session_path": "/abs/path/to/SESSION_DIR",
"slots": { "gpt-5.6-sol": "01", "grok-4.6": "02", "internal": "03" },
"next": { "status": "...", "cancel": "...", "judge": "..." }
}
```
`slots` maps each display model name to its anonymised slot id. Keep it — that id addresses
everything else on disk for that model:
| Path | Contents |
|---|---|
Voir sur GitHub