Skip to main content

ferret-kernel-system

Use when you need a NEW or optimized MPK CUDA kernel (a per-task `.cuh` under include/mirage/persistent_kernel/tasks/) that must PROVABLY beat a target WITHOUT the kernel-optimizer cheating its own correctness test. This is the frozen-gate Ferret system — a 3-agent orchestration (dispatcher → independent test-writer → optimizer) that makes "the optimizer ships a simplified-math kernel and self-reports a passing test" structurally impossible. Triggers — phrasings like "dispatch ferret to optimize a kernel", "have an agent write a kernel but guarantee it doesn't simplify the math", "generate a kernel the frozen-gate way": the user/main-thread wants a kernel optimized or generated by an agent, especially when correctness/no-simplification matters more than fire-and-forget speed.

Ir para a instalação

Informações da origem

Repositório
mirage-project/mirage
Última atividade na origem
28 de julho de 2026 às 02:04
Idioma detectado do SKILL.md
inglês
Estrelas
2.494
Forks
256

Opções de instalação

Por padrão, está selecionado o prompt que primeiro revisa a origem. Você pode mudar para um comando direto ou baixar uma cópia local.

Revise os arquivos de origem

Leia o SKILL.md e os arquivos complementares exibidos pelo SkillsMP antes de decidir se vai instalar.

Explorador de arquivos
2 arquivos

Exibindo SKILL.md

SKILL.md
Instruções da origem · Visualização somente leitura
name
ferret-kernel-system
description
Use when you need a NEW or optimized MPK CUDA kernel (a per-task `.cuh` under include/mirage/persistent_kernel/tasks/) that must PROVABLY beat a target WITHOUT the kernel-optimizer cheating its own correctness test. This is the frozen-gate Ferret system — a 3-agent orchestration (dispatcher → independent test-writer → optimizer) that makes "the optimizer ships a simplified-math kernel and self-reports a passing test" structurally impossible. Triggers — phrasings like "dispatch ferret to optimize a kernel", "have an agent write a kernel but guarantee it doesn't simplify the math", "generate a kernel the frozen-gate way": the user/main-thread wants a kernel optimized or generated by an agent, especially when correctness/no-simplification matters more than fire-and-forget speed.
# Ferret frozen-gate kernel system — how to invoke + why ## The problem it solves (read first) The OLD ferret wrote its OWN correctness test. It once shipped a DeepSeek-V3 attention kernel with SIMPLIFIED math (theta-10000 rope not YaRN, 1/sqrt(576) not the YaRN mscale scale, head-sum o_proj skipping W_UV, no kv_a_layernorm) and self-reported **cosine 1.0** — against its own simplified reference. Marking its own homework. The refactor takes judging + constraints OUT of the optimizer. ## The 3 agents (in `.claude/agents/`) | Agent | Role | You invoke? | |---|---|---| | **ferret-kernel-agent** (L1 dispatcher) | The ENTRY POINT. Pins ALL constraints, freezes the gate via the test-writer, runs the optimizer in-session, drives Codex review each round, does in-MPK faithful acceptance. | **YES — this is the only one you invoke directly.** | | **ferret-test-writer** (L2a) | Writes the FROZEN, hash-locked gate vs a CANONICAL reference (never re-derived), checking INTERMEDIATE tensors. Spawned by L1, before the optimizer. | No (L1 spawns it) | | **ferret-optimizer** (L2b) | In-session optimizer (replaces `claude -p`), judged ONLY by the frozen gate, can't simplify. Spawned by L1 each round. | No (L1 spawns it) | ## How to invoke (the main-thread → subagent contract) Invoke the **ferret-kernel-agent** subagent (via the Agent tool) with a COMPLETE constraint contract — this is where nothing gets missed and no simplification slips in. Give it, explicitly: 1. **TARGET** — the exact MPK op + what it replaces; the baseline = the kernel being replaced, benched the way MPK calls it (not an external SOTA unless it's the consumer). 2. **REAL-MATH CONTRACT** — enumerate EVERY step the kernel must compute, NO simplification (the test-writer turns this into intermediate checks). 3. **SHAPES** — exact (M,K,N) at the real TP/EP regime (derive from the builder/weights). 4. **PRODUCTION COMPILE FLAGS** — `-rdc=true` / `MPK_FORCE_RDC_TRUE=1`, arch sm_100a, single-stream / no-CUDA-graph / no-cta_group::2. State that FINAL acceptance is the in-MPK faithful build; standalone `-rdc=true` is diagnostic only. 5. **ABI** — the `__device__ task_impl` signature + NS/NE. 6. **CANONICAL REFERENCE SOURCE** — the already-trusted oracle to compare against (the in-MPK task-chain output; the official HF model; an in-tree faithful test). NEVER "let the agent derive it." If the gate already exists (hash-locked in the workspace), tell L1 to reuse it (hash-verify) and go straight to the optimizer loop. ## The invariants that make it trustworthy (Codex-hardened — gate fidelity is load-bearing) - The gate is built by the **independent test-writer**, not the optimizer. - The reference is **CANONICAL** (validated against a trusted source), never re-derived. - The gate checks **INTERMEDIATES** (golden vectors per stage), not just a final cosine — a deep simplification (a dropped layernorm, a wrong rope base) is caught at the first diverging stage (`first_failing_stage`), not washed out. - Multiple metrics + edge cases (long-context, boundary positions). - The gate is **sha256 hash-locked**; L1 re-verifies the hash before EVERY round (tamper = abort). - The optimizer is judged ONLY by `gate/check.py`. **Codex** reviews each round on two axes — Integrity (did it simplify to pass?) + Plan (is the lever sound?). - **FINAL acceptance = in-MPK faithful build** (`MPK_FORCE_RDC_TRUE=1`, compiled into the real megakernel). A standalone number never ships. - **Early stop = round incomplete, never success.** ## Proof it works Re-run on the simplified attention: the optimizer (judged by the frozen gate) built a CORRECT real-DSv3 fused attention in round 1 — `GATE_RESULT {pass:true}` on all 5 cases / all intermediates, gate untouched, dump provenance verified (W_UV BMM present, real mscale²/sqrt(192) scale), all 4 simplifications fixed, Codex Integrity+Plan PASS. The old simplified-kernel failure is now structurally impossible. ## Reference Full design + the /cd-mechanics + the hardening rationale: `DESIGN.md` (bundled alongside this skill). The ferret runtime itself lives at `~/ferret/` (its `CLAUDE.md` is the optimizer's methodology). Pre-authored task specs: `~/ferret/tasks/`.
Ver no GitHub