Skip to main content
Ejecuta cualquier Skill en Manus
con un clic

steering-coefficient-tuning

Estrellas11
Forks1
Actualizado11 de julio de 2026 a las 04:09

How to set the strength of any additive intervention on internal representations — steering vector, CAA, DAS dose-response, representation engineering, SAE feature scaling, ROME-style edits. Use whenever the plan sets a steering strength (`α`, `β`, `dose`, `magnitude`, `scale`, `coefficient`, `k`) to a fixed value or small range, especially when copied from a paper. Covers a coarse `layer × β` sweep (β in σ_proj units, scored on a target metric plus a fluency / general-ability metric), why the best coefficient is layer-dependent, the mid- vs late-layer behavior (late layers break into repetition / format-spam), and the rule to use the smallest sufficient β. Prevents two symmetric failures: TOO SMALL → effect drowned in noise → false "no causal effect"; TOO LARGE → off-distribution collapse → fluency breaks / random direction matches it → false "specificity fails". Triggers include `dose ∈ {-3..3}`, `α = 3`, "random direction beat my steering vector", "steering had no effect", "model output garbage after steer

Instalación

Instalar con Codex o Claude Copia este prompt, pégalo en Codex, Claude u otro asistente, y deja que revise la página de la skill y la instale por ti.

SKILL.md
readonly