Skip to main content
Manusで任意のスキルを実行
ワンクリックで

steering-coefficient-tuning

スター11
フォーク1
更新日2026年7月11日 04:09

How to set the strength of any additive intervention on internal representations — steering vector, CAA, DAS dose-response, representation engineering, SAE feature scaling, ROME-style edits. Use whenever the plan sets a steering strength (`α`, `β`, `dose`, `magnitude`, `scale`, `coefficient`, `k`) to a fixed value or small range, especially when copied from a paper. Covers a coarse `layer × β` sweep (β in σ_proj units, scored on a target metric plus a fluency / general-ability metric), why the best coefficient is layer-dependent, the mid- vs late-layer behavior (late layers break into repetition / format-spam), and the rule to use the smallest sufficient β. Prevents two symmetric failures: TOO SMALL → effect drowned in noise → false "no causal effect"; TOO LARGE → off-distribution collapse → fluency breaks / random direction matches it → false "specificity fails". Triggers include `dose ∈ {-3..3}`, `α = 3`, "random direction beat my steering vector", "steering had no effect", "model output garbage after steer

インストール

Codex または Claude でインストール この Prompt をコピーして Codex、Claude、または他のアシスタントに貼り付けると、Skill ページを確認してインストールできます。

SKILL.md
readonly