Skip to main content

feature-dictionary-learning

Feature Dictionary Learning methods address the polysemanticity of neuron-level units by decomposing a dense internal activation (e.g. a residual-stream state or MLP output) into a sparse weighted sum of directions drawn from a large over-complete dictionary. The dictionary contains far more directions than the activation's original dimensionality, and the decomposition is constrained to use only a small number of them at once. Each direction in the dictionary plays the role of an interpretable "feature": its weight measures how strongly that feature is present, turning a black-box vector into a small set of human-readable components.

الانتقال إلى التثبيت

معلومات المصدر

المستودع
zjunlp/Mechanist
آخر نشاط في المصدر
١١ يوليو ٢٠٢٦ في ٠٤:٠٩
لغة SKILL.md المكتشفة
الإنجليزية
النجوم
٥٠
التفرعات
٦

خيارات التثبيت

يُحدَّد Prompt الذي يراجع المصدر أولًا بشكل افتراضي. يمكنك التبديل إلى أمر مباشر أو تنزيل نسخة محلية.

مراجعة ملفات المصدر

اقرأ SKILL.md وأي ملفات مرافقة يعرضها SkillsMP قبل أن تقرر التثبيت.