Skip to main content

compute-mamba-ratio

Compute the optimal --mamba-full-memory-ratio (or --max-mamba-cache-size pin) for a hybrid attention + linear-attention (Mamba / GDN / KDA) model's two serving memory pools, from the workload and serving config. Use when a user asks what ratio to set, why concurrency is clamped, or how to size the state vs KV pools for a hybrid model.

الانتقال إلى التثبيت

معلومات المصدر

المستودع
sgl-project/sglang
آخر نشاط في المصدر
٢٩ يوليو ٢٠٢٦ في ١٤:٣٨
لغة SKILL.md المكتشفة
الإنجليزية
النجوم
٣٢٬٠٠٠
التفرعات
٧٬٩٧٢

خيارات التثبيت

يُحدَّد Prompt الذي يراجع المصدر أولًا بشكل افتراضي. يمكنك التبديل إلى أمر مباشر أو تنزيل نسخة محلية.

مراجعة ملفات المصدر

اقرأ SKILL.md وأي ملفات مرافقة يعرضها SkillsMP قبل أن تقرر التثبيت.